The engine serves as the primary component that generates power and drives vehicle movement.Given its critical role,accurately diagnosing engine faults is essential for ensuring vehicle safety and reliability.Recent a...The engine serves as the primary component that generates power and drives vehicle movement.Given its critical role,accurately diagnosing engine faults is essential for ensuring vehicle safety and reliability.Recent advances in machine learning(ML)have enabled the development of artificial intelligence(AI)-based diagnostic models with strong predictive performance.However,the lack of transparency in these models constrains user confidence in their diagnostic outcomes.While explainable AI(XAI)methods such as local interpretable model-agnostic explanations(LIME)and Shapley additive explanations(SHAP)have been introduced to improve interpretability,their reliance on visual outputs requires manual interpretation,which can be inefficient and prone to subjectivity.To address this limitation,we propose DRIVE,a novel method for explainable vehicle engine fault diagnosis.In DRIVE,LIME and SHAP are applied to an ML-based diagnostic model,and their visual outputs are translated into textual explanations using the vision-language models(VLMs).These complementary explanations are then synthesized by a large language model(LLM)into a unified diagnostic report,providing a coherent narrative of the model’s reasoning and emphasizing abnormal input features.Experiments conducted on a publicly available vehicle engine fault dataset demonstrate that DRIVE not only produces accurate and transparent diagnostic rationales but also generates structured reports that enhance usability for domain experts.By integrating multiple XAI methods with multimodal LLMs,DRIVE advances the transparency,trustworthiness,and practicality of AI-driven vehicle engine fault diagnosis.展开更多
Amid global warming,mountainous regions have emerged as critical zones of investigation owing to their heightened vulnerability to climate change,their ecological significance,and the intensified interactions between ...Amid global warming,mountainous regions have emerged as critical zones of investigation owing to their heightened vulnerability to climate change,their ecological significance,and the intensified interactions between natural stress and human activities.Land surface temperature(LST)is a fundamental indicator for assessing climatic sensitivity in these landscapes.However,a comprehensive understanding of the spatiotemporal dynamics and driving mechanisms of LST across large mountainous regions remains limited.Therefore,data from the Terra Moderate Resolution Imaging Spectroradiometer Land Surface Temperature/Emissivity Daily(MOD11A1)Version 6.1 product during 2001–2020 in Yunnan Province(a complex mountainous region),China,were analyzed.Sen's slope analysis and Mann-Kendall test were applied to detect LST trends and spatial heterogeneity at both annual and seasonal scales.Subsequently,an eXtreme Gradient Boosting(XGBoost)model coupled with SHapley Additive exPlanations(SHAP)was employed to clarify the nonlinear contributions of multiple drivers.The study revealed the following findings.LST exhibited an overall warming rate of 0.020℃/a,characterized by daytime cooling(–0.008℃/a)and nighttime warming(0.048℃/a).LST increased during spring,summer,and autumn(0.011℃/a–0.018℃/a),whereas winter LST exhibited a cooling trend(–0.011℃/a).These variations were spatially partitioned by the Ailao Mountains,with the southwest displaying stronger thermal changes than the northeast.Natural controls,including digital elevation model(DEM)and downward shortwave radiation(DSR),predominated in the northwest high mountain canyons area and south tropical rainforest area,whereas nature–human interactions were more pronounced in the central urban agglomeration warming area and southeast karst landform area.The dominant drivers consisted of DEM,DSR,Normalized Difference Moisture Index,particulate matter 2.5(PM2.5),and aerosol optical depth(AOD).The strong correlations between gross domestic product and population density(correlation coefficient(r)=0.95),as well as between PM2.5 and AOD(r=0.84),highlighted the increasing influence of socioeconomic factors on surface warming.This study can advance the understanding of how mountain topography,moisture,and anthropogenic pressures jointly regulate surface thermal regimes and provide region-specific insights for climate adaptation and sustainable ecosystem management.展开更多
Accurate reservoir permeability determination is crucial in hydrocarbon exploration and production.Conventional methods relying on empirical correlations and assumptions often result in high costs,time consumption,ina...Accurate reservoir permeability determination is crucial in hydrocarbon exploration and production.Conventional methods relying on empirical correlations and assumptions often result in high costs,time consumption,inaccuracies,and uncertainties.This study introduces a novel hybrid machine learning approach to predict the permeability of the Wangkwar formation in the Gunya oilfield,Northwestern Uganda.The group method of data handling with differential evolution(GMDH-DE)algorithm was used to predict permeability due to its capability to manage complex,nonlinear relationships between variables,reduced computation time,and parameter optimization through evolutionary algorithms.Using 1953 samples from Gunya-1 and Gunya-2 wells for training and 1563 samples from Gunya-3 for testing,the GMDH-DE outperformed the group method of data handling(GMDH)and random forest(RF)in predicting permeability with higher accuracy and lower computation time.The GMDH-DE achieved an R2of 0.9985,RMSE of 3.157,MAE of 2.366,and ME of 0.001 during training,and for testing,the ME,MAE,RMSE,and R2were 1.3508,12.503,21.3898,and 0.9534,respectively.Additionally,the GMDH-DE demonstrated a 41%reduction in processing time compared to GMDH and RF.The model was also used to predict the permeability of the Mita Gamma well in the Mandawa basin,Tanzania,which lacks core data.Shapley additive explanations(SHAP)analysis identified thermal neutron porosity(TNPH),effective porosity(PHIE),and spectral gamma-ray(SGR)as the most critical parameters in permeability prediction.Therefore,the GMDH-DE model offers a novel,efficient,and accurate approach for fast permeability prediction,enhancing hydrocarbon exploration and production.展开更多
This study provides an in-depth comparative evaluation of landslide susceptibility using two distinct spatial units:and slope units(SUs)and hydrological response units(HRUs),within Goesan County,South Korea.Leveraging...This study provides an in-depth comparative evaluation of landslide susceptibility using two distinct spatial units:and slope units(SUs)and hydrological response units(HRUs),within Goesan County,South Korea.Leveraging the capabilities of the extreme gradient boosting(XGB)algorithm combined with Shapley Additive Explanations(SHAP),this work assesses the precision and clarity with which each unit predicts areas vulnerable to landslides.SUs focus on the geomorphological features like ridges and valleys,focusing on slope stability and landslide triggers.Conversely,HRUs are established based on a variety of hydrological factors,including land cover,soil type and slope gradients,to encapsulate the dynamic water processes of the region.The methodological framework includes the systematic gathering,preparation and analysis of data,ranging from historical landslide occurrences to topographical and environmental variables like elevation,slope angle and land curvature etc.The XGB algorithm used to construct the Landslide Susceptibility Model(LSM)was combined with SHAP for model interpretation and the results were evaluated using Random Cross-validation(RCV)to ensure accuracy and reliability.To ensure optimal model performance,the XGB algorithm’s hyperparameters were tuned using Differential Evolution,considering multicollinearity-free variables.The results show that SU and HRU are effective for LSM,but their effectiveness varies depending on landscape characteristics.The XGB algorithm demonstrates strong predictive power and SHAP enhances model transparency of the influential variables involved.This work underscores the importance of selecting appropriate assessment units tailored to specific landscape characteristics for accurate LSM.The integration of advanced machine learning techniques with interpretative tools offers a robust framework for landslide susceptibility assessment,improving both predictive capabilities and model interpretability.Future research should integrate broader data sets and explore hybrid analytical models to strengthen the generalizability of these findings across varied geographical settings.展开更多
BACKGROUND Diabetic foot ulcer(DFU)is a serious and destructive complication of diabetes,which has a high amputation rate and carries a huge social burden.Early detection of risk factors and intervention are essential...BACKGROUND Diabetic foot ulcer(DFU)is a serious and destructive complication of diabetes,which has a high amputation rate and carries a huge social burden.Early detection of risk factors and intervention are essential to reduce amputation rates.With the development of artificial intelligence technology,efficient interpretable predictive models can be generated in clinical practice to improve DFU care.AIM To develop and validate an interpretable model for predicting amputation risk in DFU patients.METHODS This retrospective study collected basic data from 599 patients with DFU in Beijing Shijitan Hospital between January 2015 and June 2024.The data set was randomly divided into a training set and test set with fivefold cross-validation.Three binary variable models were built with the eXtreme Gradient Boosting(XGBoost)algorithm to input risk factors that predict amputation probability.The model performance was optimized by adjusting the super parameters.The pre-dictive performance of the three models was expressed by sensitivity,specificity,positive predictive value,negative predictive value and area under the curve(AUC).Visualization of the prediction results was realized through SHapley Additive exPlanation(SHAP).RESULTS A total of 157(26.2%)patients underwent minor amputation during hospitalization and 50(8.3%)had major amputation.All three XGBoost models demonstrated good discriminative ability,with AUC values>0.7.The model for predicting major amputation achieved the highest performance[AUC=0.977,95%confidence interval(CI):0.956-0.998],followed by the minor amputation model(AUC=0.800,95%CI:0.762-0.838)and the non-amputation model(AUC=0.772,95%CI:0.730-0.814).Feature importance ranking of the three models revealed the risk factors for minor and major amputation.Wagner grade 4/5,osteomyelitis,and high C-reactive protein were all considered important predictive variables.CONCLUSION XGBoost effectively predicts diabetic foot amputation risk and provides interpretable insights to support person-alized treatment decisions.展开更多
The methods of network attacks have become increasingly sophisticated,rendering traditional cybersecurity defense mechanisms insufficient to address novel and complex threats effectively.In recent years,artificial int...The methods of network attacks have become increasingly sophisticated,rendering traditional cybersecurity defense mechanisms insufficient to address novel and complex threats effectively.In recent years,artificial intelligence has achieved significant progress in the field of network security.However,many challenges and issues remain,particularly regarding the interpretability of deep learning and ensemble learning algorithms.To address the challenge of enhancing the interpretability of network attack prediction models,this paper proposes a method that combines Light Gradient Boosting Machine(LGBM)and SHapley Additive exPlanations(SHAP).LGBM is employed to model anomalous fluctuations in various network indicators,enabling the rapid and accurate identification and prediction of potential network attack types,thereby facilitating the implementation of timely defense measures,the model achieved an accuracy of 0.977,precision of 0.985,recall of 0.975,and an F1 score of 0.979,demonstrating better performance compared to other models in the domain of network attack prediction.SHAP is utilized to analyze the black-box decision-making process of the model,providing interpretability by quantifying the contribution of each feature to the prediction results and elucidating the relationships between features.The experimental results demonstrate that the network attack predictionmodel based on LGBM exhibits superior accuracy and outstanding predictive capabilities.Moreover,the SHAP-based interpretability analysis significantly improves the model’s transparency and interpretability.展开更多
Predicting molecular properties is essential for advancing for advancing drug discovery and design. Recently, Graph Neural Networks (GNNs) have gained prominence due to their ability to capture the complex structural ...Predicting molecular properties is essential for advancing for advancing drug discovery and design. Recently, Graph Neural Networks (GNNs) have gained prominence due to their ability to capture the complex structural and relational information inherent in molecular graphs. Despite their effectiveness, the “black-box” nature of GNNs remains a significant obstacle to their widespread adoption in chemistry, as it hinders interpretability and trust. In this context, several explanation methods based on factual reasoning have emerged. These methods aim to interpret the predictions made by GNNs by analyzing the key features contributing to the prediction. However, these approaches fail to answer critical questions: “How to ensure that the structure-property mapping learned by GNNs is consistent with established domain knowledge”. In this paper, we propose MMGCF, a novel counterfactual explanation framework designed specifically for the prediction of GNN-based molecular properties. MMGCF constructs a hierarchical tree structure on molecular motifs, enabling the systematic generation of counterfactuals through motif perturbations. This framework identifies causally significant motifs and elucidates their impact on model predictions, offering insights into the relationship between structural modifications and predicted properties. Our method demonstrates its effectiveness through comprehensive quantitative and qualitative evaluations of four real-world molecular datasets.展开更多
Deep learning models have become a core technological tool in the field of medical image analysis.However,these models often suffer from a lack of transparency in their decision-making processes,leading to challenges ...Deep learning models have become a core technological tool in the field of medical image analysis.However,these models often suffer from a lack of transparency in their decision-making processes,leading to challenges related to trust and interpret ability in clinical applications.To address this issue,explainable artificial intelligence(XAI)techniques have been applied to medical image analysis.While showing promising potential,XAI also brings significant ethical risks in practice—most notably,the problem of spurious explanations.Such explanations may rise further concerns regarding patient privacy,data security,and the attribution of decisionmaking authority in medical contexts.This paper analyzes the application of XAI methods—particularly saliency aps—in medical image interpretation,identifies the underlying causes of spurious explanations,and proposes possible mitigation strategies.The aim is to contribute to the responsible and sustainable integration of explainable AI into clinical practice.展开更多
Accurate prediction of shield tunneling-induced settlement is a complex problem that requires consideration of many influential parameters.Recent studies reveal that machine learning(ML)algorithms can predict the sett...Accurate prediction of shield tunneling-induced settlement is a complex problem that requires consideration of many influential parameters.Recent studies reveal that machine learning(ML)algorithms can predict the settlement caused by tunneling.However,well-performing ML models are usually less interpretable.Irrelevant input features decrease the performance and interpretability of an ML model.Nonetheless,feature selection,a critical step in the ML pipeline,is usually ignored in most studies that focused on predicting tunneling-induced settlement.This study applies four techniques,i.e.Pearson correlation method,sequential forward selection(SFS),sequential backward selection(SBS)and Boruta algorithm,to investigate the effect of feature selection on the model’s performance when predicting the tunneling-induced maximum surface settlement(Smax).The data set used in this study was compiled from two metro tunnel projects excavated in Hangzhou,China using earth pressure balance(EPB)shields and consists of 14 input features and a single output(i.e.Smax).The ML model that is trained on features selected from the Boruta algorithm demonstrates the best performance in both the training and testing phases.The relevant features chosen from the Boruta algorithm further indicate that tunneling-induced settlement is affected by parameters related to tunnel geometry,geological conditions and shield operation.The recently proposed Shapley additive explanations(SHAP)method explores how the input features contribute to the output of a complex ML model.It is observed that the larger settlements are induced during shield tunneling in silty clay.Moreover,the SHAP analysis reveals that the low magnitudes of face pressure at the top of the shield increase the model’s output。展开更多
The flow regimes of GLCC with horizon inlet and a vertical pipe are investigated in experiments,and the velocities and pressure drops data labeled by the corresponding flow regimes are collected.Combined with the flow...The flow regimes of GLCC with horizon inlet and a vertical pipe are investigated in experiments,and the velocities and pressure drops data labeled by the corresponding flow regimes are collected.Combined with the flow regimes data of other GLCC positions from other literatures in existence,the gas and liquid superficial velocities and pressure drops are used as the input of the machine learning algorithms respectively which are applied to identify the flow regimes.The choosing of input data types takes the availability of data for practical industry fields into consideration,and the twelve machine learning algorithms are chosen from the classical and popular algorithms in the area of classification,including the typical ensemble models,SVM,KNN,Bayesian Model and MLP.The results of flow regimes identification show that gas and liquid superficial velocities are the ideal type of input data for the flow regimes identification by machine learning.Most of the ensemble models can identify the flow regimes of GLCC by gas and liquid velocities with the accuracy of 0.99 and more.For the pressure drops as the input of each algorithm,it is not the suitable as gas and liquid velocities,and only XGBoost and Bagging Tree can identify the GLCC flow regimes accurately.The success and confusion of each algorithm are analyzed and explained based on the experimental phenomena of flow regimes evolution processes,the flow regimes map,and the principles of algorithms.The applicability and feasibility of each algorithm according to different types of data for GLCC flow regimes identification are proposed.展开更多
Collaborative Filtering(CF) is a leading approach to build recommender systems which has gained considerable development and popularity. A predominant approach to CF is rating prediction recommender algorithm, aiming ...Collaborative Filtering(CF) is a leading approach to build recommender systems which has gained considerable development and popularity. A predominant approach to CF is rating prediction recommender algorithm, aiming to predict a user's rating for those items which were not rated yet by the user. However, with the increasing number of items and users, thedata is sparse.It is difficult to detectlatent closely relation among the items or users for predicting the user behaviors. In this paper,we enhance the rating prediction approach leading to substantial improvement of prediction accuracy by categorizing according to the genres of movies. Then the probabilities that users are interested in the genres are computed to integrate the prediction of each genre cluster. A novel probabilistic approach based on the sentiment analysis of the user reviews is also proposed to give intuitional explanations of why an item is recommended.To test the novel recommendation approach, a new corpus of user reviews on movies obtained from the Internet Movies Database(IMDB) has been generated. Experimental results show that the proposed framework is effective and achieves a better prediction performance.展开更多
The Kirk test has good precision for measuring stray light in optical lithography and is the usual method of measuring stray light.However,Kirk did not provide a theoretical explanation to his simulation model.We atte...The Kirk test has good precision for measuring stray light in optical lithography and is the usual method of measuring stray light.However,Kirk did not provide a theoretical explanation to his simulation model.We attempt to give Kirk's model a kind of theoretical explanation and a little improvement based on the model of point spread function of scattering and the theory of statistical optics.It is indicated by simulation that the improved model fits Kirk's measurement data better.展开更多
Finding an attribute to explain the relationships between a given pair of entities is valuable in many applications.However,many direct solutions fail,owing to its low precision caused by heavy dependence on text and ...Finding an attribute to explain the relationships between a given pair of entities is valuable in many applications.However,many direct solutions fail,owing to its low precision caused by heavy dependence on text and low recall by evidence scarcity.Thus,we propose a generalization-and-inference framework and implement it to build a system:entity-relationship finder(ERF).Our main idea is conceptualizing entity pairs into proper concept pairs,as intermediate random variables to form the explanation.Although entity conceptualization has been studied,it has new challenges of collective optimization for multiple relationship instances,joint optimization for both entities,and aggregation of diluted observations into the head concepts defining the relationship.We propose conceptualization solutions and validate them as well as the framework with extensive experiments.展开更多
In the letter to the editor, Dr. Comings et al. proposed a potential explanation of our findings that the L allele rather than S allele of 5-HTTLPR was associated with higher anxiety levels and reduced amygdala-prefro...In the letter to the editor, Dr. Comings et al. proposed a potential explanation of our findings that the L allele rather than S allele of 5-HTTLPR was associated with higher anxiety levels and reduced amygdala-prefrontal cortex (PFC) connectivity in Han Chinese[1], which demonstrated an 'allele reversal' in the genetics of the 5-HTTLPR gene in Asians versus Caucasians. The authors alleged that this 'allele reversal' might simply result from maternal age and suggested that we test this on our datasets. Unfortunately,展开更多
Existing explanation methods for Convolutional Neural Networks(CNNs)lack the pixel-level visualization explanations to generate the reliable fine-grained decision features.Since there are inconsistencies between the e...Existing explanation methods for Convolutional Neural Networks(CNNs)lack the pixel-level visualization explanations to generate the reliable fine-grained decision features.Since there are inconsistencies between the explanation and the actual behavior of the model to be interpreted,we propose a Fine-Grained Visual Explanation for CNN,namely F-GVE,which produces a fine-grained explanation with higher consistency to the decision of the original model.The exact backward class-specific gradients with respect to the input image is obtained to highlight the object-related pixels the model used to make prediction.In addition,for better visualization and less noise,F-GVE selects an appropriate threshold to filter the gradient during the calculation and the explanation map is obtained by element-wise multiplying the gradient and the input image to show fine-grained classification decision features.Experimental results demonstrate that F-GVE has good visual performances and highlights the importance of fine-grained decision features.Moreover,the faithfulness of the explanation in this paper is high and it is effective and practical on troubleshooting and debugging detection.展开更多
Supernova 1987 A is a core collapse supernova in the Large Magellanic Cloud, inside which the product is most likely a neutron star. Despite the most sensitive available detection instruments from radio to γ-ray wave...Supernova 1987 A is a core collapse supernova in the Large Magellanic Cloud, inside which the product is most likely a neutron star. Despite the most sensitive available detection instruments from radio to γ-ray wavebands being exploited in the pass thirty years, there have not yet been any pulse signals detected. By considering the density of the medium plasma in the remnant of 1987 A, we find that the plasma cut-off frequency is approximately7 GHz, a value higher than the conventional observational waveband of radio pulsars. As derived, with the expansion of the supernova remnant, the radio signal will be detected in 2073 A.D. at 3 GHz.展开更多
Majorana zero modes in the hybrid semiconductor-superconductornanowire is one of the promising candidates for topologicalquantum computing. Recently, in nanowires with a superconductingisland, the signature of Majoran...Majorana zero modes in the hybrid semiconductor-superconductornanowire is one of the promising candidates for topologicalquantum computing. Recently, in nanowires with a superconductingisland, the signature of Majorana zero modescan be revealed as a subgap state whose energy oscillatesaround zero in magnetic field. This oscillation was interpretedas overlapping Majoranas. However, the oscillation amplitudeeither dies away after an overshoot or decays, sharply oppositeto the theoretically predicted enhanced oscillations for Majoranabound states, as the magnetic field increases. Several theoreticalstudies have tried to address this discrepancy, but arepartially successful. This discrepancy has raised the concernson the conclusive identification of Majorana bound states, andhas even endangered the scheme of Majorana qubits basedon the nanowires.展开更多
Recently,convolutional neural network(CNN)-based visual inspec-tion has been developed to detect defects on building surfaces automatically.The CNN model demonstrates remarkable accuracy in image data analysis;however...Recently,convolutional neural network(CNN)-based visual inspec-tion has been developed to detect defects on building surfaces automatically.The CNN model demonstrates remarkable accuracy in image data analysis;however,the predicted results have uncertainty in providing accurate informa-tion to users because of the“black box”problem in the deep learning model.Therefore,this study proposes a visual explanation method to overcome the uncertainty limitation of CNN-based defect identification.The visual repre-sentative gradient-weights class activation mapping(Grad-CAM)method is adopted to provide visually explainable information.A visualizing evaluation index is proposed to quantitatively analyze visual representations;this index reflects a rough estimate of the concordance rate between the visualized heat map and intended defects.In addition,an ablation study,adopting three-branch combinations with the VGG16,is implemented to identify perfor-mance variations by visualizing predicted results.Experiments reveal that the proposed model,combined with hybrid pooling,batch normalization,and multi-attention modules,achieves the best performance with an accuracy of 97.77%,corresponding to an improvement of 2.49%compared with the baseline model.Consequently,this study demonstrates that reliable results from an automatic defect classification model can be provided to an inspector through the visual representation of the predicted results using CNN models.展开更多
基金supported by the National Research Foundation of Korea(NRF)grant funded by the Korea government(MSIT)(RS-2025-00516023)supported by Korea Institute of Planning and Evaluation for Technology in Food,Agriculture and Forestry(IPET)through the High Value-added Food Technology Development Program,funded by the Ministry of Agriculture,Food and Rural Affairs(MAFRA)(RS-2024-00403286).
摘要The engine serves as the primary component that generates power and drives vehicle movement.Given its critical role,accurately diagnosing engine faults is essential for ensuring vehicle safety and reliability.Recent advances in machine learning(ML)have enabled the development of artificial intelligence(AI)-based diagnostic models with strong predictive performance.However,the lack of transparency in these models constrains user confidence in their diagnostic outcomes.While explainable AI(XAI)methods such as local interpretable model-agnostic explanations(LIME)and Shapley additive explanations(SHAP)have been introduced to improve interpretability,their reliance on visual outputs requires manual interpretation,which can be inefficient and prone to subjectivity.To address this limitation,we propose DRIVE,a novel method for explainable vehicle engine fault diagnosis.In DRIVE,LIME and SHAP are applied to an ML-based diagnostic model,and their visual outputs are translated into textual explanations using the vision-language models(VLMs).These complementary explanations are then synthesized by a large language model(LLM)into a unified diagnostic report,providing a coherent narrative of the model’s reasoning and emphasizing abnormal input features.Experiments conducted on a publicly available vehicle engine fault dataset demonstrate that DRIVE not only produces accurate and transparent diagnostic rationales but also generates structured reports that enhance usability for domain experts.By integrating multiple XAI methods with multimodal LLMs,DRIVE advances the transparency,trustworthiness,and practicality of AI-driven vehicle engine fault diagnosis.
基金supported by the National Natural Science Foundation of China(42061004)the Youth Special Project of Xing Dian Talent Support Program of Yunnan Province,China(XDYC-QNRC-2022-0230)+1 种基金the Open Subjects of First-class Disciplines in Soil and Water Conservation and Desertification Control in Yunnan Province,China(SBK20240021)the Special Project for Building a Science and Technology Innovation Center for South and Southeast Asia,China(202503AP140004)。
摘要Amid global warming,mountainous regions have emerged as critical zones of investigation owing to their heightened vulnerability to climate change,their ecological significance,and the intensified interactions between natural stress and human activities.Land surface temperature(LST)is a fundamental indicator for assessing climatic sensitivity in these landscapes.However,a comprehensive understanding of the spatiotemporal dynamics and driving mechanisms of LST across large mountainous regions remains limited.Therefore,data from the Terra Moderate Resolution Imaging Spectroradiometer Land Surface Temperature/Emissivity Daily(MOD11A1)Version 6.1 product during 2001–2020 in Yunnan Province(a complex mountainous region),China,were analyzed.Sen's slope analysis and Mann-Kendall test were applied to detect LST trends and spatial heterogeneity at both annual and seasonal scales.Subsequently,an eXtreme Gradient Boosting(XGBoost)model coupled with SHapley Additive exPlanations(SHAP)was employed to clarify the nonlinear contributions of multiple drivers.The study revealed the following findings.LST exhibited an overall warming rate of 0.020℃/a,characterized by daytime cooling(–0.008℃/a)and nighttime warming(0.048℃/a).LST increased during spring,summer,and autumn(0.011℃/a–0.018℃/a),whereas winter LST exhibited a cooling trend(–0.011℃/a).These variations were spatially partitioned by the Ailao Mountains,with the southwest displaying stronger thermal changes than the northeast.Natural controls,including digital elevation model(DEM)and downward shortwave radiation(DSR),predominated in the northwest high mountain canyons area and south tropical rainforest area,whereas nature–human interactions were more pronounced in the central urban agglomeration warming area and southeast karst landform area.The dominant drivers consisted of DEM,DSR,Normalized Difference Moisture Index,particulate matter 2.5(PM2.5),and aerosol optical depth(AOD).The strong correlations between gross domestic product and population density(correlation coefficient(r)=0.95),as well as between PM2.5 and AOD(r=0.84),highlighted the increasing influence of socioeconomic factors on surface warming.This study can advance the understanding of how mountain topography,moisture,and anthropogenic pressures jointly regulate surface thermal regimes and provide region-specific insights for climate adaptation and sustainable ecosystem management.
基金supported by the Major National Science and Technology Programs in the“Thirteenth Five-Year”Plan period(Grant No.2017ZX05032-002-004)the Innovation Team Funding of Natural Science Foundation of Hubei Province,China(Grant No.2021CFA031)the Chinese Scholarship Council(CSC)and Silk Road Institute for their support in terms of stipend.
摘要Accurate reservoir permeability determination is crucial in hydrocarbon exploration and production.Conventional methods relying on empirical correlations and assumptions often result in high costs,time consumption,inaccuracies,and uncertainties.This study introduces a novel hybrid machine learning approach to predict the permeability of the Wangkwar formation in the Gunya oilfield,Northwestern Uganda.The group method of data handling with differential evolution(GMDH-DE)algorithm was used to predict permeability due to its capability to manage complex,nonlinear relationships between variables,reduced computation time,and parameter optimization through evolutionary algorithms.Using 1953 samples from Gunya-1 and Gunya-2 wells for training and 1563 samples from Gunya-3 for testing,the GMDH-DE outperformed the group method of data handling(GMDH)and random forest(RF)in predicting permeability with higher accuracy and lower computation time.The GMDH-DE achieved an R2of 0.9985,RMSE of 3.157,MAE of 2.366,and ME of 0.001 during training,and for testing,the ME,MAE,RMSE,and R2were 1.3508,12.503,21.3898,and 0.9534,respectively.Additionally,the GMDH-DE demonstrated a 41%reduction in processing time compared to GMDH and RF.The model was also used to predict the permeability of the Mita Gamma well in the Mandawa basin,Tanzania,which lacks core data.Shapley additive explanations(SHAP)analysis identified thermal neutron porosity(TNPH),effective porosity(PHIE),and spectral gamma-ray(SGR)as the most critical parameters in permeability prediction.Therefore,the GMDH-DE model offers a novel,efficient,and accurate approach for fast permeability prediction,enhancing hydrocarbon exploration and production.
基金supported by a National Research Foundation of Korea(NRF)grant funded by the Korean government(MSIT)(RS-2023-00222536).
摘要This study provides an in-depth comparative evaluation of landslide susceptibility using two distinct spatial units:and slope units(SUs)and hydrological response units(HRUs),within Goesan County,South Korea.Leveraging the capabilities of the extreme gradient boosting(XGB)algorithm combined with Shapley Additive Explanations(SHAP),this work assesses the precision and clarity with which each unit predicts areas vulnerable to landslides.SUs focus on the geomorphological features like ridges and valleys,focusing on slope stability and landslide triggers.Conversely,HRUs are established based on a variety of hydrological factors,including land cover,soil type and slope gradients,to encapsulate the dynamic water processes of the region.The methodological framework includes the systematic gathering,preparation and analysis of data,ranging from historical landslide occurrences to topographical and environmental variables like elevation,slope angle and land curvature etc.The XGB algorithm used to construct the Landslide Susceptibility Model(LSM)was combined with SHAP for model interpretation and the results were evaluated using Random Cross-validation(RCV)to ensure accuracy and reliability.To ensure optimal model performance,the XGB algorithm’s hyperparameters were tuned using Differential Evolution,considering multicollinearity-free variables.The results show that SU and HRU are effective for LSM,but their effectiveness varies depending on landscape characteristics.The XGB algorithm demonstrates strong predictive power and SHAP enhances model transparency of the influential variables involved.This work underscores the importance of selecting appropriate assessment units tailored to specific landscape characteristics for accurate LSM.The integration of advanced machine learning techniques with interpretative tools offers a robust framework for landslide susceptibility assessment,improving both predictive capabilities and model interpretability.Future research should integrate broader data sets and explore hybrid analytical models to strengthen the generalizability of these findings across varied geographical settings.
摘要BACKGROUND Diabetic foot ulcer(DFU)is a serious and destructive complication of diabetes,which has a high amputation rate and carries a huge social burden.Early detection of risk factors and intervention are essential to reduce amputation rates.With the development of artificial intelligence technology,efficient interpretable predictive models can be generated in clinical practice to improve DFU care.AIM To develop and validate an interpretable model for predicting amputation risk in DFU patients.METHODS This retrospective study collected basic data from 599 patients with DFU in Beijing Shijitan Hospital between January 2015 and June 2024.The data set was randomly divided into a training set and test set with fivefold cross-validation.Three binary variable models were built with the eXtreme Gradient Boosting(XGBoost)algorithm to input risk factors that predict amputation probability.The model performance was optimized by adjusting the super parameters.The pre-dictive performance of the three models was expressed by sensitivity,specificity,positive predictive value,negative predictive value and area under the curve(AUC).Visualization of the prediction results was realized through SHapley Additive exPlanation(SHAP).RESULTS A total of 157(26.2%)patients underwent minor amputation during hospitalization and 50(8.3%)had major amputation.All three XGBoost models demonstrated good discriminative ability,with AUC values>0.7.The model for predicting major amputation achieved the highest performance[AUC=0.977,95%confidence interval(CI):0.956-0.998],followed by the minor amputation model(AUC=0.800,95%CI:0.762-0.838)and the non-amputation model(AUC=0.772,95%CI:0.730-0.814).Feature importance ranking of the three models revealed the risk factors for minor and major amputation.Wagner grade 4/5,osteomyelitis,and high C-reactive protein were all considered important predictive variables.CONCLUSION XGBoost effectively predicts diabetic foot amputation risk and provides interpretable insights to support person-alized treatment decisions.
基金supported by the National Natural Science Foundation of China Project(No.62302540)please visit their website at http://gffzzf112c495998e46desbcvbvpxvu9qp6o6c.ffgz.tsg.suse.edu.cn/(accessed on 18 June 2024).
摘要The methods of network attacks have become increasingly sophisticated,rendering traditional cybersecurity defense mechanisms insufficient to address novel and complex threats effectively.In recent years,artificial intelligence has achieved significant progress in the field of network security.However,many challenges and issues remain,particularly regarding the interpretability of deep learning and ensemble learning algorithms.To address the challenge of enhancing the interpretability of network attack prediction models,this paper proposes a method that combines Light Gradient Boosting Machine(LGBM)and SHapley Additive exPlanations(SHAP).LGBM is employed to model anomalous fluctuations in various network indicators,enabling the rapid and accurate identification and prediction of potential network attack types,thereby facilitating the implementation of timely defense measures,the model achieved an accuracy of 0.977,precision of 0.985,recall of 0.975,and an F1 score of 0.979,demonstrating better performance compared to other models in the domain of network attack prediction.SHAP is utilized to analyze the black-box decision-making process of the model,providing interpretability by quantifying the contribution of each feature to the prediction results and elucidating the relationships between features.The experimental results demonstrate that the network attack predictionmodel based on LGBM exhibits superior accuracy and outstanding predictive capabilities.Moreover,the SHAP-based interpretability analysis significantly improves the model’s transparency and interpretability.
摘要Predicting molecular properties is essential for advancing for advancing drug discovery and design. Recently, Graph Neural Networks (GNNs) have gained prominence due to their ability to capture the complex structural and relational information inherent in molecular graphs. Despite their effectiveness, the “black-box” nature of GNNs remains a significant obstacle to their widespread adoption in chemistry, as it hinders interpretability and trust. In this context, several explanation methods based on factual reasoning have emerged. These methods aim to interpret the predictions made by GNNs by analyzing the key features contributing to the prediction. However, these approaches fail to answer critical questions: “How to ensure that the structure-property mapping learned by GNNs is consistent with established domain knowledge”. In this paper, we propose MMGCF, a novel counterfactual explanation framework designed specifically for the prediction of GNN-based molecular properties. MMGCF constructs a hierarchical tree structure on molecular motifs, enabling the systematic generation of counterfactuals through motif perturbations. This framework identifies causally significant motifs and elucidates their impact on model predictions, offering insights into the relationship between structural modifications and predicted properties. Our method demonstrates its effectiveness through comprehensive quantitative and qualitative evaluations of four real-world molecular datasets.
摘要Deep learning models have become a core technological tool in the field of medical image analysis.However,these models often suffer from a lack of transparency in their decision-making processes,leading to challenges related to trust and interpret ability in clinical applications.To address this issue,explainable artificial intelligence(XAI)techniques have been applied to medical image analysis.While showing promising potential,XAI also brings significant ethical risks in practice—most notably,the problem of spurious explanations.Such explanations may rise further concerns regarding patient privacy,data security,and the attribution of decisionmaking authority in medical contexts.This paper analyzes the application of XAI methods—particularly saliency aps—in medical image interpretation,identifies the underlying causes of spurious explanations,and proposes possible mitigation strategies.The aim is to contribute to the responsible and sustainable integration of explainable AI into clinical practice.
基金support provided by The Science and Technology Development Fund,Macao SAR,China(File Nos.0057/2020/AGJ and SKL-IOTSC-2021-2023)Science and Technology Program of Guangdong Province,China(Grant No.2021A0505080009).
摘要Accurate prediction of shield tunneling-induced settlement is a complex problem that requires consideration of many influential parameters.Recent studies reveal that machine learning(ML)algorithms can predict the settlement caused by tunneling.However,well-performing ML models are usually less interpretable.Irrelevant input features decrease the performance and interpretability of an ML model.Nonetheless,feature selection,a critical step in the ML pipeline,is usually ignored in most studies that focused on predicting tunneling-induced settlement.This study applies four techniques,i.e.Pearson correlation method,sequential forward selection(SFS),sequential backward selection(SBS)and Boruta algorithm,to investigate the effect of feature selection on the model’s performance when predicting the tunneling-induced maximum surface settlement(Smax).The data set used in this study was compiled from two metro tunnel projects excavated in Hangzhou,China using earth pressure balance(EPB)shields and consists of 14 input features and a single output(i.e.Smax).The ML model that is trained on features selected from the Boruta algorithm demonstrates the best performance in both the training and testing phases.The relevant features chosen from the Boruta algorithm further indicate that tunneling-induced settlement is affected by parameters related to tunnel geometry,geological conditions and shield operation.The recently proposed Shapley additive explanations(SHAP)method explores how the input features contribute to the output of a complex ML model.It is observed that the larger settlements are induced during shield tunneling in silty clay.Moreover,the SHAP analysis reveals that the low magnitudes of face pressure at the top of the shield increase the model’s output。
摘要The flow regimes of GLCC with horizon inlet and a vertical pipe are investigated in experiments,and the velocities and pressure drops data labeled by the corresponding flow regimes are collected.Combined with the flow regimes data of other GLCC positions from other literatures in existence,the gas and liquid superficial velocities and pressure drops are used as the input of the machine learning algorithms respectively which are applied to identify the flow regimes.The choosing of input data types takes the availability of data for practical industry fields into consideration,and the twelve machine learning algorithms are chosen from the classical and popular algorithms in the area of classification,including the typical ensemble models,SVM,KNN,Bayesian Model and MLP.The results of flow regimes identification show that gas and liquid superficial velocities are the ideal type of input data for the flow regimes identification by machine learning.Most of the ensemble models can identify the flow regimes of GLCC by gas and liquid velocities with the accuracy of 0.99 and more.For the pressure drops as the input of each algorithm,it is not the suitable as gas and liquid velocities,and only XGBoost and Bagging Tree can identify the GLCC flow regimes accurately.The success and confusion of each algorithm are analyzed and explained based on the experimental phenomena of flow regimes evolution processes,the flow regimes map,and the principles of algorithms.The applicability and feasibility of each algorithm according to different types of data for GLCC flow regimes identification are proposed.
基金supported in part by National Science Foundation of China under Grants No.61303105 and 61402304the Humanity&Social Science general project of Ministry of Education under Grants No.14YJAZH046+2 种基金the Beijing Natural Science Foundation under Grants No.4154065the Beijing Educational Committee Science and Technology Development Planned under Grants No.KM201410028017Academic Degree Graduate Courses group projects
摘要Collaborative Filtering(CF) is a leading approach to build recommender systems which has gained considerable development and popularity. A predominant approach to CF is rating prediction recommender algorithm, aiming to predict a user's rating for those items which were not rated yet by the user. However, with the increasing number of items and users, thedata is sparse.It is difficult to detectlatent closely relation among the items or users for predicting the user behaviors. In this paper,we enhance the rating prediction approach leading to substantial improvement of prediction accuracy by categorizing according to the genres of movies. Then the probabilities that users are interested in the genres are computed to integrate the prediction of each genre cluster. A novel probabilistic approach based on the sentiment analysis of the user reviews is also proposed to give intuitional explanations of why an item is recommended.To test the novel recommendation approach, a new corpus of user reviews on movies obtained from the Internet Movies Database(IMDB) has been generated. Experimental results show that the proposed framework is effective and achieves a better prediction performance.
基金by the National Basic Research Program of China under Grant No 2007AA01Z333the National Special Program of China under Grant No 2009ZX02204-008.
摘要The Kirk test has good precision for measuring stray light in optical lithography and is the usual method of measuring stray light.However,Kirk did not provide a theoretical explanation to his simulation model.We attempt to give Kirk's model a kind of theoretical explanation and a little improvement based on the model of point spread function of scattering and the theory of statistical optics.It is indicated by simulation that the improved model fits Kirk's measurement data better.
基金the Shanghai Science and Technology Innovation Action Plan(No.19511120400)the National Key Research and Development Project(No.2020AAA0109302)the Shanghai Municipal Science and Technology Major Project(No.2021SHZDZX0103)。
摘要Finding an attribute to explain the relationships between a given pair of entities is valuable in many applications.However,many direct solutions fail,owing to its low precision caused by heavy dependence on text and low recall by evidence scarcity.Thus,we propose a generalization-and-inference framework and implement it to build a system:entity-relationship finder(ERF).Our main idea is conceptualizing entity pairs into proper concept pairs,as intermediate random variables to form the explanation.Although entity conceptualization has been studied,it has new challenges of collective optimization for multiple relationship instances,joint optimization for both entities,and aggregation of diluted observations into the head concepts defining the relationship.We propose conceptualization solutions and validate them as well as the framework with extensive experiments.
摘要In the letter to the editor, Dr. Comings et al. proposed a potential explanation of our findings that the L allele rather than S allele of 5-HTTLPR was associated with higher anxiety levels and reduced amygdala-prefrontal cortex (PFC) connectivity in Han Chinese[1], which demonstrated an 'allele reversal' in the genetics of the 5-HTTLPR gene in Asians versus Caucasians. The authors alleged that this 'allele reversal' might simply result from maternal age and suggested that we test this on our datasets. Unfortunately,
基金This work was partially supported by Beijing Natural Science Foundation(No.4222038)by Open Research Project of the State Key Laboratory of Media Convergence and Communication(Communication University of China),by the National Key RD Program of China(No.2021YFF0307600)and by Fundamental Research Funds for the Central Universities.
摘要Existing explanation methods for Convolutional Neural Networks(CNNs)lack the pixel-level visualization explanations to generate the reliable fine-grained decision features.Since there are inconsistencies between the explanation and the actual behavior of the model to be interpreted,we propose a Fine-Grained Visual Explanation for CNN,namely F-GVE,which produces a fine-grained explanation with higher consistency to the decision of the original model.The exact backward class-specific gradients with respect to the input image is obtained to highlight the object-related pixels the model used to make prediction.In addition,for better visualization and less noise,F-GVE selects an appropriate threshold to filter the gradient during the calculation and the explanation map is obtained by element-wise multiplying the gradient and the input image to show fine-grained classification decision features.Experimental results demonstrate that F-GVE has good visual performances and highlights the importance of fine-grained decision features.Moreover,the faithfulness of the explanation in this paper is high and it is effective and practical on troubleshooting and debugging detection.
基金Supported by the National Basic Research Program of China under Grant No 2015CB857100the National Key Research and Development Program of China under Grant No 2017YFA0402600the National Natural Science Foundation of China under Grant Nos 11173034,11703003 and U1731238
摘要Supernova 1987 A is a core collapse supernova in the Large Magellanic Cloud, inside which the product is most likely a neutron star. Despite the most sensitive available detection instruments from radio to γ-ray wavebands being exploited in the pass thirty years, there have not yet been any pulse signals detected. By considering the density of the medium plasma in the remnant of 1987 A, we find that the plasma cut-off frequency is approximately7 GHz, a value higher than the conventional observational waveband of radio pulsars. As derived, with the expansion of the supernova remnant, the radio signal will be detected in 2073 A.D. at 3 GHz.
摘要Majorana zero modes in the hybrid semiconductor-superconductornanowire is one of the promising candidates for topologicalquantum computing. Recently, in nanowires with a superconductingisland, the signature of Majorana zero modescan be revealed as a subgap state whose energy oscillatesaround zero in magnetic field. This oscillation was interpretedas overlapping Majoranas. However, the oscillation amplitudeeither dies away after an overshoot or decays, sharply oppositeto the theoretically predicted enhanced oscillations for Majoranabound states, as the magnetic field increases. Several theoreticalstudies have tried to address this discrepancy, but arepartially successful. This discrepancy has raised the concernson the conclusive identification of Majorana bound states, andhas even endangered the scheme of Majorana qubits basedon the nanowires.
基金supported by a Korea Agency for Infrastructure Technology Advancement(KAIA)grant funded by the Ministry of Land,Infrastructure,and Transport(Grant 22CTAP-C163951-02).
摘要Recently,convolutional neural network(CNN)-based visual inspec-tion has been developed to detect defects on building surfaces automatically.The CNN model demonstrates remarkable accuracy in image data analysis;however,the predicted results have uncertainty in providing accurate informa-tion to users because of the“black box”problem in the deep learning model.Therefore,this study proposes a visual explanation method to overcome the uncertainty limitation of CNN-based defect identification.The visual repre-sentative gradient-weights class activation mapping(Grad-CAM)method is adopted to provide visually explainable information.A visualizing evaluation index is proposed to quantitatively analyze visual representations;this index reflects a rough estimate of the concordance rate between the visualized heat map and intended defects.In addition,an ablation study,adopting three-branch combinations with the VGG16,is implemented to identify perfor-mance variations by visualizing predicted results.Experiments reveal that the proposed model,combined with hybrid pooling,batch normalization,and multi-attention modules,achieves the best performance with an accuracy of 97.77%,corresponding to an improvement of 2.49%compared with the baseline model.Consequently,this study demonstrates that reliable results from an automatic defect classification model can be provided to an inspector through the visual representation of the predicted results using CNN models.