Objectives To develop and validate a machine learning(ML)model for real-time monitoring of in-hospital mortality(IHM)risk and identification of major risk factors in heart failure(HF)patients admitted to intensive car...Objectives To develop and validate a machine learning(ML)model for real-time monitoring of in-hospital mortality(IHM)risk and identification of major risk factors in heart failure(HF)patients admitted to intensive care units(ICUs).Methods Data from ICU HF patients were extracted from the multicenter eICU-Collaborative Research Database(eICU-CRD.External validation used MIMIC-IV and a real-world Chinese dataset(CHN-dataset).Daily measurements from MIMIC-IV patients staying≥3 days formed a Daily Measurement(DM)dataset.After rigorous preprocessing and feature selection,five ML algorithms were trained and optimized using eICU-CRD data.Model performance was evaluated using AUC,sensitivity,specificity,and balanced accuracy.The optimal model was benchmarked against APACHE and SOFA scores.SHapley Additive ex-Planations(SHAP)interpreted feature contributions.A Windows application was developed for clinical deployment.Results XGBoost emerged as the optimal model(Final-ML model),which incorporated only 17 routinely collected clinical variables:age,non-invasive systolic blood pressure(NI-SBP),heart rate,respiratory rate,Glasgow Coma Scale eye opening score,white blood cell count(WBC),creatinine,bicarbonate,red cell distribution width(RDW),platelet count,glucose,calcium,mean corpuscular hemoglobin concentration(MCHC),sodium,mean corpuscular volume,red blood cell count,and potassium.It achieved high AUCs:0.876(95%CI:0.836-0.915;eICU-CRD test data),0.932(95%CI:0.921-0.942;MIMIC-IV),and 0.879(95%CI:0.846-0.912;CHN-dataset).It significantly outperformed APACHE(AUC=0.740,95%CI:0.720-0.761)and SOFA(AUC=0.717,95%CI:0.694-0.740)scores.The model demonstrated strong generalizability across ethnicities,ward types,and genders within MIMIC-IV.Using daily data(DM dataset),predicted IHM risk accurately tracked patient trajectories:risk decreased progressively for survivors and increased for non-survivors throughout the ICU stay.SHAP analysis identified key predictors:NI-SBP,age,heart rate,WBC,glucose,and notably,RDW and MCHC.Time-dependent Cox regression confirmed RDW increase(HR=3.783,95%CI:2.237-6.398)and MCHC decrease(HR=0.173,95%CI:0.040-0.741)as significant independent risk factors for IHM.Conclusions The developed XGBoost model provides a reliable,generalizable tool for real-time IHM risk quantification and monitoring in ICU HF patients,using only 17 routinely collected clinical variables.It surpasses traditional scoring systems and enables dynamic risk assessment throughout the ICU stay.By identifying patient-specific major risk factors via SHAP values,the model facilitates timely,personalized treatment adjustments.展开更多
Amid global warming,mountainous regions have emerged as critical zones of investigation owing to their heightened vulnerability to climate change,their ecological significance,and the intensified interactions between ...Amid global warming,mountainous regions have emerged as critical zones of investigation owing to their heightened vulnerability to climate change,their ecological significance,and the intensified interactions between natural stress and human activities.Land surface temperature(LST)is a fundamental indicator for assessing climatic sensitivity in these landscapes.However,a comprehensive understanding of the spatiotemporal dynamics and driving mechanisms of LST across large mountainous regions remains limited.Therefore,data from the Terra Moderate Resolution Imaging Spectroradiometer Land Surface Temperature/Emissivity Daily(MOD11A1)Version 6.1 product during 2001–2020 in Yunnan Province(a complex mountainous region),China,were analyzed.Sen's slope analysis and Mann-Kendall test were applied to detect LST trends and spatial heterogeneity at both annual and seasonal scales.Subsequently,an eXtreme Gradient Boosting(XGBoost)model coupled with SHapley Additive exPlanations(SHAP)was employed to clarify the nonlinear contributions of multiple drivers.The study revealed the following findings.LST exhibited an overall warming rate of 0.020℃/a,characterized by daytime cooling(–0.008℃/a)and nighttime warming(0.048℃/a).LST increased during spring,summer,and autumn(0.011℃/a–0.018℃/a),whereas winter LST exhibited a cooling trend(–0.011℃/a).These variations were spatially partitioned by the Ailao Mountains,with the southwest displaying stronger thermal changes than the northeast.Natural controls,including digital elevation model(DEM)and downward shortwave radiation(DSR),predominated in the northwest high mountain canyons area and south tropical rainforest area,whereas nature–human interactions were more pronounced in the central urban agglomeration warming area and southeast karst landform area.The dominant drivers consisted of DEM,DSR,Normalized Difference Moisture Index,particulate matter 2.5(PM2.5),and aerosol optical depth(AOD).The strong correlations between gross domestic product and population density(correlation coefficient(r)=0.95),as well as between PM2.5 and AOD(r=0.84),highlighted the increasing influence of socioeconomic factors on surface warming.This study can advance the understanding of how mountain topography,moisture,and anthropogenic pressures jointly regulate surface thermal regimes and provide region-specific insights for climate adaptation and sustainable ecosystem management.展开更多
This study provides an in-depth comparative evaluation of landslide susceptibility using two distinct spatial units:and slope units(SUs)and hydrological response units(HRUs),within Goesan County,South Korea.Leveraging...This study provides an in-depth comparative evaluation of landslide susceptibility using two distinct spatial units:and slope units(SUs)and hydrological response units(HRUs),within Goesan County,South Korea.Leveraging the capabilities of the extreme gradient boosting(XGB)algorithm combined with Shapley Additive Explanations(SHAP),this work assesses the precision and clarity with which each unit predicts areas vulnerable to landslides.SUs focus on the geomorphological features like ridges and valleys,focusing on slope stability and landslide triggers.Conversely,HRUs are established based on a variety of hydrological factors,including land cover,soil type and slope gradients,to encapsulate the dynamic water processes of the region.The methodological framework includes the systematic gathering,preparation and analysis of data,ranging from historical landslide occurrences to topographical and environmental variables like elevation,slope angle and land curvature etc.The XGB algorithm used to construct the Landslide Susceptibility Model(LSM)was combined with SHAP for model interpretation and the results were evaluated using Random Cross-validation(RCV)to ensure accuracy and reliability.To ensure optimal model performance,the XGB algorithm’s hyperparameters were tuned using Differential Evolution,considering multicollinearity-free variables.The results show that SU and HRU are effective for LSM,but their effectiveness varies depending on landscape characteristics.The XGB algorithm demonstrates strong predictive power and SHAP enhances model transparency of the influential variables involved.This work underscores the importance of selecting appropriate assessment units tailored to specific landscape characteristics for accurate LSM.The integration of advanced machine learning techniques with interpretative tools offers a robust framework for landslide susceptibility assessment,improving both predictive capabilities and model interpretability.Future research should integrate broader data sets and explore hybrid analytical models to strengthen the generalizability of these findings across varied geographical settings.展开更多
BACKGROUND Severe esophagogastric varices(EGVs)significantly affect prognosis of patients with hepatitis B because of the risk of life-threatening hemorrhage.Endoscopy is the gold standard for EGV detection but it is ...BACKGROUND Severe esophagogastric varices(EGVs)significantly affect prognosis of patients with hepatitis B because of the risk of life-threatening hemorrhage.Endoscopy is the gold standard for EGV detection but it is invasive,costly and carries risks.Noninvasive predictive models using ultrasound and serological markers are essential for identifying high-risk patients and optimizing endoscopy utilization.Machine learning(ML)offers a powerful approach to analyze complex clinical data and improve predictive accuracy.This study hypothesized that ML models,utilizing noninvasive ultrasound and serological markers,can accurately predict the risk of EGVs in hepatitis B patients,thereby improving clinical decisionmaking.AIM To construct and validate a noninvasive predictive model using ML for EGVs in hepatitis B patients.METHODS We retrospectively collected ultrasound and serological data from 310 eligible cases,randomly dividing them into training(80%)and validation(20%)groups.Eleven ML algorithms were used to build predictive models.The performance of the models was evaluated using the area under the curve and decision curve analysis.The best-performing model was further analyzed using SHapley Additive exPlanation to interpret feature importance.RESULTS Among the 310 patients,124 were identified as high-risk for EGVs.The extreme gradient boosting model demonstrated the best performance,achieving an area under the curve of 0.96 in the validation set.The model also exhibited high sensitivity(78%),specificity(94%),positive predictive value(84%),negative predictive value(88%),F1 score(83%),and overall accuracy(86%).The top four predictive variables were albumin,prothrombin time,portal vein flow velocity and spleen stiffness.A web-based version of the model was developed for clinical use,providing real-time predictions for high-risk patients.CONCLUSION We identified an efficient noninvasive predictive model using extreme gradient boosting for EGVs among hepatitis B patients.The model,presented as a web application,has potential for screening high-risk EGV patients and can aid clinicians in optimizing the use of endoscopy.展开更多
BACKGROUND Diabetic foot ulcer(DFU)is a serious and destructive complication of diabetes,which has a high amputation rate and carries a huge social burden.Early detection of risk factors and intervention are essential...BACKGROUND Diabetic foot ulcer(DFU)is a serious and destructive complication of diabetes,which has a high amputation rate and carries a huge social burden.Early detection of risk factors and intervention are essential to reduce amputation rates.With the development of artificial intelligence technology,efficient interpretable predictive models can be generated in clinical practice to improve DFU care.AIM To develop and validate an interpretable model for predicting amputation risk in DFU patients.METHODS This retrospective study collected basic data from 599 patients with DFU in Beijing Shijitan Hospital between January 2015 and June 2024.The data set was randomly divided into a training set and test set with fivefold cross-validation.Three binary variable models were built with the eXtreme Gradient Boosting(XGBoost)algorithm to input risk factors that predict amputation probability.The model performance was optimized by adjusting the super parameters.The pre-dictive performance of the three models was expressed by sensitivity,specificity,positive predictive value,negative predictive value and area under the curve(AUC).Visualization of the prediction results was realized through SHapley Additive exPlanation(SHAP).RESULTS A total of 157(26.2%)patients underwent minor amputation during hospitalization and 50(8.3%)had major amputation.All three XGBoost models demonstrated good discriminative ability,with AUC values>0.7.The model for predicting major amputation achieved the highest performance[AUC=0.977,95%confidence interval(CI):0.956-0.998],followed by the minor amputation model(AUC=0.800,95%CI:0.762-0.838)and the non-amputation model(AUC=0.772,95%CI:0.730-0.814).Feature importance ranking of the three models revealed the risk factors for minor and major amputation.Wagner grade 4/5,osteomyelitis,and high C-reactive protein were all considered important predictive variables.CONCLUSION XGBoost effectively predicts diabetic foot amputation risk and provides interpretable insights to support person-alized treatment decisions.展开更多
Accurate assessment of undrained shear strength(USS)for soft sensitive clays is a great concern in geotechnical engineering practice.This study applies novel data-driven extreme gradient boosting(XGBoost)and random fo...Accurate assessment of undrained shear strength(USS)for soft sensitive clays is a great concern in geotechnical engineering practice.This study applies novel data-driven extreme gradient boosting(XGBoost)and random forest(RF)ensemble learning methods for capturing the relationships between the USS and various basic soil parameters.Based on the soil data sets from TC304 database,a general approach is developed to predict the USS of soft clays using the two machine learning methods above,where five feature variables including the preconsolidation stress(PS),vertical effective stress(VES),liquid limit(LL),plastic limit(PL)and natural water content(W)are adopted.To reduce the dependence on the rule of thumb and inefficient brute-force search,the Bayesian optimization method is applied to determine the appropriate model hyper-parameters of both XGBoost and RF.The developed models are comprehensively compared with three comparison machine learning methods and two transformation models with respect to predictive accuracy and robustness under 5-fold cross-validation(CV).It is shown that XGBoost-based and RF-based methods outperform these approaches.Besides,the XGBoostbased model provides feature importance ranks,which makes it a promising tool in the prediction of geotechnical parameters and enhances the interpretability of model.展开更多
To enhance the accuracy and efficiency of bridge damage identification,a novel data-driven damage identification method was proposed.First,convolutional autoencoder(CAE)was used to extract key features from the accele...To enhance the accuracy and efficiency of bridge damage identification,a novel data-driven damage identification method was proposed.First,convolutional autoencoder(CAE)was used to extract key features from the acceleration signal of the bridge structure through data reconstruction.The extreme gradient boosting tree(XGBoost)was then used to perform analysis on the feature data to achieve damage detection with high accuracy and high performance.The proposed method was applied in a numerical simulation study on a three-span continuous girder and further validated experimentally on a scaled model of a cable-stayed bridge.The numerical simulation results show that the identification errors remain within 2.9%for six single-damage cases and within 3.1%for four double-damage cases.The experimental validation results demonstrate that when the tension in a single cable of the cable-stayed bridge decreases by 20%,the method accurately identifies damage at different cable locations using only sensors installed on the main girder,achieving identification accuracies above 95.8%in all cases.The proposed method shows high identification accuracy and generalization ability across various damage scenarios.展开更多
It is important for regional water resources management to know the agricultural water consumption information several months in advance.Forecasting reference evapotranspiration(ET0)in the next few months is import...It is important for regional water resources management to know the agricultural water consumption information several months in advance.Forecasting reference evapotranspiration(ET0)in the next few months is important for irrigation and reservoir management.Studies on forecasting of multiple-month ahead ET0 using machine learning models have not been reported yet.Besides,machine learning models such as the XGBoost model has multiple parameters that need to be tuned,and traditional methods can get stuck in a regional optimal solution and fail to obtain a global optimal solution.This study investigated the performance of the hybrid extreme gradient boosting(XGBoost)model coupled with the Grey Wolf Optimizer(GWO)algorithm for forecasting multi-step ahead ET0(1-3 months ahead),compared with three conventional machine learning models,i.e.,standalone XGBoost,multi-layer perceptron(MLP)and M5 model tree(M5)models in the subtropical zone of China.The results showed that theGWO-XGB model generally performed better than the other three machine learning models in forecasting 1-3 months ahead ET0,followed by the XGB,M5 and MLP models with very small differences among the three models.The GWO-XGB model performed best in autumn,while the MLP model performed slightly better than the other three models in summer.It is thus suggested to apply the MLP model for ET0 forecasting in summer but use the GWO-XGB model in other seasons.展开更多
Efficient water quality monitoring and ensuring the safety of drinking water by government agencies in areas where the resource is constantly depleted due to anthropogenic or natural factors cannot be overemphasized. ...Efficient water quality monitoring and ensuring the safety of drinking water by government agencies in areas where the resource is constantly depleted due to anthropogenic or natural factors cannot be overemphasized. The above statement holds for West Texas, Midland, and Odessa Precisely. Two machine learning regression algorithms (Random Forest and XGBoost) were employed to develop models for the prediction of total dissolved solids (TDS) and sodium absorption ratio (SAR) for efficient water quality monitoring of two vital aquifers: Edward-Trinity (plateau), and Ogallala aquifers. These two aquifers have contributed immensely to providing water for different uses ranging from domestic, agricultural, industrial, etc. The data was obtained from the Texas Water Development Board (TWDB). The XGBoost and Random Forest models used in this study gave an accurate prediction of observed data (TDS and SAR) for both the Edward-Trinity (plateau) and Ogallala aquifers with the R2 values consistently greater than 0.83. The Random Forest model gave a better prediction of TDS and SAR concentration with an average R, MAE, RMSE and MSE of 0.977, 0.015, 0.029 and 0.00, respectively. For the XGBoost, an average R, MAE, RMSE, and MSE of 0.953, 0.016, 0.037 and 0.00, respectively, were achieved. The overall performance of the models produced was impressive. From this study, we can clearly understand that Random Forest and XGBoost are appropriate for water quality prediction and monitoring in an area of high hydrocarbon activities like Midland and Odessa and West Texas at large.展开更多
The Sentinel-2 satellites are providing an unparalleled wealth of high-resolution remotely sensed information with a short revisit cycle, which is ideal for mapping burned areas both accurately and timely. This paper ...The Sentinel-2 satellites are providing an unparalleled wealth of high-resolution remotely sensed information with a short revisit cycle, which is ideal for mapping burned areas both accurately and timely. This paper proposes an automated methodology for mapping burn scars using pairs of Sentinel-2 imagery, exploiting the state-of-the-art eXtreme Gradient Boosting (XGB) machine learning framework. A large database of 64 reference wildfire perimeters in Greece from 2016 to 2019 is used to train the classifier. An empirical methodology for appropriately sampling the training patterns from this database is formulated, which guarantees the effectiveness of the approach and its computational efficiency. A difference (pre-fire minus post-fire) spectral index is used for this purpose, upon which we appropriately identify the clear and fuzzy value ranges. To reduce the data volume, a super-pixel segmentation of the images is also employed, implemented via the QuickShift algorithm. The cross-validation results showcase the effectiveness of the proposed algorithm, with the average commission and omission errors being 9% and 2%, respectively, and the average Matthews correlation coefficient (MCC) equal to 0.93.展开更多
Nuclear masses are investigated for the first time using the eXtreme Gradient Boosting(XGBoost)method.Nucleon numbers,valence nucleon numbers,and physical quantities related to the magic number are used as input featu...Nuclear masses are investigated for the first time using the eXtreme Gradient Boosting(XGBoost)method.Nucleon numbers,valence nucleon numbers,and physical quantities related to the magic number are used as input features for the decision tree,which learns the residuals of experimental binding energies with respect to the Bethe-Weizsäcker(BW2)formula predictions,and the XGBoost method can achieve high accuracy predictions of nuclear binding energy.For nuclear masses of magic number nuclei with prediction challenges,XGBoost can better capture the physical information associated with the magic number compared to that using BW2,and the root mean square deviation of its predicted nuclear mass ranges from 2.769 to 0.732 MeV.Comparing the results of BW2*and XGBoost* with the pseudo-experimental data of Finite-Range Droplet Model(FRDM12)suggests that the XGBoost* method may have better extrapolation abilities.展开更多
Precisely estimating the remaining mileage of electric vehicles is highly important for vehicle control and battery recharging determinations.Remaining mileage estimation(RME)is a technique difficulty in practice sinc...Precisely estimating the remaining mileage of electric vehicles is highly important for vehicle control and battery recharging determinations.Remaining mileage estimation(RME)is a technique difficulty in practice since it is impacted by many factors,including the battery state of charge(SOC),state of health(SOH),ambient temperature,and traffic condition,etc.In this study,an online RME method is proposed based on dual extended Kalman filter(DEKF)and extreme gradient boosting(XGB)algorithms.Firstly,the battery SOC and SOH are co-estimated based on DEKF with considering the impacts of ambient temperature.Secondly,the current traffic condition are analyzed by using a historical data segement,and then the energy consumpation rate is predicted by XGB algorithm.The XGB algorithm's accuracy under the varying length of data segment is analyzed for determining the proper algorithm parameters.The presented method is evaluated by a simulation study.The results under several typical driving cycles indicate that the precise RME can be achieved with the maximum error less than 1.2%.The method is expected to be useful in providing credible mileage estimation in electric vehiecle applications.展开更多
Concrete is the most commonly used construction material.However,its production leads to high carbon dioxide(CO2)emissions and energy consumption.Therefore,developing waste-substitutable concrete components is nece...Concrete is the most commonly used construction material.However,its production leads to high carbon dioxide(CO2)emissions and energy consumption.Therefore,developing waste-substitutable concrete components is necessary.Improving the sustainability and greenness of concrete is the focus of this research.In this regard,899 data points were collected from existing studies where cement,slag,fly ash,superplasticizer,coarse aggregate,and fine aggregate were considered potential influential factors.The complex relationship between influential factors and concrete compressive strength makes the prediction and estimation of compressive strength difficult.Instead of the traditional compressive strength test,this study combines five novel metaheuristic algorithms with extreme gradient boosting(XGB)to predict the compressive strength of green concrete based on fly ash and blast furnace slag.The intelligent prediction models were assessed using the root mean square error(RMSE),coefficient of determination(R2),mean absolute error(MAE),and variance accounted for(VAF).The results indicated that the squirrel search algorithm-extreme gradient boosting(SSA-XGB)yielded the best overall prediction performance with R2 values of 0.9930 and 0.9576,VAF values of 99.30 and 95.79,MAE values of 0.52 and 2.50,RMSE of 1.34 and 3.31 for the training and testing sets,respectively.The remaining five prediction methods yield promising results.Therefore,the developed hybrid XGB model can be introduced as an accurate and fast technique for the performance prediction of green concrete.Finally,the developed SSA-XGB considered the effects of all the input factors on the compressive strength.The ability of the model to predict the performance of concrete with unknown proportions can play a significant role in accelerating the development and application of sustainable concrete and furthering a sustainable economy.展开更多
Background:Accurate risk stratification of critically ill patients with coronavirus disease 2019(COVID-19)is essential for optimizing resource allocation,delivering targeted interventions,and maximizing patient surviv...Background:Accurate risk stratification of critically ill patients with coronavirus disease 2019(COVID-19)is essential for optimizing resource allocation,delivering targeted interventions,and maximizing patient survival probability.Machine learning(ML)techniques are attracting increased interest for the development of prediction models as they excel in the analysis of complex signals in data-rich environments such as critical care.Methods:We retrieved data on patients with COVID-19 admitted to an intensive care unit(ICU)between March and October 2020 from the RIsk Stratification in COVID-19 patients in the Intensive Care Unit(RISC-19-ICU)registry.We applied the Extreme Gradient Boosting(XGBoost)algorithm to the data to predict as a binary out-come the increase or decrease in patients’Sequential Organ Failure Assessment(SOFA)score on day 5 after ICU admission.The model was iteratively cross-validated in different subsets of the study cohort.Results:The final study population consisted of 675 patients.The XGBoost model correctly predicted a decrease in SOFA score in 320/385(83%)critically ill COVID-19 patients,and an increase in the score in 210/290(72%)patients.The area under the mean receiver operating characteristic curve for XGBoost was significantly higher than that for the logistic regression model(0.86 vs.0.69,P<0.01[paired t-test with 95%confidence interval]).Conclusions:The XGBoost model predicted the change in SOFA score in critically ill COVID-19 patients admitted to the ICU and can guide clinical decision support systems(CDSSs)aimed at optimizing available resources.展开更多
The study aims to develop machine learning-based mechanisms that can accurately predict the axial capacity of high-strength concrete-filled steel tube(CFST)columns.Precisely predicting the axial capacity of a CFST col...The study aims to develop machine learning-based mechanisms that can accurately predict the axial capacity of high-strength concrete-filled steel tube(CFST)columns.Precisely predicting the axial capacity of a CFST column is always challenging for engineers.Using artificial neural networks(ANNs),random forest(RF),and extreme gradient boosting(XG-Boost),a total of 165 experimental data sets were analyzed.The selected input parameters included the steel tensile strength,concrete compressive strength,tube diameter,tube thickness,and column length.The results indicated that the ANN and RF demonstrated a coefficient of determination(R2)value of 0.965 and 0.952 during the training and 0.923 and 0.793 during the testing phase.The most effective technique was the XG-Boost due to its high efficiency,optimizing the gradient boosting,capturing complex patterns,and incorporating regularization to prevent overfitting.The outstanding R2 values of 0.991 and 0.946 during the training and testing were achieved.Due to flexibility in model hyperparameter tuning and customization options,the XG-Boost model demonstrated the lowest values of root mean square error and mean absolute error compared to the other methods.According to the findings,the diameter of CFST columns has the greatest impact on the output,while the column length has the least influence on the ultimate bearing capacity.展开更多
Complex modulus(G*)is one of the important criteria for asphalt classification according to AASHTO M320-10,and is often used to predict the linear viscoelastic behavior of asphalt binders.In addition,phase angle(φ...Complex modulus(G*)is one of the important criteria for asphalt classification according to AASHTO M320-10,and is often used to predict the linear viscoelastic behavior of asphalt binders.In addition,phase angle(φ)characterizes the deformation resilience of asphalt and is used to assess the ratio between the viscous and elastic components.It is thus important to quickly and accurately estimate these two indicators.The purpose of this investigation is to construct an extreme gradient boosting(XGB)model to predict G*andφof graphene oxide(GO)modified asphaltat medium and high temperatures.Two data sets are gathered from previously published experiments,consisting of 357 samples for G*and 339 samples forφ,and the se are used to develop the XGB model using nine inputs representing theasphalt binder components.The findings show that XGB is an excellent predictor of G*andφof GO-modified asphalt,evaluated by the coefficient of determination R2(R2=0.990 and 0.9903 for G*andφ,respectively)and root mean square error(RMSE=31.499 and 1.08 for G*andφ,respectively).In addition,the model’s performance is compared with experimental results and five other machine learning(ML)models to highlight its accuracy.In the final step,the Shapley additive explanations(SHAP)value analysis is conducted to assess the impact of each input and the correlation between pairs of important features on asphalt’s two physical properties.展开更多
Soil liquefaction under strong earthquakes is the primary cause of damage to foundations and superstructures.Predicting the potential for seismic-induced soil liquefaction is key to preventing related disasters.This s...Soil liquefaction under strong earthquakes is the primary cause of damage to foundations and superstructures.Predicting the potential for seismic-induced soil liquefaction is key to preventing related disasters.This study compares four machine learning(ML)models for soil liquefaction potential based on cone penetration test datasets:decision tree,random forest,gradient boosting,and extreme gradient boosting.The database was collected from previously published research and includes information on earthquake moment magnitude,peak ground acceleration,depth of soil layer,total vertical stresses,effective vertical stresses,and cone tip stresses.The predictive capabilities of the developed models were evaluated using overall accuracy,precision,recall,F-measure,and receiver operating characteristic curves.The results showed that the extreme gradient boosting model exhibited the highest efficacy.A subsequent analysis of feature importance demonstrated that cone tip stresses exerted the most significant influence on soil liquefaction potential.In a final comparative assessment with conventional liquefaction discrimination theory methods,the study revealed that the Robertson method yielded a higher success rate for liquefaction cases,the Olsen method was more successful in non-liquefaction cases,and the ML approach manifested superior success rates in both liquefaction and non-liquefaction cases.展开更多
Shield attitude control is a critical aspect that must be continuously monitored during shield tunneling.To achieve scientifically rational settings for shield tunneling parameters,this study constructed multiple mach...Shield attitude control is a critical aspect that must be continuously monitored during shield tunneling.To achieve scientifically rational settings for shield tunneling parameters,this study constructed multiple machine learning prediction models,including shield attitude deviations and tunneling speed,and optimized the hyperparameters of these models using Bayesian algorithms.Subsequently,a constrained grey wolf optimization(GWO)algorithm was employed to establish a real-time safety control method for attitude that considers tunneling efficiency,by dynamically updating the upper and lower bounds for adjustable parameters.The results indicate that the k-nearest neighbors(KNN)model achieved the highest prediction accuracy;however,due to its specific algorithmic principles,KNN is unsuitable for optimization tasks.Embedding the extreme gradient boosting model into the GWO algorithm yielded the best attitude control performance:the absolute attitude deviations were reduced by an average of 45.1%compared to actual values,while the rate of change for adjustable parameters did not exceed 30%.This approach ensures safety and tunneling efficiency during attitude correction and exhibits universal applicability.Compared with other optimization algorithms,GWO demonstrated significant advantages in both optimization effectiveness and computational time.展开更多
Accurately determining the effective fracture toughness(Keff)of rock-concrete(R-C)bi-materials,governed by interface inclination and ambient temperature,is a prerequisite for assessing their structural stability.Th...Accurately determining the effective fracture toughness(Keff)of rock-concrete(R-C)bi-materials,governed by interface inclination and ambient temperature,is a prerequisite for assessing their structural stability.This study developed a hybrid NRBO-XGBoost prediction model using the Newton-RaphsonBased Optimizer(NRBO)to tune the hyperparameters of Extreme Gradient Boosting(XGBoost)model.The established model was developed based on 154 datasets obtained from laboratory tests and numerical simulations with the cracked straight-through Brazilian disc(CSTBD)specimens,including twelve input parameters.The NRBO-XGBoost model for Keffprediction was investigated and compared with seven more models.Furthermore,the Shapley Additive exPlanations(SHAP)method was employed to quantify the contributions of inputs to Keffto improve the interpretability of the developed model.Finally,new data were used to validate the model.Evaluation results demonstrate that metaheuristic optimization algorithms significantly enhance the performance of XGBoost,with NRBO-XGBoost performing the best.The models rank from highest to lowest prediction performance as follows:NRBO-XGBoost,WOA-XGBoost,PSO-XGBoost,XGBoost,RF,CatBoost,LightGBM,and AdaBoost.The interpretable analysis shows that the interface inclination angle exerts the dominant influence.The validation results demonstrate that NRBO-XGBoost achieves high predictive accuracy on a new dataset,showing promising implications for practical applications.展开更多
To curb the worsening tropospheric ozone(O3)pollution problem in China,a rapid and accurate identification of O3-precursor sensitivity(OPS)is a crucial prerequisite for formulating effective contingency O3 po...To curb the worsening tropospheric ozone(O3)pollution problem in China,a rapid and accurate identification of O3-precursor sensitivity(OPS)is a crucial prerequisite for formulating effective contingency O3 pollution control strategies.However,currently widely-used methods,such as statistical models and numerical models,exhibit inherent limitations in identifying OPS in a timely and accurate manner.In this study,we developed a novel approach to identify OPS based on eXtreme Gradient Boosting model,Shapley additive explanation(SHAP)al-gorithm,and volatile organic compound(VOC)photochemical decay adjustment,using the meteorology and speciated pollutant monitoring data as the input.By comparing the difference in SHAP values between base sce-nario and precursor reduction scenario for nitrogen oxides(NOx)and VOCs,OPS was divided into NOx-limited,VOCs-limited and transition regime.Using the long-lasting O3 pollution episode in the autumn of 2022 at the Guangdong-Hong Kong-Macao Greater Bay Area(GBA)as an example,we demonstrated large spatiotemporal heterogeneities of OPS over the GBA,which were generally shifted from NOx-limited to VOCs-limited from September to October and more inclined to be VOCs-limited at the central and NOx-limited in the peripheral areas.This study developed an innovative OPS identification method by comparing the difference in SHAP value before and after precursor emission reduction.Our method enables the accurate identification of OPS in the time scale of seconds,thereby providing a state-of-the-art tool for the rapid guidance of spatial-specific O3 control strategies.展开更多
基金supported by the Student’s Innovation Capability Improvement Plan Project of Guangzhou Medical University(2023 to B.L.)Guangzhou science and technology plan projects[2023A03J0400 to W.C.O and 2024A03J0940 to B.L.]Guangdong Basic and Applied Basic Research Foundation[2021A1515011364 to W.L.].
摘要Objectives To develop and validate a machine learning(ML)model for real-time monitoring of in-hospital mortality(IHM)risk and identification of major risk factors in heart failure(HF)patients admitted to intensive care units(ICUs).Methods Data from ICU HF patients were extracted from the multicenter eICU-Collaborative Research Database(eICU-CRD.External validation used MIMIC-IV and a real-world Chinese dataset(CHN-dataset).Daily measurements from MIMIC-IV patients staying≥3 days formed a Daily Measurement(DM)dataset.After rigorous preprocessing and feature selection,five ML algorithms were trained and optimized using eICU-CRD data.Model performance was evaluated using AUC,sensitivity,specificity,and balanced accuracy.The optimal model was benchmarked against APACHE and SOFA scores.SHapley Additive ex-Planations(SHAP)interpreted feature contributions.A Windows application was developed for clinical deployment.Results XGBoost emerged as the optimal model(Final-ML model),which incorporated only 17 routinely collected clinical variables:age,non-invasive systolic blood pressure(NI-SBP),heart rate,respiratory rate,Glasgow Coma Scale eye opening score,white blood cell count(WBC),creatinine,bicarbonate,red cell distribution width(RDW),platelet count,glucose,calcium,mean corpuscular hemoglobin concentration(MCHC),sodium,mean corpuscular volume,red blood cell count,and potassium.It achieved high AUCs:0.876(95%CI:0.836-0.915;eICU-CRD test data),0.932(95%CI:0.921-0.942;MIMIC-IV),and 0.879(95%CI:0.846-0.912;CHN-dataset).It significantly outperformed APACHE(AUC=0.740,95%CI:0.720-0.761)and SOFA(AUC=0.717,95%CI:0.694-0.740)scores.The model demonstrated strong generalizability across ethnicities,ward types,and genders within MIMIC-IV.Using daily data(DM dataset),predicted IHM risk accurately tracked patient trajectories:risk decreased progressively for survivors and increased for non-survivors throughout the ICU stay.SHAP analysis identified key predictors:NI-SBP,age,heart rate,WBC,glucose,and notably,RDW and MCHC.Time-dependent Cox regression confirmed RDW increase(HR=3.783,95%CI:2.237-6.398)and MCHC decrease(HR=0.173,95%CI:0.040-0.741)as significant independent risk factors for IHM.Conclusions The developed XGBoost model provides a reliable,generalizable tool for real-time IHM risk quantification and monitoring in ICU HF patients,using only 17 routinely collected clinical variables.It surpasses traditional scoring systems and enables dynamic risk assessment throughout the ICU stay.By identifying patient-specific major risk factors via SHAP values,the model facilitates timely,personalized treatment adjustments.
基金supported by the National Natural Science Foundation of China(42061004)the Youth Special Project of Xing Dian Talent Support Program of Yunnan Province,China(XDYC-QNRC-2022-0230)+1 种基金the Open Subjects of First-class Disciplines in Soil and Water Conservation and Desertification Control in Yunnan Province,China(SBK20240021)the Special Project for Building a Science and Technology Innovation Center for South and Southeast Asia,China(202503AP140004)。
摘要Amid global warming,mountainous regions have emerged as critical zones of investigation owing to their heightened vulnerability to climate change,their ecological significance,and the intensified interactions between natural stress and human activities.Land surface temperature(LST)is a fundamental indicator for assessing climatic sensitivity in these landscapes.However,a comprehensive understanding of the spatiotemporal dynamics and driving mechanisms of LST across large mountainous regions remains limited.Therefore,data from the Terra Moderate Resolution Imaging Spectroradiometer Land Surface Temperature/Emissivity Daily(MOD11A1)Version 6.1 product during 2001–2020 in Yunnan Province(a complex mountainous region),China,were analyzed.Sen's slope analysis and Mann-Kendall test were applied to detect LST trends and spatial heterogeneity at both annual and seasonal scales.Subsequently,an eXtreme Gradient Boosting(XGBoost)model coupled with SHapley Additive exPlanations(SHAP)was employed to clarify the nonlinear contributions of multiple drivers.The study revealed the following findings.LST exhibited an overall warming rate of 0.020℃/a,characterized by daytime cooling(–0.008℃/a)and nighttime warming(0.048℃/a).LST increased during spring,summer,and autumn(0.011℃/a–0.018℃/a),whereas winter LST exhibited a cooling trend(–0.011℃/a).These variations were spatially partitioned by the Ailao Mountains,with the southwest displaying stronger thermal changes than the northeast.Natural controls,including digital elevation model(DEM)and downward shortwave radiation(DSR),predominated in the northwest high mountain canyons area and south tropical rainforest area,whereas nature–human interactions were more pronounced in the central urban agglomeration warming area and southeast karst landform area.The dominant drivers consisted of DEM,DSR,Normalized Difference Moisture Index,particulate matter 2.5(PM2.5),and aerosol optical depth(AOD).The strong correlations between gross domestic product and population density(correlation coefficient(r)=0.95),as well as between PM2.5 and AOD(r=0.84),highlighted the increasing influence of socioeconomic factors on surface warming.This study can advance the understanding of how mountain topography,moisture,and anthropogenic pressures jointly regulate surface thermal regimes and provide region-specific insights for climate adaptation and sustainable ecosystem management.
基金supported by a National Research Foundation of Korea(NRF)grant funded by the Korean government(MSIT)(RS-2023-00222536).
摘要This study provides an in-depth comparative evaluation of landslide susceptibility using two distinct spatial units:and slope units(SUs)and hydrological response units(HRUs),within Goesan County,South Korea.Leveraging the capabilities of the extreme gradient boosting(XGB)algorithm combined with Shapley Additive Explanations(SHAP),this work assesses the precision and clarity with which each unit predicts areas vulnerable to landslides.SUs focus on the geomorphological features like ridges and valleys,focusing on slope stability and landslide triggers.Conversely,HRUs are established based on a variety of hydrological factors,including land cover,soil type and slope gradients,to encapsulate the dynamic water processes of the region.The methodological framework includes the systematic gathering,preparation and analysis of data,ranging from historical landslide occurrences to topographical and environmental variables like elevation,slope angle and land curvature etc.The XGB algorithm used to construct the Landslide Susceptibility Model(LSM)was combined with SHAP for model interpretation and the results were evaluated using Random Cross-validation(RCV)to ensure accuracy and reliability.To ensure optimal model performance,the XGB algorithm’s hyperparameters were tuned using Differential Evolution,considering multicollinearity-free variables.The results show that SU and HRU are effective for LSM,but their effectiveness varies depending on landscape characteristics.The XGB algorithm demonstrates strong predictive power and SHAP enhances model transparency of the influential variables involved.This work underscores the importance of selecting appropriate assessment units tailored to specific landscape characteristics for accurate LSM.The integration of advanced machine learning techniques with interpretative tools offers a robust framework for landslide susceptibility assessment,improving both predictive capabilities and model interpretability.Future research should integrate broader data sets and explore hybrid analytical models to strengthen the generalizability of these findings across varied geographical settings.
基金Supported by the Agency Natural Science Foundation of Fujian Province,China,No.2022J011285 and No.2023J011480.
摘要BACKGROUND Severe esophagogastric varices(EGVs)significantly affect prognosis of patients with hepatitis B because of the risk of life-threatening hemorrhage.Endoscopy is the gold standard for EGV detection but it is invasive,costly and carries risks.Noninvasive predictive models using ultrasound and serological markers are essential for identifying high-risk patients and optimizing endoscopy utilization.Machine learning(ML)offers a powerful approach to analyze complex clinical data and improve predictive accuracy.This study hypothesized that ML models,utilizing noninvasive ultrasound and serological markers,can accurately predict the risk of EGVs in hepatitis B patients,thereby improving clinical decisionmaking.AIM To construct and validate a noninvasive predictive model using ML for EGVs in hepatitis B patients.METHODS We retrospectively collected ultrasound and serological data from 310 eligible cases,randomly dividing them into training(80%)and validation(20%)groups.Eleven ML algorithms were used to build predictive models.The performance of the models was evaluated using the area under the curve and decision curve analysis.The best-performing model was further analyzed using SHapley Additive exPlanation to interpret feature importance.RESULTS Among the 310 patients,124 were identified as high-risk for EGVs.The extreme gradient boosting model demonstrated the best performance,achieving an area under the curve of 0.96 in the validation set.The model also exhibited high sensitivity(78%),specificity(94%),positive predictive value(84%),negative predictive value(88%),F1 score(83%),and overall accuracy(86%).The top four predictive variables were albumin,prothrombin time,portal vein flow velocity and spleen stiffness.A web-based version of the model was developed for clinical use,providing real-time predictions for high-risk patients.CONCLUSION We identified an efficient noninvasive predictive model using extreme gradient boosting for EGVs among hepatitis B patients.The model,presented as a web application,has potential for screening high-risk EGV patients and can aid clinicians in optimizing the use of endoscopy.
摘要BACKGROUND Diabetic foot ulcer(DFU)is a serious and destructive complication of diabetes,which has a high amputation rate and carries a huge social burden.Early detection of risk factors and intervention are essential to reduce amputation rates.With the development of artificial intelligence technology,efficient interpretable predictive models can be generated in clinical practice to improve DFU care.AIM To develop and validate an interpretable model for predicting amputation risk in DFU patients.METHODS This retrospective study collected basic data from 599 patients with DFU in Beijing Shijitan Hospital between January 2015 and June 2024.The data set was randomly divided into a training set and test set with fivefold cross-validation.Three binary variable models were built with the eXtreme Gradient Boosting(XGBoost)algorithm to input risk factors that predict amputation probability.The model performance was optimized by adjusting the super parameters.The pre-dictive performance of the three models was expressed by sensitivity,specificity,positive predictive value,negative predictive value and area under the curve(AUC).Visualization of the prediction results was realized through SHapley Additive exPlanation(SHAP).RESULTS A total of 157(26.2%)patients underwent minor amputation during hospitalization and 50(8.3%)had major amputation.All three XGBoost models demonstrated good discriminative ability,with AUC values>0.7.The model for predicting major amputation achieved the highest performance[AUC=0.977,95%confidence interval(CI):0.956-0.998],followed by the minor amputation model(AUC=0.800,95%CI:0.762-0.838)and the non-amputation model(AUC=0.772,95%CI:0.730-0.814).Feature importance ranking of the three models revealed the risk factors for minor and major amputation.Wagner grade 4/5,osteomyelitis,and high C-reactive protein were all considered important predictive variables.CONCLUSION XGBoost effectively predicts diabetic foot amputation risk and provides interpretable insights to support person-alized treatment decisions.
基金financial support from High-end Foreign Expert Introduction program(No.G20190022002)Chongqing Construction Science and Technology Plan Project(2019-0045)as well as Chongqing Engineering Research Center of Disaster Prevention&Control for Banks and Structures in Three Gorges Reservoir Area(Nos.SXAPGC18ZD01 and SXAPGC18YB03)。
摘要Accurate assessment of undrained shear strength(USS)for soft sensitive clays is a great concern in geotechnical engineering practice.This study applies novel data-driven extreme gradient boosting(XGBoost)and random forest(RF)ensemble learning methods for capturing the relationships between the USS and various basic soil parameters.Based on the soil data sets from TC304 database,a general approach is developed to predict the USS of soft clays using the two machine learning methods above,where five feature variables including the preconsolidation stress(PS),vertical effective stress(VES),liquid limit(LL),plastic limit(PL)and natural water content(W)are adopted.To reduce the dependence on the rule of thumb and inefficient brute-force search,the Bayesian optimization method is applied to determine the appropriate model hyper-parameters of both XGBoost and RF.The developed models are comprehensively compared with three comparison machine learning methods and two transformation models with respect to predictive accuracy and robustness under 5-fold cross-validation(CV).It is shown that XGBoost-based and RF-based methods outperform these approaches.Besides,the XGBoostbased model provides feature importance ranks,which makes it a promising tool in the prediction of geotechnical parameters and enhances the interpretability of model.
基金The National Natural Science Foundation of China(No.52361165658,52378318,52078459).
摘要To enhance the accuracy and efficiency of bridge damage identification,a novel data-driven damage identification method was proposed.First,convolutional autoencoder(CAE)was used to extract key features from the acceleration signal of the bridge structure through data reconstruction.The extreme gradient boosting tree(XGBoost)was then used to perform analysis on the feature data to achieve damage detection with high accuracy and high performance.The proposed method was applied in a numerical simulation study on a three-span continuous girder and further validated experimentally on a scaled model of a cable-stayed bridge.The numerical simulation results show that the identification errors remain within 2.9%for six single-damage cases and within 3.1%for four double-damage cases.The experimental validation results demonstrate that when the tension in a single cable of the cable-stayed bridge decreases by 20%,the method accurately identifies damage at different cable locations using only sensors installed on the main girder,achieving identification accuracies above 95.8%in all cases.The proposed method shows high identification accuracy and generalization ability across various damage scenarios.
基金This study was jointly supported by the National Natural Science Foundation of China(Nos.51879196,51790533,51709143)Jiangxi Natural Science Foundation of China(No.20181BAB206045).
摘要It is important for regional water resources management to know the agricultural water consumption information several months in advance.Forecasting reference evapotranspiration(ET0)in the next few months is important for irrigation and reservoir management.Studies on forecasting of multiple-month ahead ET0 using machine learning models have not been reported yet.Besides,machine learning models such as the XGBoost model has multiple parameters that need to be tuned,and traditional methods can get stuck in a regional optimal solution and fail to obtain a global optimal solution.This study investigated the performance of the hybrid extreme gradient boosting(XGBoost)model coupled with the Grey Wolf Optimizer(GWO)algorithm for forecasting multi-step ahead ET0(1-3 months ahead),compared with three conventional machine learning models,i.e.,standalone XGBoost,multi-layer perceptron(MLP)and M5 model tree(M5)models in the subtropical zone of China.The results showed that theGWO-XGB model generally performed better than the other three machine learning models in forecasting 1-3 months ahead ET0,followed by the XGB,M5 and MLP models with very small differences among the three models.The GWO-XGB model performed best in autumn,while the MLP model performed slightly better than the other three models in summer.It is thus suggested to apply the MLP model for ET0 forecasting in summer but use the GWO-XGB model in other seasons.
摘要Efficient water quality monitoring and ensuring the safety of drinking water by government agencies in areas where the resource is constantly depleted due to anthropogenic or natural factors cannot be overemphasized. The above statement holds for West Texas, Midland, and Odessa Precisely. Two machine learning regression algorithms (Random Forest and XGBoost) were employed to develop models for the prediction of total dissolved solids (TDS) and sodium absorption ratio (SAR) for efficient water quality monitoring of two vital aquifers: Edward-Trinity (plateau), and Ogallala aquifers. These two aquifers have contributed immensely to providing water for different uses ranging from domestic, agricultural, industrial, etc. The data was obtained from the Texas Water Development Board (TWDB). The XGBoost and Random Forest models used in this study gave an accurate prediction of observed data (TDS and SAR) for both the Edward-Trinity (plateau) and Ogallala aquifers with the R2 values consistently greater than 0.83. The Random Forest model gave a better prediction of TDS and SAR concentration with an average R, MAE, RMSE and MSE of 0.977, 0.015, 0.029 and 0.00, respectively. For the XGBoost, an average R, MAE, RMSE, and MSE of 0.953, 0.016, 0.037 and 0.00, respectively, were achieved. The overall performance of the models produced was impressive. From this study, we can clearly understand that Random Forest and XGBoost are appropriate for water quality prediction and monitoring in an area of high hydrocarbon activities like Midland and Odessa and West Texas at large.
摘要The Sentinel-2 satellites are providing an unparalleled wealth of high-resolution remotely sensed information with a short revisit cycle, which is ideal for mapping burned areas both accurately and timely. This paper proposes an automated methodology for mapping burn scars using pairs of Sentinel-2 imagery, exploiting the state-of-the-art eXtreme Gradient Boosting (XGB) machine learning framework. A large database of 64 reference wildfire perimeters in Greece from 2016 to 2019 is used to train the classifier. An empirical methodology for appropriately sampling the training patterns from this database is formulated, which guarantees the effectiveness of the approach and its computational efficiency. A difference (pre-fire minus post-fire) spectral index is used for this purpose, upon which we appropriately identify the clear and fuzzy value ranges. To reduce the data volume, a super-pixel segmentation of the images is also employed, implemented via the QuickShift algorithm. The cross-validation results showcase the effectiveness of the proposed algorithm, with the average commission and omission errors being 9% and 2%, respectively, and the average Matthews correlation coefficient (MCC) equal to 0.93.
基金supported by the Scientific Research Foundation for High-level Talents of Anhui University of Science and Technology(2023yjrc107)。
摘要Nuclear masses are investigated for the first time using the eXtreme Gradient Boosting(XGBoost)method.Nucleon numbers,valence nucleon numbers,and physical quantities related to the magic number are used as input features for the decision tree,which learns the residuals of experimental binding energies with respect to the Bethe-Weizsäcker(BW2)formula predictions,and the XGBoost method can achieve high accuracy predictions of nuclear binding energy.For nuclear masses of magic number nuclei with prediction challenges,XGBoost can better capture the physical information associated with the magic number compared to that using BW2,and the root mean square deviation of its predicted nuclear mass ranges from 2.769 to 0.732 MeV.Comparing the results of BW2*and XGBoost* with the pseudo-experimental data of Finite-Range Droplet Model(FRDM12)suggests that the XGBoost* method may have better extrapolation abilities.
基金supported by the National Natural Science Foundation of China(U23B20139,52172401)the Fundamental Research Funds for the Central Universities(Grant No.N2403013).
摘要Precisely estimating the remaining mileage of electric vehicles is highly important for vehicle control and battery recharging determinations.Remaining mileage estimation(RME)is a technique difficulty in practice since it is impacted by many factors,including the battery state of charge(SOC),state of health(SOH),ambient temperature,and traffic condition,etc.In this study,an online RME method is proposed based on dual extended Kalman filter(DEKF)and extreme gradient boosting(XGB)algorithms.Firstly,the battery SOC and SOH are co-estimated based on DEKF with considering the impacts of ambient temperature.Secondly,the current traffic condition are analyzed by using a historical data segement,and then the energy consumpation rate is predicted by XGB algorithm.The XGB algorithm's accuracy under the varying length of data segment is analyzed for determining the proper algorithm parameters.The presented method is evaluated by a simulation study.The results under several typical driving cycles indicate that the precise RME can be achieved with the maximum error less than 1.2%.The method is expected to be useful in providing credible mileage estimation in electric vehiecle applications.
基金funding provided by the China Scholarship Council (Nos.202008440524 and 202006370006)supported by the Distinguished Youth Science Foundation of Hunan Province of China (No.2022JJ10073)+1 种基金Innovation Driven Project of Central South University (No.2020CX040)Shenzhen Sciencee and Technology Plan (No.JCYJ20190808123013260).
摘要Concrete is the most commonly used construction material.However,its production leads to high carbon dioxide(CO2)emissions and energy consumption.Therefore,developing waste-substitutable concrete components is necessary.Improving the sustainability and greenness of concrete is the focus of this research.In this regard,899 data points were collected from existing studies where cement,slag,fly ash,superplasticizer,coarse aggregate,and fine aggregate were considered potential influential factors.The complex relationship between influential factors and concrete compressive strength makes the prediction and estimation of compressive strength difficult.Instead of the traditional compressive strength test,this study combines five novel metaheuristic algorithms with extreme gradient boosting(XGB)to predict the compressive strength of green concrete based on fly ash and blast furnace slag.The intelligent prediction models were assessed using the root mean square error(RMSE),coefficient of determination(R2),mean absolute error(MAE),and variance accounted for(VAF).The results indicated that the squirrel search algorithm-extreme gradient boosting(SSA-XGB)yielded the best overall prediction performance with R2 values of 0.9930 and 0.9576,VAF values of 99.30 and 95.79,MAE values of 0.52 and 2.50,RMSE of 1.34 and 3.31 for the training and testing sets,respectively.The remaining five prediction methods yield promising results.Therefore,the developed hybrid XGB model can be introduced as an accurate and fast technique for the performance prediction of green concrete.Finally,the developed SSA-XGB considered the effects of all the input factors on the compressive strength.The ability of the model to predict the performance of concrete with unknown proportions can play a significant role in accelerating the development and application of sustainable concrete and furthering a sustainable economy.
基金supported by the“Microsoft Grant Award:AI for Health COVID-19″The RISC-19-ICU reg-istry is supported by the Swiss Society of Intensive Care Medicine and funded by internal resources of the Institute of Intensive Care Medicine,of the University Hospital Zurich and by unrestricted grants from CytoSorbents Europe GmbH(Berlin,Germany)+1 种基金Union Bancaire Privée(Zurich,Switzerland)The sponsors had no role in the design of the study,the collection and analysis of the data,or the preparation of the manuscript.
摘要Background:Accurate risk stratification of critically ill patients with coronavirus disease 2019(COVID-19)is essential for optimizing resource allocation,delivering targeted interventions,and maximizing patient survival probability.Machine learning(ML)techniques are attracting increased interest for the development of prediction models as they excel in the analysis of complex signals in data-rich environments such as critical care.Methods:We retrieved data on patients with COVID-19 admitted to an intensive care unit(ICU)between March and October 2020 from the RIsk Stratification in COVID-19 patients in the Intensive Care Unit(RISC-19-ICU)registry.We applied the Extreme Gradient Boosting(XGBoost)algorithm to the data to predict as a binary out-come the increase or decrease in patients’Sequential Organ Failure Assessment(SOFA)score on day 5 after ICU admission.The model was iteratively cross-validated in different subsets of the study cohort.Results:The final study population consisted of 675 patients.The XGBoost model correctly predicted a decrease in SOFA score in 320/385(83%)critically ill COVID-19 patients,and an increase in the score in 210/290(72%)patients.The area under the mean receiver operating characteristic curve for XGBoost was significantly higher than that for the logistic regression model(0.86 vs.0.69,P<0.01[paired t-test with 95%confidence interval]).Conclusions:The XGBoost model predicted the change in SOFA score in critically ill COVID-19 patients admitted to the ICU and can guide clinical decision support systems(CDSSs)aimed at optimizing available resources.
基金The Second Century Fund(C2F),Chulalongkorn University.
摘要The study aims to develop machine learning-based mechanisms that can accurately predict the axial capacity of high-strength concrete-filled steel tube(CFST)columns.Precisely predicting the axial capacity of a CFST column is always challenging for engineers.Using artificial neural networks(ANNs),random forest(RF),and extreme gradient boosting(XG-Boost),a total of 165 experimental data sets were analyzed.The selected input parameters included the steel tensile strength,concrete compressive strength,tube diameter,tube thickness,and column length.The results indicated that the ANN and RF demonstrated a coefficient of determination(R2)value of 0.965 and 0.952 during the training and 0.923 and 0.793 during the testing phase.The most effective technique was the XG-Boost due to its high efficiency,optimizing the gradient boosting,capturing complex patterns,and incorporating regularization to prevent overfitting.The outstanding R2 values of 0.991 and 0.946 during the training and testing were achieved.Due to flexibility in model hyperparameter tuning and customization options,the XG-Boost model demonstrated the lowest values of root mean square error and mean absolute error compared to the other methods.According to the findings,the diameter of CFST columns has the greatest impact on the output,while the column length has the least influence on the ultimate bearing capacity.
摘要Complex modulus(G*)is one of the important criteria for asphalt classification according to AASHTO M320-10,and is often used to predict the linear viscoelastic behavior of asphalt binders.In addition,phase angle(φ)characterizes the deformation resilience of asphalt and is used to assess the ratio between the viscous and elastic components.It is thus important to quickly and accurately estimate these two indicators.The purpose of this investigation is to construct an extreme gradient boosting(XGB)model to predict G*andφof graphene oxide(GO)modified asphaltat medium and high temperatures.Two data sets are gathered from previously published experiments,consisting of 357 samples for G*and 339 samples forφ,and the se are used to develop the XGB model using nine inputs representing theasphalt binder components.The findings show that XGB is an excellent predictor of G*andφof GO-modified asphalt,evaluated by the coefficient of determination R2(R2=0.990 and 0.9903 for G*andφ,respectively)and root mean square error(RMSE=31.499 and 1.08 for G*andφ,respectively).In addition,the model’s performance is compared with experimental results and five other machine learning(ML)models to highlight its accuracy.In the final step,the Shapley additive explanations(SHAP)value analysis is conducted to assess the impact of each input and the correlation between pairs of important features on asphalt’s two physical properties.
基金National Natural Science Foundation of China under Grant Nos.42271088 and 42102311Fundamental Research Funds for the Central Universities under Grant No.LH2022D016+1 种基金Key Laboratory of Coal Gangue Resource Utilization and Energy-Saving Building Materials of Liaoning under Grant No.LNTUCEM-2304Heilongjiang Provincial Science and Technology Innovation Base Award Project under Grant No.JD25B010。
摘要Soil liquefaction under strong earthquakes is the primary cause of damage to foundations and superstructures.Predicting the potential for seismic-induced soil liquefaction is key to preventing related disasters.This study compares four machine learning(ML)models for soil liquefaction potential based on cone penetration test datasets:decision tree,random forest,gradient boosting,and extreme gradient boosting.The database was collected from previously published research and includes information on earthquake moment magnitude,peak ground acceleration,depth of soil layer,total vertical stresses,effective vertical stresses,and cone tip stresses.The predictive capabilities of the developed models were evaluated using overall accuracy,precision,recall,F-measure,and receiver operating characteristic curves.The results showed that the extreme gradient boosting model exhibited the highest efficacy.A subsequent analysis of feature importance demonstrated that cone tip stresses exerted the most significant influence on soil liquefaction potential.In a final comparative assessment with conventional liquefaction discrimination theory methods,the study revealed that the Robertson method yielded a higher success rate for liquefaction cases,the Olsen method was more successful in non-liquefaction cases,and the ML approach manifested superior success rates in both liquefaction and non-liquefaction cases.
基金supported by the National Natural Science Foundation of China(Grant Nos.52378386 and 52178336).
摘要Shield attitude control is a critical aspect that must be continuously monitored during shield tunneling.To achieve scientifically rational settings for shield tunneling parameters,this study constructed multiple machine learning prediction models,including shield attitude deviations and tunneling speed,and optimized the hyperparameters of these models using Bayesian algorithms.Subsequently,a constrained grey wolf optimization(GWO)algorithm was employed to establish a real-time safety control method for attitude that considers tunneling efficiency,by dynamically updating the upper and lower bounds for adjustable parameters.The results indicate that the k-nearest neighbors(KNN)model achieved the highest prediction accuracy;however,due to its specific algorithmic principles,KNN is unsuitable for optimization tasks.Embedding the extreme gradient boosting model into the GWO algorithm yielded the best attitude control performance:the absolute attitude deviations were reduced by an average of 45.1%compared to actual values,while the rate of change for adjustable parameters did not exceed 30%.This approach ensures safety and tunneling efficiency during attitude correction and exhibits universal applicability.Compared with other optimization algorithms,GWO demonstrated significant advantages in both optimization effectiveness and computational time.
基金financially supported by the National Natural Science Foundation of China(Nos.52274167 and 52304123)the Hunan Province’s technology research project“Revealing the List and Taking Command”(No.2021SK1050)+2 种基金the Young Talent Lifting Project of the China Association for Science and Technology(No.2024QNRC001)Natural Science Foundation of University of South China(No.5525QD012)Sichuan-Chongqing Science and Technology Innovation Cooperation Program Project(No.CSTB2024TIAD-CYKJCXX0016)。
摘要Accurately determining the effective fracture toughness(Keff)of rock-concrete(R-C)bi-materials,governed by interface inclination and ambient temperature,is a prerequisite for assessing their structural stability.This study developed a hybrid NRBO-XGBoost prediction model using the Newton-RaphsonBased Optimizer(NRBO)to tune the hyperparameters of Extreme Gradient Boosting(XGBoost)model.The established model was developed based on 154 datasets obtained from laboratory tests and numerical simulations with the cracked straight-through Brazilian disc(CSTBD)specimens,including twelve input parameters.The NRBO-XGBoost model for Keffprediction was investigated and compared with seven more models.Furthermore,the Shapley Additive exPlanations(SHAP)method was employed to quantify the contributions of inputs to Keffto improve the interpretability of the developed model.Finally,new data were used to validate the model.Evaluation results demonstrate that metaheuristic optimization algorithms significantly enhance the performance of XGBoost,with NRBO-XGBoost performing the best.The models rank from highest to lowest prediction performance as follows:NRBO-XGBoost,WOA-XGBoost,PSO-XGBoost,XGBoost,RF,CatBoost,LightGBM,and AdaBoost.The interpretable analysis shows that the interface inclination angle exerts the dominant influence.The validation results demonstrate that NRBO-XGBoost achieves high predictive accuracy on a new dataset,showing promising implications for practical applications.
基金supported by the Key-Area Research and Development Program of Guangdong Province(No.2020B1111360003)the National Natural Science Foundation of China(Nos.42465008 and 42105164)+2 种基金Yunnan Science and Technology Department Project(No.202501AT070239)Yunnan Science and Technology Department Youth Project(No.202401AU070202)Xianyang Rapid Response Decision Support Project for Ozone(No.YZ2024-ZB019).
摘要To curb the worsening tropospheric ozone(O3)pollution problem in China,a rapid and accurate identification of O3-precursor sensitivity(OPS)is a crucial prerequisite for formulating effective contingency O3 pollution control strategies.However,currently widely-used methods,such as statistical models and numerical models,exhibit inherent limitations in identifying OPS in a timely and accurate manner.In this study,we developed a novel approach to identify OPS based on eXtreme Gradient Boosting model,Shapley additive explanation(SHAP)al-gorithm,and volatile organic compound(VOC)photochemical decay adjustment,using the meteorology and speciated pollutant monitoring data as the input.By comparing the difference in SHAP values between base sce-nario and precursor reduction scenario for nitrogen oxides(NOx)and VOCs,OPS was divided into NOx-limited,VOCs-limited and transition regime.Using the long-lasting O3 pollution episode in the autumn of 2022 at the Guangdong-Hong Kong-Macao Greater Bay Area(GBA)as an example,we demonstrated large spatiotemporal heterogeneities of OPS over the GBA,which were generally shifted from NOx-limited to VOCs-limited from September to October and more inclined to be VOCs-limited at the central and NOx-limited in the peripheral areas.This study developed an innovative OPS identification method by comparing the difference in SHAP value before and after precursor emission reduction.Our method enables the accurate identification of OPS in the time scale of seconds,thereby providing a state-of-the-art tool for the rapid guidance of spatial-specific O3 control strategies.