Slope units are divided according to the real topography and have clear geological characteristics,making them ideal units for evaluating the susceptibility to geological disasters.Based on the results of automaticall...Slope units are divided according to the real topography and have clear geological characteristics,making them ideal units for evaluating the susceptibility to geological disasters.Based on the results of automatically and manually corrected hydrological slope unit division,the Longhua District,Shenzhen City,Guangdong Province,was selected as the study area.A total of 15 influencing factors,namely Fluctuation,slope,slope aspect,curvature,topographic witness index(TWI),stream power index(SPI),topographic roughness index(TRI),annual average rainfall,distance to water system,engineering rock group,distance to fault,land use,normalized difference vegetation index(NDVI),nighttime light,and distance to road,were selected as evaluation indicators.The information volume model(IV)and random points were used to select non-geological disaster units,and then the random forest model(RF)was used to evaluate the susceptibility to geological disasters.The automatic slope unit and the hydrological slope unit were compared and analyzed in the random forest and information volume random forest models.The results show that the area under the curve(AUC)values of the automatic slope unit evaluation results are 0.931 for the IV-RF model and 0.716 for the RF model,which are 0.6%(IV-RF model)and 1.9%(RF model)higher than those for the hydrological slope unit.Based on a comparison of the evaluation methods based on the two types of slope units,the hydrological slope unit evaluation method based on manual correction is highly subjective,is complicated to operate,and has a low evaluation accuracy,whereas the evaluation method based on automatic slope unit division is efficient and accurate,is suitable for large-scale efficient geological disaster evaluation,and can better deal with the problem of geological disaster susceptibility evaluation.展开更多
Detecting cyber attacks in networks connected to the Internet of Things(IoT)is of utmost importance because of the growing vulnerabilities in the smart environment.Conventional models,such as Naive Bayes and support v...Detecting cyber attacks in networks connected to the Internet of Things(IoT)is of utmost importance because of the growing vulnerabilities in the smart environment.Conventional models,such as Naive Bayes and support vector machine(SVM),as well as ensemble methods,such as Gradient Boosting and eXtreme gradient boosting(XGBoost),are often plagued by high computational costs,which makes it challenging for them to perform real-time detection.In this regard,we suggested an attack detection approach that integrates Visual Geometry Group 16(VGG16),Artificial Rabbits Optimizer(ARO),and Random Forest Model to increase detection accuracy and operational efficiency in Internet of Things(IoT)networks.In the suggested model,the extraction of features from malware pictures was accomplished with the help of VGG16.The prediction process is carried out by the random forest model using the extracted features from the VGG16.Additionally,ARO is used to improve the hyper-parameters of the random forest model of the random forest.With an accuracy of 96.36%,the suggested model outperforms the standard models in terms of accuracy,F1-score,precision,and recall.The comparative research highlights our strategy’s success,which improves performance while maintaining a lower computational cost.This method is ideal for real-time applications,but it is effective.展开更多
Random forest model is the mainstream research method used to accurately describe the distribution law and impact mechanism of regional population.We took Shijiazhuang as the research area,with comprehensive zoning ba...Random forest model is the mainstream research method used to accurately describe the distribution law and impact mechanism of regional population.We took Shijiazhuang as the research area,with comprehensive zoning based on endowments as the modeling unit,conducted stratified sampling on a hectare grid cell,and systematically carried out incremental selection experiments of population density impact factors,optimizing the population density random forest model throughout the process(zonal modeling,stratified sampling,factor selection,weighted output).The results are as follows:(1)Zonal modeling addresses the issue of confusion in population distribution laws caused by a single model.Sampling on a grid cell not only ensures the quality of training data by avoiding the modifiable areal unit problem(MAUP)but also attempts to mitigate the adverse effects of the ecological fallacy.Stratified sampling ensures the stability of population density label values(target variable)in the training sample.(2)Zonal selection experiments on population density impact factors help identify suitable combinations of factors,leading to a significant improvement in the goodness of fit(R2)of the zonal models.(3)Weighted combination output of the population density prediction dataset substantially enhances the model's robustness.(4)The population density dataset exhibits multi-scale superposition characteristics.On a large scale,the population density in plains is higher than that in mountainous areas,while on a small scale,urban areas have higher density compared to rural areas.The optimization scheme for the population density random forest model that we propose offers a unified technical framework for uncovering local population distribution law and the impact mechanisms.展开更多
Potential of the Random Forest Model on mapping of different desertification processes was studied in Muttuma watershed of mid-Murrumbidgee river region of New South Wales,Australia.Desertification vulnerability index...Potential of the Random Forest Model on mapping of different desertification processes was studied in Muttuma watershed of mid-Murrumbidgee river region of New South Wales,Australia.Desertification vulnerability index was developed using climate,terrain,vegetation,soil and land quality indices to identify environmentally sensitive areas for desertification.Random Forest Model(RFM)was used to predict the different desertification processes such as soil erosion,salinization and waterlogging in the watershed and the information needed to train classification algorithms was obtained from satellite imagery interpretation and ground truth data.Climatic factors(evaporation,rainfall,temperature),terrain factors(aspect,slope,slope length,steepness,and wetness index),soil properties(pH,organic carbon,clay and sand content)and vulnerability indices were used as an explanatory variable.Classification accuracy and kappa index were calculated for training and testing datasets.We recorded an overall accuracy rate of 87.7%and 72.1%for training and testing sites,respectively.We found larger discrepancies between overall accuracy rate and kappa index for testing datasets(72.2%and 27.5%,respectively)suggesting that all the classes are not predicted well.The prediction of soil erosion and no desertification process was good and poor for salinization and water-logging process.Overall,the results observed give a new idea of using the knowledge of desertification process in training areas that can be used to predict the desertification processes at unvisited areas.展开更多
Objective Body fluid mixtures are complex biological samples that frequently occur in crime scenes,and can provide important clues for criminal case analysis.DNA methylation assay has been applied in the identificatio...Objective Body fluid mixtures are complex biological samples that frequently occur in crime scenes,and can provide important clues for criminal case analysis.DNA methylation assay has been applied in the identification of human body fluids,and has exhibited excellent performance in predicting single-source body fluids.The present study aims to develop a methylation SNaPshot multiplex system for body fluid identification,and accurately predict the mixture samples.In addition,the value of DNA methylation in the prediction of body fluid mixtures was further explored.Methods In the present study,420 samples of body fluid mixtures and 250 samples of single body fluids were tested using an optimized multiplex methylation system.Each kind of body fluid sample presented the specific methylation profiles of the 10 markers.Results Significant differences in methylation levels were observed between the mixtures and single body fluids.For all kinds of mixtures,the Spearman’s correlation analysis revealed a significantly strong correlation between the methylation levels and component proportions(1:20,1:10,1:5,1:1,5:1,10:1 and 20:1).Two random forest classification models were trained for the prediction of mixture types and the prediction of the mixture proportion of 2 components,based on the methylation levels of 10 markers.For the mixture prediction,Model-1 presented outstanding prediction accuracy,which reached up to 99.3%in 427 training samples,and had a remarkable accuracy of 100%in 243 independent test samples.For the mixture proportion prediction,Model-2 demonstrated an excellent accuracy of 98.8%in 252 training samples,and 98.2%in 168 independent test samples.The total prediction accuracy reached 99.3%for body fluid mixtures and 98.6%for the mixture proportions.Conclusion These results indicate the excellent capability and powerful value of the multiplex methylation system in the identification of forensic body fluid mixtures.展开更多
This study proposes a framework to automate the evaluation of traditional village preservation status and analyze the major influential factors(MIFs)of preservation status and influencing mechanisms through YOLOv10 mo...This study proposes a framework to automate the evaluation of traditional village preservation status and analyze the major influential factors(MIFs)of preservation status and influencing mechanisms through YOLOv10 model and Random Forest model,taking Tibetan-Qiang region of northwest Sichuan as the study area.The framework adopts satellite maps,based on the YOLOv10 model,to comprehensively detect the preservation status of houses in the traditional villages,and calculates the preservation score of the corresponding villages as an evaluation of their preservation status.Further,through the Feature importance of Random Forest model,the MIFs of the village preservation status are filtered from the multiple environmental factors,and the SHAP value resolves the influencing intensity of the MIFs on the preservation status.Finally,for villages with poor preservation status,targeted preservation strategies are proposed.The contribution of this framework is saving the cost of traditional field research and significantly improving the efficiency and scope of the evaluations.Besides,the results also fill the gap in evaluating the preservation status and analyzing their influencing mechanisms of the traditional villages in Tibetan-Qiang region,and support the decision makers to propose more targeted optimization strategies.展开更多
Modeling the spatial distribution of soil heavy metals is important in determining the safety of contaminated soils for agricultural use. This study utilized 60 topsoil samples (0 - 30 cm), multispectral images (Senti...Modeling the spatial distribution of soil heavy metals is important in determining the safety of contaminated soils for agricultural use. This study utilized 60 topsoil samples (0 - 30 cm), multispectral images (Sentinel-2), spectral indices, and ancillary data to model the spatial distribution of heavy metals in the soils along the Nairobi River. The model was generated using the Random Forest package in R. Using R2 to assess the prediction accuracy, the Random Forest model generated satisfactory results for all the elements. It also ranked the variables in order of their importance in the overall prediction. Spectral indices were the most important variables within the rankings. From the predicted topsoil maps, there were high concentrations of Cadmium on the easterly end of the river. Cadmium is an impurity in detergents, and this section is in close proximity to the Nairobi water sewerage plant, which could be a direct source of Cadmium. Some farms had Zinc levels which were above the World Health Organization recommended limit. The Random Forest model performed satisfactorily. However, the predictions can be improved further if the spatial resolutions of the various variables are increased and through the addition of more predictor variables.展开更多
The“Yarlung Zangbo River,Lhasa River and Nyangqu River”(YLN)region is the main grain producing area on which the Tibetan people depend for survival.The densities of soil organic carbon(SOC),total nitrogen(TN)and tot...The“Yarlung Zangbo River,Lhasa River and Nyangqu River”(YLN)region is the main grain producing area on which the Tibetan people depend for survival.The densities of soil organic carbon(SOC),total nitrogen(TN)and total phosphorus(TP)in farmlands are closely related to grain production.Scientific management and regulation of these nutrient densities are of great significance for ensuring food security.However,accurate simulations of spatial variations in the densities of SOC(SOCD),TN(TND)and TP(TPD)and the spatial distributions of SOCD,TND and TPD are still unclear.In this study,388 samples of cultivated soils at 0–10 and 10–20 cm in the YLN region were collected to determine the SOC,TN,and TP contents,as well as pH and bulk density(BD).Random forest models of SOCD,TND and TPD were constructed using longitude,latitude,elevation,mean annual temperature,mean annual precipitation,mean annual radiation and vegetation index,which were then used to obtain the spatial distribution maps of SOCD,TND and TPD,and the storages of SOC(SOCS),TN(TNS)and TP(TPS).Mean annual radiation can partially explain the spatial variations of SOCD and TND,in addition to temperature and precipitation.The relative biases between modelled and observed SOCD,TND,TPD,SOCS,TNS and TPS ranged from–9.43%to 7.57%.The SOCD and TND increased from west to east,but they were both low in the middle and high in the north and south.The SOCD and TND decreased with increasing pH and BD.SOCD,TND and TPD were low at mid-elevations but high at low and high elevations.The SOCD,TND,TPD,SOCS,TNS and TPS were 2.72 kg m-2,0.30 kg m-2,0.18 kg m-2,4.88 Tg,0.54 Tg and 0.32 Tg,respectively,at 0–20 cm over the cultivated lands of the YLN region.Based on these results,the random forest models constructed in this study can be used for subsequent related studies.Besides warming and precipitation changes,radiation changes can also affect SOCD and TND.In terms of the production of food crops such as highland barley,the farmland soils in the YLN region currently can have relative deficiencies of nitrogen and phosphorus nutrients.In the future,measures such as increasing the application of organic fertilizers should be taken to improve the carbon sequestration capacity and nitrogen and phosphorus nutrition of the soil.These findings have important guiding significance for the fertilization management of cultivated lands in the YLN region and other alpine regions similar to the YLN region.展开更多
A machine learning-based APP may quickly and non-destructively evaluate the quality of parameters,such as hardness and anthocyanin content in blue honeysuckle berries(Lonicera caerulea L.,BHB),based on changes in peri...A machine learning-based APP may quickly and non-destructively evaluate the quality of parameters,such as hardness and anthocyanin content in blue honeysuckle berries(Lonicera caerulea L.,BHB),based on changes in pericarp color characteristics.The color feature information of the BHB pericarp was extracted,and the corresponding hardness and anthocyanin content were determined at various growing stages.Correlation analysis of BHB quality indexes was conducted by single and combined components of BHB epidermal color features.The results showed that fruit hardness had a significantly negative correlation with color feature parameter R-G,and its anthocyanin content had a significantly positive correlation with color feature parameter R.Comparing the eight models,random forest(RF)was established to evaluate the hardness and anthocyanin content of BHB according to the correlation between pericarp color features and hardness and anthocyanin content on BHB quality evaluation APP on the WeChat platform.The credibility of APP embedding RF model for evaluating hardness and anthocyanin content in BHB was validated with the determination coefficient of 0.89 and 0.93 in practice.This approach could efficiently and conveniently evaluate the quality indexes of BHB in real time and serve as a technical reference for the detection of quality indicators of other berries using smartphones.展开更多
Canopy Nitrogen Concentration (CNC) is a key indicator of crop yields. It is feasible to establish a real- time regional model to estimate CNC by upscaling the field-scale spectral model. This study focuses on monit...Canopy Nitrogen Concentration (CNC) is a key indicator of crop yields. It is feasible to establish a real- time regional model to estimate CNC by upscaling the field-scale spectral model. This study focuses on monitoring the CNC in rice on a large scale in real-time. The Random Forest (RF) algorithm is used to establish the CNC spectral inversion model, and some vegetation indexes that are sensitive to nitrogen were selected as input parameters for the RF. CNC was selected as an output parameter. The hyperspectral and biochemical data were collected in a paddy in Changchun City, Jilin Province, China, and the data in Suzhou was used to test the model's universality and effectiveness. Two regional-scale models were developed by applying scale transformation based on the input and output variables respectively. The results show that the RFCNC model (CNC spectral inversion model based on the RF algorithm) performed accurately and significantly improved upon existing methods. R2, used to validate method accuracy in Changchun and Suzhou, was 0.82 and 0.73 respectively. The regional application accuracy increased (R2= 0. 81) through the two upscaling methods using hyperspectral remote sensing satellite images. This study suggests that this method is promising for estimating regional CNC in rice by upscaling a field-scale spectral model if the strategy is appropriately selected.展开更多
To improve the efficiency of air quality analysis and the accuracy of predictions, this paper proposes a composite method based on Vector Autoregressive (VAR) and Random Forest (RF) models. In the theoretical section,...To improve the efficiency of air quality analysis and the accuracy of predictions, this paper proposes a composite method based on Vector Autoregressive (VAR) and Random Forest (RF) models. In the theoretical section, the model introduction and estimation algorithms are provided. In the empirical analysis section, global air quality data from 2022 to 2024 are used, and the proposed method is applied. Specifically, principal component analysis (PCA) is first conducted, and then VAR and Random Forest methods are used for prediction on the reduced-dimensional data. The results show that the RMSE of the hybrid model is 45.27, significantly lower than the 49.11 of the VAR model alone, verifying its superiority. The stability and predictive performance of the model are effectively enhanced.展开更多
Evapotranspiration(ET)is a core parameter of the hydrology and carbon cycles,and its accurate estimation is crucial for water resource management.Satellite-based ET products provide an effective means for large-scale ...Evapotranspiration(ET)is a core parameter of the hydrology and carbon cycles,and its accurate estimation is crucial for water resource management.Satellite-based ET products provide an effective means for large-scale monitoring.However,due to limitations in the spatial and temporal resolution,the use of these products at regional and field scales is limited.In this study,daily ET in the Baoding Plain—a key groundwater resource recharge area—was estimated at a 500 m spatial resolution using the Bayesian Model Averaging(BMA)method.The model was driven by a synthesis of remote sensing datasets,reanalysis products,and interpolated data from meteorological stations.Validation results from in-situ observations indicated that the BMA ET had better performance(R=0.83,RMSE=1.25 mm/d)than each model in the BMA scheme.The spatiotemporal analysis revealed that the average annual BMA ET in the Baoding Plain was 683 mm/year from 2000 to 2019.Seasonal and monthly variations in the BMA ET captured the irrigation and water consumption patterns of the local crop rotation systems.A significant increasing trend of BMA ET(2.40 mm/year2)was observed in the Baoding Plain over the study period.At the regional scale,ET over more than 50%of the plain exhibited a significant positive trend.Further analysis identified water availability,solar radiation,and temperature as the primary drivers of ET variation.The BMA ET product generated in this study is characterized by high spatiotemporal resolution and accuracy.This reliable,highresolution dataset offers valuable support for precision agricultural water management and hydrological studies,including groundwater investigations,in this predominantly agricultural region.展开更多
The car-following models are the research basis of traffic flow theory and microscopic traffic simulation. Among the previous work, the theory-driven models are dominant, while the data-driven ones are relatively rare...The car-following models are the research basis of traffic flow theory and microscopic traffic simulation. Among the previous work, the theory-driven models are dominant, while the data-driven ones are relatively rare. In recent years, the related technologies of Intelligent Transportation System (ITS) re- presented by the Vehicles to Everything (V2X) technology have been developing rapidly. Utilizing the related technologies of ITS, the large-scale vehicle microscopic trajectory data with high quality can be acquired, which provides the research foundation for modeling the car-following behavior based on the data-driven methods. According to this point, a data-driven car-following model based on the Random Forest (RF) method was constructed in this work, and the Next Generation Simulation (NGSIM) dataset was used to calibrate and train the constructed model. The Artificial Neural Network (ANN) model, GM model, and Full Velocity Difference (FVD) model are em- ployed to comparatively verify the proposed model. The research results suggest that the model proposed in this work can accurately describe the car- following behavior with better performance under multiple performance indicators.展开更多
Accurate estimation of understory terrain has significant scientific importance for maintaining ecosystem balance and biodiversity conservation.Addressing the issue of inadequate representation of spatial heterogeneit...Accurate estimation of understory terrain has significant scientific importance for maintaining ecosystem balance and biodiversity conservation.Addressing the issue of inadequate representation of spatial heterogeneity when traditional forest topographic inversion methods consider the entire forest as the inversion unit,this study pro⁃poses a differentiated modeling approach to forest types based on refined land cover classification.Taking Puerto Ri⁃co and Maryland as study areas,a multi-dimensional feature system is constructed by integrating multi-source re⁃mote sensing data:ICESat-2 spaceborne LiDAR is used to obtain benchmark values for understory terrain,topo⁃graphic factors such as slope and aspect are extracted based on SRTM data,and vegetation cover characteristics are analyzed using Landsat-8 multispectral imagery.This study incorporates forest type as a classification modeling con⁃dition and applies the random forest algorithm to build differentiated topographic inversion models.Experimental re⁃sults indicate that,compared to traditional whole-area modeling methods(RMSE=5.06 m),forest type-based classi⁃fication modeling significantly improves the accuracy of understory terrain estimation(RMSE=2.94 m),validating the effectiveness of spatial heterogeneity modeling.Further sensitivity analysis reveals that canopy structure parame⁃ters(with RMSE variation reaching 4.11 m)exert a stronger regulatory effect on estimation accuracy compared to forest cover,providing important theoretical support for optimizing remote sensing models of forest topography.展开更多
Height–diameter relationships are essential elements of forest assessment and modeling efforts.In this work,two linear and eighteen nonlinear height–diameter equations were evaluated to find a local model for Orient...Height–diameter relationships are essential elements of forest assessment and modeling efforts.In this work,two linear and eighteen nonlinear height–diameter equations were evaluated to find a local model for Oriental beech(Fagus orientalis Lipsky) in the Hyrcanian Forest in Iran.The predictive performance of these models was first assessed by different evaluation criteria: adjusted R^2(R^2_(adj)),root mean square error(RMSE),relative RMSE(%RMSE),bias,and relative bias(%bias) criteria.The best model was selected for use as the base mixed-effects model.Random parameters for test plots were estimated with different tree selection options.Results show that the Chapman–Richards model had better predictive ability in terms of adj R^2(0.81),RMSE(3.7 m),%RMSE(12.9),bias(0.8),%Bias(2.79) than the other models.Furthermore,the calibration response,based on a selection of four trees from the sample plots,resulted in a reduction percentage for bias and RMSE of about 1.6–2.7%.Our results indicate that the calibrated model produced the most accurate results.展开更多
Most conventional oilfields in Eastern China with waterflood operations have reached ultra-high water cut in recent decade.The high water injection demand and produced water treatment cost pose significant environment...Most conventional oilfields in Eastern China with waterflood operations have reached ultra-high water cut in recent decade.The high water injection demand and produced water treatment cost pose significant environmental threats.Therefore,optimizing waterflood performance is key to improving production efficiency.This study performs data analysis on waterflood operations of all the oilfields operated by Sinopec across Eastern China.The production mechanisms and most effective operations for different reservoir types at diverse production stages are identified using data-driven methods.Random Forest models(RFMs)are constructed and integrated with Shapley Additive exPlanations(SHAP)analysis to quantify the weights and patterns of key geological and engineering features.A comparison of the estimated ultimate recovery factors for different blocks shows that geological factors play dominant roles in medium-to-high permeability reservoirs while development parameters are more critical for low-permeability reservoirs.The analysis of temporal data regarding field development and production history is conducted to select oil production-increasing operations in blocks.The results show that the most influential field operations vary for the diverse production stages,and well patterns should be carefully designed to improve production efficiency and reduce ineffective water circulation.展开更多
Rapid urban expansion has exerted substantial adverse effects on ecosystems.Improving urban land use eco-efficiency(ULUEE)is crucial to achieving Sustainable Development Goal(SDG)11.While previous endeavors have predo...Rapid urban expansion has exerted substantial adverse effects on ecosystems.Improving urban land use eco-efficiency(ULUEE)is crucial to achieving Sustainable Development Goal(SDG)11.While previous endeavors have predominantly concentrated on the linear correlations of the drivers,there has been a conspicuous absence of discourse on their nonlinear and heterogeneous impacts,which are more reflective of real-world complexities.In this study,the ULUEE across the Yellow River Basin,China was assessed from 2005 to 2020,its spatial-temporal dynamic evolution was revealed,and the nonlinear impact mechanisms were elucidated.The results indicate that the ULUEE increased from 0.4843 to 0.7822,and the spatial-temporal pattern exhibited strong stability and transfer inertia.Among the determinants,terrain exerted the most pronounced impact,whereas industrial structure had the least influence.Overall,terrain and economic development vitality presented W-shaped and U-shaped relationships,respectively.Urban land use structure had a detrimental effect,whereas altitude,industrial structure,economic output,social consumption,financial support,and science and technology development exerted positive influences.Notably,the influencing factors also displayed heterogeneous effects across the sub-basins.This study highlights the need for localized and targeted policy implementation to achieve efficient and sustainable urban land use.展开更多
There is a lack of studies when dealing with the comparison between regression methods and machine learning(ML)-type methods in terms of their ability to interpret and describe how the components of a bituminous mixtu...There is a lack of studies when dealing with the comparison between regression methods and machine learning(ML)-type methods in terms of their ability to interpret and describe how the components of a bituminous mixture affect mechanistic performance.At the same time,artificial intelligence(AI)-driven approaches are becoming more popular in analysing asphalt mixtures,yet there are limited comparisons of regression and machine learning(ML)models for mechanistic performance interpretation.Consequently,a comparison of AI and statistical approaches is presented in this study for predicting bituminous mixture properties such as stiffness,fatigue resistance,and tensile strength.Some of the important input features are bitumen content,crumb rubber content,and air void content.The research uses random forest model(RFM),linear regression model(LRM),and polynomial regression model(PRM).RFM and PRM achieved an R2 as high as 0.94,with mean absolute error(MAE)less than 2.5,and are,therefore,good predictive models.Interestingly,RFM works best in one-third of instances,particularly when dealing with outliers,whereas traditional statistical models work better in two-thirds of instances.The results highlight AI's value in bituminous mixture optimisation,where RFM showed good prediction accuracy.In 30%of the cases,AI models outperformed the conventional statistical approaches.At the same time,analyses show that model performance varies significantly with scenarios and that even if AI models capture complex nonlinear relationships,they must not override DOE principles.展开更多
摘要Slope units are divided according to the real topography and have clear geological characteristics,making them ideal units for evaluating the susceptibility to geological disasters.Based on the results of automatically and manually corrected hydrological slope unit division,the Longhua District,Shenzhen City,Guangdong Province,was selected as the study area.A total of 15 influencing factors,namely Fluctuation,slope,slope aspect,curvature,topographic witness index(TWI),stream power index(SPI),topographic roughness index(TRI),annual average rainfall,distance to water system,engineering rock group,distance to fault,land use,normalized difference vegetation index(NDVI),nighttime light,and distance to road,were selected as evaluation indicators.The information volume model(IV)and random points were used to select non-geological disaster units,and then the random forest model(RF)was used to evaluate the susceptibility to geological disasters.The automatic slope unit and the hydrological slope unit were compared and analyzed in the random forest and information volume random forest models.The results show that the area under the curve(AUC)values of the automatic slope unit evaluation results are 0.931 for the IV-RF model and 0.716 for the RF model,which are 0.6%(IV-RF model)and 1.9%(RF model)higher than those for the hydrological slope unit.Based on a comparison of the evaluation methods based on the two types of slope units,the hydrological slope unit evaluation method based on manual correction is highly subjective,is complicated to operate,and has a low evaluation accuracy,whereas the evaluation method based on automatic slope unit division is efficient and accurate,is suitable for large-scale efficient geological disaster evaluation,and can better deal with the problem of geological disaster susceptibility evaluation.
基金funded by Institutional Fund Projects under grant no.(IFPDP-261-22)。
摘要Detecting cyber attacks in networks connected to the Internet of Things(IoT)is of utmost importance because of the growing vulnerabilities in the smart environment.Conventional models,such as Naive Bayes and support vector machine(SVM),as well as ensemble methods,such as Gradient Boosting and eXtreme gradient boosting(XGBoost),are often plagued by high computational costs,which makes it challenging for them to perform real-time detection.In this regard,we suggested an attack detection approach that integrates Visual Geometry Group 16(VGG16),Artificial Rabbits Optimizer(ARO),and Random Forest Model to increase detection accuracy and operational efficiency in Internet of Things(IoT)networks.In the suggested model,the extraction of features from malware pictures was accomplished with the help of VGG16.The prediction process is carried out by the random forest model using the extracted features from the VGG16.Additionally,ARO is used to improve the hyper-parameters of the random forest model of the random forest.With an accuracy of 96.36%,the suggested model outperforms the standard models in terms of accuracy,F1-score,precision,and recall.The comparative research highlights our strategy’s success,which improves performance while maintaining a lower computational cost.This method is ideal for real-time applications,but it is effective.
基金National Natural Science Foundation of China,No.42071167,No.42201197,No.40871073The Second Tibetan Plateau Scientific Expedition and Research Program,No.2019QZKK0406Natural Science Foundation of Hebei Province,No.D2007000272。
摘要Random forest model is the mainstream research method used to accurately describe the distribution law and impact mechanism of regional population.We took Shijiazhuang as the research area,with comprehensive zoning based on endowments as the modeling unit,conducted stratified sampling on a hectare grid cell,and systematically carried out incremental selection experiments of population density impact factors,optimizing the population density random forest model throughout the process(zonal modeling,stratified sampling,factor selection,weighted output).The results are as follows:(1)Zonal modeling addresses the issue of confusion in population distribution laws caused by a single model.Sampling on a grid cell not only ensures the quality of training data by avoiding the modifiable areal unit problem(MAUP)but also attempts to mitigate the adverse effects of the ecological fallacy.Stratified sampling ensures the stability of population density label values(target variable)in the training sample.(2)Zonal selection experiments on population density impact factors help identify suitable combinations of factors,leading to a significant improvement in the goodness of fit(R2)of the zonal models.(3)Weighted combination output of the population density prediction dataset substantially enhances the model's robustness.(4)The population density dataset exhibits multi-scale superposition characteristics.On a large scale,the population density in plains is higher than that in mountainous areas,while on a small scale,urban areas have higher density compared to rural areas.The optimization scheme for the population density random forest model that we propose offers a unified technical framework for uncovering local population distribution law and the impact mechanisms.
摘要Potential of the Random Forest Model on mapping of different desertification processes was studied in Muttuma watershed of mid-Murrumbidgee river region of New South Wales,Australia.Desertification vulnerability index was developed using climate,terrain,vegetation,soil and land quality indices to identify environmentally sensitive areas for desertification.Random Forest Model(RFM)was used to predict the different desertification processes such as soil erosion,salinization and waterlogging in the watershed and the information needed to train classification algorithms was obtained from satellite imagery interpretation and ground truth data.Climatic factors(evaporation,rainfall,temperature),terrain factors(aspect,slope,slope length,steepness,and wetness index),soil properties(pH,organic carbon,clay and sand content)and vulnerability indices were used as an explanatory variable.Classification accuracy and kappa index were calculated for training and testing datasets.We recorded an overall accuracy rate of 87.7%and 72.1%for training and testing sites,respectively.We found larger discrepancies between overall accuracy rate and kappa index for testing datasets(72.2%and 27.5%,respectively)suggesting that all the classes are not predicted well.The prediction of soil erosion and no desertification process was good and poor for salinization and water-logging process.Overall,the results observed give a new idea of using the knowledge of desertification process in training areas that can be used to predict the desertification processes at unvisited areas.
基金supported by the grants from the Natural Science Foundation of Hubei Province(No.2020CFB780)the Fundamental Research Funds for the Central Universities(No.2017KFYXJJ020).
摘要Objective Body fluid mixtures are complex biological samples that frequently occur in crime scenes,and can provide important clues for criminal case analysis.DNA methylation assay has been applied in the identification of human body fluids,and has exhibited excellent performance in predicting single-source body fluids.The present study aims to develop a methylation SNaPshot multiplex system for body fluid identification,and accurately predict the mixture samples.In addition,the value of DNA methylation in the prediction of body fluid mixtures was further explored.Methods In the present study,420 samples of body fluid mixtures and 250 samples of single body fluids were tested using an optimized multiplex methylation system.Each kind of body fluid sample presented the specific methylation profiles of the 10 markers.Results Significant differences in methylation levels were observed between the mixtures and single body fluids.For all kinds of mixtures,the Spearman’s correlation analysis revealed a significantly strong correlation between the methylation levels and component proportions(1:20,1:10,1:5,1:1,5:1,10:1 and 20:1).Two random forest classification models were trained for the prediction of mixture types and the prediction of the mixture proportion of 2 components,based on the methylation levels of 10 markers.For the mixture prediction,Model-1 presented outstanding prediction accuracy,which reached up to 99.3%in 427 training samples,and had a remarkable accuracy of 100%in 243 independent test samples.For the mixture proportion prediction,Model-2 demonstrated an excellent accuracy of 98.8%in 252 training samples,and 98.2%in 168 independent test samples.The total prediction accuracy reached 99.3%for body fluid mixtures and 98.6%for the mixture proportions.Conclusion These results indicate the excellent capability and powerful value of the multiplex methylation system in the identification of forensic body fluid mixtures.
摘要This study proposes a framework to automate the evaluation of traditional village preservation status and analyze the major influential factors(MIFs)of preservation status and influencing mechanisms through YOLOv10 model and Random Forest model,taking Tibetan-Qiang region of northwest Sichuan as the study area.The framework adopts satellite maps,based on the YOLOv10 model,to comprehensively detect the preservation status of houses in the traditional villages,and calculates the preservation score of the corresponding villages as an evaluation of their preservation status.Further,through the Feature importance of Random Forest model,the MIFs of the village preservation status are filtered from the multiple environmental factors,and the SHAP value resolves the influencing intensity of the MIFs on the preservation status.Finally,for villages with poor preservation status,targeted preservation strategies are proposed.The contribution of this framework is saving the cost of traditional field research and significantly improving the efficiency and scope of the evaluations.Besides,the results also fill the gap in evaluating the preservation status and analyzing their influencing mechanisms of the traditional villages in Tibetan-Qiang region,and support the decision makers to propose more targeted optimization strategies.
摘要Modeling the spatial distribution of soil heavy metals is important in determining the safety of contaminated soils for agricultural use. This study utilized 60 topsoil samples (0 - 30 cm), multispectral images (Sentinel-2), spectral indices, and ancillary data to model the spatial distribution of heavy metals in the soils along the Nairobi River. The model was generated using the Random Forest package in R. Using R2 to assess the prediction accuracy, the Random Forest model generated satisfactory results for all the elements. It also ranked the variables in order of their importance in the overall prediction. Spectral indices were the most important variables within the rankings. From the predicted topsoil maps, there were high concentrations of Cadmium on the easterly end of the river. Cadmium is an impurity in detergents, and this section is in close proximity to the Nairobi water sewerage plant, which could be a direct source of Cadmium. Some farms had Zinc levels which were above the World Health Organization recommended limit. The Random Forest model performed satisfactorily. However, the predictions can be improved further if the spatial resolutions of the various variables are increased and through the addition of more predictor variables.
基金The Lhasa Science and Technology Plan Project(LSKJ202422)The Xizang Autonomous Region Science and Technology Project(XZ202501ZY0056)。
摘要The“Yarlung Zangbo River,Lhasa River and Nyangqu River”(YLN)region is the main grain producing area on which the Tibetan people depend for survival.The densities of soil organic carbon(SOC),total nitrogen(TN)and total phosphorus(TP)in farmlands are closely related to grain production.Scientific management and regulation of these nutrient densities are of great significance for ensuring food security.However,accurate simulations of spatial variations in the densities of SOC(SOCD),TN(TND)and TP(TPD)and the spatial distributions of SOCD,TND and TPD are still unclear.In this study,388 samples of cultivated soils at 0–10 and 10–20 cm in the YLN region were collected to determine the SOC,TN,and TP contents,as well as pH and bulk density(BD).Random forest models of SOCD,TND and TPD were constructed using longitude,latitude,elevation,mean annual temperature,mean annual precipitation,mean annual radiation and vegetation index,which were then used to obtain the spatial distribution maps of SOCD,TND and TPD,and the storages of SOC(SOCS),TN(TNS)and TP(TPS).Mean annual radiation can partially explain the spatial variations of SOCD and TND,in addition to temperature and precipitation.The relative biases between modelled and observed SOCD,TND,TPD,SOCS,TNS and TPS ranged from–9.43%to 7.57%.The SOCD and TND increased from west to east,but they were both low in the middle and high in the north and south.The SOCD and TND decreased with increasing pH and BD.SOCD,TND and TPD were low at mid-elevations but high at low and high elevations.The SOCD,TND,TPD,SOCS,TNS and TPS were 2.72 kg m-2,0.30 kg m-2,0.18 kg m-2,4.88 Tg,0.54 Tg and 0.32 Tg,respectively,at 0–20 cm over the cultivated lands of the YLN region.Based on these results,the random forest models constructed in this study can be used for subsequent related studies.Besides warming and precipitation changes,radiation changes can also affect SOCD and TND.In terms of the production of food crops such as highland barley,the farmland soils in the YLN region currently can have relative deficiencies of nitrogen and phosphorus nutrients.In the future,measures such as increasing the application of organic fertilizers should be taken to improve the carbon sequestration capacity and nitrogen and phosphorus nutrition of the soil.These findings have important guiding significance for the fertilization management of cultivated lands in the YLN region and other alpine regions similar to the YLN region.
基金Supported by the National Natural Science Foundation of China(32072352)the National Key Research and Development Program Project of China(2022YFD1600500)。
摘要A machine learning-based APP may quickly and non-destructively evaluate the quality of parameters,such as hardness and anthocyanin content in blue honeysuckle berries(Lonicera caerulea L.,BHB),based on changes in pericarp color characteristics.The color feature information of the BHB pericarp was extracted,and the corresponding hardness and anthocyanin content were determined at various growing stages.Correlation analysis of BHB quality indexes was conducted by single and combined components of BHB epidermal color features.The results showed that fruit hardness had a significantly negative correlation with color feature parameter R-G,and its anthocyanin content had a significantly positive correlation with color feature parameter R.Comparing the eight models,random forest(RF)was established to evaluate the hardness and anthocyanin content of BHB according to the correlation between pericarp color features and hardness and anthocyanin content on BHB quality evaluation APP on the WeChat platform.The credibility of APP embedding RF model for evaluating hardness and anthocyanin content in BHB was validated with the determination coefficient of 0.89 and 0.93 in practice.This approach could efficiently and conveniently evaluate the quality indexes of BHB in real time and serve as a technical reference for the detection of quality indicators of other berries using smartphones.
基金Foundation: National Natural Science Foundation of China (No.41371407, 40771155)
摘要Canopy Nitrogen Concentration (CNC) is a key indicator of crop yields. It is feasible to establish a real- time regional model to estimate CNC by upscaling the field-scale spectral model. This study focuses on monitoring the CNC in rice on a large scale in real-time. The Random Forest (RF) algorithm is used to establish the CNC spectral inversion model, and some vegetation indexes that are sensitive to nitrogen were selected as input parameters for the RF. CNC was selected as an output parameter. The hyperspectral and biochemical data were collected in a paddy in Changchun City, Jilin Province, China, and the data in Suzhou was used to test the model's universality and effectiveness. Two regional-scale models were developed by applying scale transformation based on the input and output variables respectively. The results show that the RFCNC model (CNC spectral inversion model based on the RF algorithm) performed accurately and significantly improved upon existing methods. R2, used to validate method accuracy in Changchun and Suzhou, was 0.82 and 0.73 respectively. The regional application accuracy increased (R2= 0. 81) through the two upscaling methods using hyperspectral remote sensing satellite images. This study suggests that this method is promising for estimating regional CNC in rice by upscaling a field-scale spectral model if the strategy is appropriately selected.
摘要To improve the efficiency of air quality analysis and the accuracy of predictions, this paper proposes a composite method based on Vector Autoregressive (VAR) and Random Forest (RF) models. In the theoretical section, the model introduction and estimation algorithms are provided. In the empirical analysis section, global air quality data from 2022 to 2024 are used, and the proposed method is applied. Specifically, principal component analysis (PCA) is first conducted, and then VAR and Random Forest methods are used for prediction on the reduced-dimensional data. The results show that the RMSE of the hybrid model is 45.27, significantly lower than the 49.11 of the VAR model alone, verifying its superiority. The stability and predictive performance of the model are effectively enhanced.
基金supported by the National Natural Science Foundation of China grant ID U2244214by Tianjin Normal University Doctor Foundation(043-135202XB1605)Our thanks are also extended to AE and anonymous reviews for their insightful and constructive comments to improve the quality of this article,and Tian-Chyi Jim Yeh of the University of Arizona for his technical edit of this paper.
摘要Evapotranspiration(ET)is a core parameter of the hydrology and carbon cycles,and its accurate estimation is crucial for water resource management.Satellite-based ET products provide an effective means for large-scale monitoring.However,due to limitations in the spatial and temporal resolution,the use of these products at regional and field scales is limited.In this study,daily ET in the Baoding Plain—a key groundwater resource recharge area—was estimated at a 500 m spatial resolution using the Bayesian Model Averaging(BMA)method.The model was driven by a synthesis of remote sensing datasets,reanalysis products,and interpolated data from meteorological stations.Validation results from in-situ observations indicated that the BMA ET had better performance(R=0.83,RMSE=1.25 mm/d)than each model in the BMA scheme.The spatiotemporal analysis revealed that the average annual BMA ET in the Baoding Plain was 683 mm/year from 2000 to 2019.Seasonal and monthly variations in the BMA ET captured the irrigation and water consumption patterns of the local crop rotation systems.A significant increasing trend of BMA ET(2.40 mm/year2)was observed in the Baoding Plain over the study period.At the regional scale,ET over more than 50%of the plain exhibited a significant positive trend.Further analysis identified water availability,solar radiation,and temperature as the primary drivers of ET variation.The BMA ET product generated in this study is characterized by high spatiotemporal resolution and accuracy.This reliable,highresolution dataset offers valuable support for precision agricultural water management and hydrological studies,including groundwater investigations,in this predominantly agricultural region.
摘要The car-following models are the research basis of traffic flow theory and microscopic traffic simulation. Among the previous work, the theory-driven models are dominant, while the data-driven ones are relatively rare. In recent years, the related technologies of Intelligent Transportation System (ITS) re- presented by the Vehicles to Everything (V2X) technology have been developing rapidly. Utilizing the related technologies of ITS, the large-scale vehicle microscopic trajectory data with high quality can be acquired, which provides the research foundation for modeling the car-following behavior based on the data-driven methods. According to this point, a data-driven car-following model based on the Random Forest (RF) method was constructed in this work, and the Next Generation Simulation (NGSIM) dataset was used to calibrate and train the constructed model. The Artificial Neural Network (ANN) model, GM model, and Full Velocity Difference (FVD) model are em- ployed to comparatively verify the proposed model. The research results suggest that the model proposed in this work can accurately describe the car- following behavior with better performance under multiple performance indicators.
基金Supported by the National Natural Science Foundation of China(42401488,42071351)the National Key Research and Development Program of China(2020YFA0608501,2017YFB0504204)+4 种基金the Liaoning Revitalization Talents Program(XLYC1802027)the Talent Recruited Program of the Chinese Academy of Science(Y938091)the Project Supported Discipline Innovation Team of the Liaoning Technical University(LNTU20TD-23)the Liaoning Province Doctoral Research Initiation Fund Program(2023-BS-202)the Basic Research Projects of Liaoning Department of Education(JYTQN2023202)。
摘要Accurate estimation of understory terrain has significant scientific importance for maintaining ecosystem balance and biodiversity conservation.Addressing the issue of inadequate representation of spatial heterogeneity when traditional forest topographic inversion methods consider the entire forest as the inversion unit,this study pro⁃poses a differentiated modeling approach to forest types based on refined land cover classification.Taking Puerto Ri⁃co and Maryland as study areas,a multi-dimensional feature system is constructed by integrating multi-source re⁃mote sensing data:ICESat-2 spaceborne LiDAR is used to obtain benchmark values for understory terrain,topo⁃graphic factors such as slope and aspect are extracted based on SRTM data,and vegetation cover characteristics are analyzed using Landsat-8 multispectral imagery.This study incorporates forest type as a classification modeling con⁃dition and applies the random forest algorithm to build differentiated topographic inversion models.Experimental re⁃sults indicate that,compared to traditional whole-area modeling methods(RMSE=5.06 m),forest type-based classi⁃fication modeling significantly improves the accuracy of understory terrain estimation(RMSE=2.94 m),validating the effectiveness of spatial heterogeneity modeling.Further sensitivity analysis reveals that canopy structure parame⁃ters(with RMSE variation reaching 4.11 m)exert a stronger regulatory effect on estimation accuracy compared to forest cover,providing important theoretical support for optimizing remote sensing models of forest topography.
基金This research received no specific grant from any funding agency in the public,commercial,or not-for-profit sectors
摘要Height–diameter relationships are essential elements of forest assessment and modeling efforts.In this work,two linear and eighteen nonlinear height–diameter equations were evaluated to find a local model for Oriental beech(Fagus orientalis Lipsky) in the Hyrcanian Forest in Iran.The predictive performance of these models was first assessed by different evaluation criteria: adjusted R^2(R^2_(adj)),root mean square error(RMSE),relative RMSE(%RMSE),bias,and relative bias(%bias) criteria.The best model was selected for use as the base mixed-effects model.Random parameters for test plots were estimated with different tree selection options.Results show that the Chapman–Richards model had better predictive ability in terms of adj R^2(0.81),RMSE(3.7 m),%RMSE(12.9),bias(0.8),%Bias(2.79) than the other models.Furthermore,the calibration response,based on a selection of four trees from the sample plots,resulted in a reduction percentage for bias and RMSE of about 1.6–2.7%.Our results indicate that the calibrated model produced the most accurate results.
摘要Most conventional oilfields in Eastern China with waterflood operations have reached ultra-high water cut in recent decade.The high water injection demand and produced water treatment cost pose significant environmental threats.Therefore,optimizing waterflood performance is key to improving production efficiency.This study performs data analysis on waterflood operations of all the oilfields operated by Sinopec across Eastern China.The production mechanisms and most effective operations for different reservoir types at diverse production stages are identified using data-driven methods.Random Forest models(RFMs)are constructed and integrated with Shapley Additive exPlanations(SHAP)analysis to quantify the weights and patterns of key geological and engineering features.A comparison of the estimated ultimate recovery factors for different blocks shows that geological factors play dominant roles in medium-to-high permeability reservoirs while development parameters are more critical for low-permeability reservoirs.The analysis of temporal data regarding field development and production history is conducted to select oil production-increasing operations in blocks.The results show that the most influential field operations vary for the diverse production stages,and well patterns should be carefully designed to improve production efficiency and reduce ineffective water circulation.
基金Under the auspices of National Natural Science Foundation of China(No.42293271)。
摘要Rapid urban expansion has exerted substantial adverse effects on ecosystems.Improving urban land use eco-efficiency(ULUEE)is crucial to achieving Sustainable Development Goal(SDG)11.While previous endeavors have predominantly concentrated on the linear correlations of the drivers,there has been a conspicuous absence of discourse on their nonlinear and heterogeneous impacts,which are more reflective of real-world complexities.In this study,the ULUEE across the Yellow River Basin,China was assessed from 2005 to 2020,its spatial-temporal dynamic evolution was revealed,and the nonlinear impact mechanisms were elucidated.The results indicate that the ULUEE increased from 0.4843 to 0.7822,and the spatial-temporal pattern exhibited strong stability and transfer inertia.Among the determinants,terrain exerted the most pronounced impact,whereas industrial structure had the least influence.Overall,terrain and economic development vitality presented W-shaped and U-shaped relationships,respectively.Urban land use structure had a detrimental effect,whereas altitude,industrial structure,economic output,social consumption,financial support,and science and technology development exerted positive influences.Notably,the influencing factors also displayed heterogeneous effects across the sub-basins.This study highlights the need for localized and targeted policy implementation to achieve efficient and sustainable urban land use.
基金sustained them with this research(including Eng.Giuseppe Colicchio)and the European Commission for its financial contribution to the LIFE SILENT project“Sustainable Innovations for Long-life Environmental Noise Technologies”(LIFE22-ENV-IT-LIFE-SILENT/101114310.Acronym:LIFE22-ENV-ITLIFE SILENT)the LIFE SNEAK Project“Optimised Surfaces Against Noise and Vibrations Produced by Tramway Track and Road Traffic”(LIFE20 ENV/IT/000181.Acronym:LIFE SNEAK).
摘要There is a lack of studies when dealing with the comparison between regression methods and machine learning(ML)-type methods in terms of their ability to interpret and describe how the components of a bituminous mixture affect mechanistic performance.At the same time,artificial intelligence(AI)-driven approaches are becoming more popular in analysing asphalt mixtures,yet there are limited comparisons of regression and machine learning(ML)models for mechanistic performance interpretation.Consequently,a comparison of AI and statistical approaches is presented in this study for predicting bituminous mixture properties such as stiffness,fatigue resistance,and tensile strength.Some of the important input features are bitumen content,crumb rubber content,and air void content.The research uses random forest model(RFM),linear regression model(LRM),and polynomial regression model(PRM).RFM and PRM achieved an R2 as high as 0.94,with mean absolute error(MAE)less than 2.5,and are,therefore,good predictive models.Interestingly,RFM works best in one-third of instances,particularly when dealing with outliers,whereas traditional statistical models work better in two-thirds of instances.The results highlight AI's value in bituminous mixture optimisation,where RFM showed good prediction accuracy.In 30%of the cases,AI models outperformed the conventional statistical approaches.At the same time,analyses show that model performance varies significantly with scenarios and that even if AI models capture complex nonlinear relationships,they must not override DOE principles.