Prediction plays a vital role in decision making. Correct prediction leads to right decision making to save the life, energy,efforts, money and time. The right decision prevents physical and material losses and it is ...Prediction plays a vital role in decision making. Correct prediction leads to right decision making to save the life, energy,efforts, money and time. The right decision prevents physical and material losses and it is practiced in all the fields including medical,finance, environmental studies, engineering and emerging technologies. Prediction is carried out by a model called classifier. The predictive accuracy of the classifier highly depends on the training datasets utilized for training the classifier. The irrelevant and redundant features of the training dataset reduce the accuracy of the classifier. Hence, the irrelevant and redundant features must be removed from the training dataset through the process known as feature selection. This paper proposes a feature selection algorithm namely unsupervised learning with ranking based feature selection(FSULR). It removes redundant features by clustering and eliminates irrelevant features by statistical measures to select the most significant features from the training dataset. The performance of this proposed algorithm is compared with the other seven feature selection algorithms by well known classifiers namely naive Bayes(NB),instance based(IB1) and tree based J48. Experimental results show that the proposed algorithm yields better prediction accuracy for classifiers.展开更多
Epilepsy is a chronic neurological disorder characterized by recurrent seizures,posing significant challenges to patients’quality of life.Accurate classification of seizure states is crucial for effective interventio...Epilepsy is a chronic neurological disorder characterized by recurrent seizures,posing significant challenges to patients’quality of life.Accurate classification of seizure states is crucial for effective intervention.This paper presents a deep learning-based approach for epileptic seizure classification by integrating multi-feature analysis of electroencephalogram(EEG)signals.The proposed method begins with signal preprocessing,including denoising,segmentation,and label construction.Subsequently,a comprehensive set of temporal,spectral,and wavelet-based features—such as signal mean,power,heart rate,and wavelet coefficients—is extracted.Feature selection is then performed using the Maximal Information Coefficient(MIC)to identify the most discriminative inputs.A hybrid model combining a Transformer encoder and a Long Short-Term Memory(LSTM)network is developed to effectively capture both long-range dependencies and temporal dynamics in EEG sequences for seizure classification.Evaluated on the Bonn dataset using 5-fold cross-validation,the proposed method achieves an accuracy of 96.43%in distinguishing between epileptic patients and healthy subjects,with a sensitivity of 97.53%in detecting seizure states.It also attains a multi-class classification accuracy of 90.14%across different epileptic signal types.Ablation studies confirm that MICbased feature selection improves accuracy by over 2O%compared to using raw features without selection.The results demonstrate that the integration of multi-feature analysis with the Transformer-LSTM architecture offers an effective and reliable solution for EEG-based seizure classification.展开更多
With the increasing dimensionality of the data,High-dimensional Feature Selection(HFS)becomes an increasingly dif-ficult task.It is not simple to find the best subset of features due to the breadth of the search space...With the increasing dimensionality of the data,High-dimensional Feature Selection(HFS)becomes an increasingly dif-ficult task.It is not simple to find the best subset of features due to the breadth of the search space and the intricacy of the interactions between features.Many of the Feature Selection(FS)approaches now in use for these problems perform sig-nificantly less well when faced with such intricate situations involving high-dimensional search spaces.It is demonstrated that meta-heuristic algorithms can provide sub-optimal results in an acceptable amount of time.This paper presents a new binary Boosted version of the Spider Wasp Optimizer(BSWO)called Binary Boosted SWO(BBSWO),which combines a number of successful and promising strategies,in order to deal with HFS.The shortcomings of the original BSWO,including early convergence,settling into local optimums,limited exploration and exploitation,and lack of population diversity,were addressed by the proposal of this new variant of SWO.The concept of chaos optimization is introduced in BSWO,where initialization is consistently produced by utilizing the properties of sine chaos mapping.A new convergence parameter was then incorporated into BSWO to achieve a promising balance between exploration and exploitation.Multiple exploration mechanisms were then applied in conjunction with several exploitation strategies to effectively enrich the search process of BSWO within the search space.Finally,quantum-based optimization was added to enhance the diversity of the search agents in BSWO.The proposed BBSWO not only offers the most suitable subset of features located,but it also lessens the data's redundancy structure.BBSWO was evaluated using the k-Nearest Neighbor(k-NN)classifier on 23 HFS problems from the biomedical domain taken from the UCI repository.The results were compared with those of traditional BSWO and other well-known meta-heuristics-based FS.The findings indicate that,in comparison to other competing techniques,the proposed BBSWO can,on average,identify the least significant subsets of features with efficient classification accuracy of the k-NN classifier.展开更多
Soil organic matter(SOM)is a core indicator of soil fertility and ecosystem function.However,in regions where Mollisol and non-Mollisol coexist,high-precision spatial mapping faces significant challenges due to pronou...Soil organic matter(SOM)is a core indicator of soil fertility and ecosystem function.However,in regions where Mollisol and non-Mollisol coexist,high-precision spatial mapping faces significant challenges due to pronounced terrain heterogeneity and redundancy in high-dimensional covariates.This study proposes a"remote sensing zoning-feature selection optimization-random forest(RSZ-FSO-RF)"framework.By integrating Landsat-8 multi-temporal imagery from 2014-2023 with topographic and climatic factors,and leveraging the Google Earth Engine(GEE)platform,it achieves highprecision remote sensing zoning of Mollisol and non-Mollisol areas(overall accuracy:92.13%,Kappa coefficient:0.70).Subsequently,local Random Forest(RF)regression models were established within each zone for SOM prediction,with predictive variables optimized using recursive feature elimination(RFE).Results demonstrate that compared to FAOzone-based modeling,the RSZ-FSO-RF framework significantly enhances prediction accuracy(R2=0.619,RMSE=6.849 g kg-1).And further feature optimization continued to enhance model performance(R2=0.627,RMSE=6.781 g kg-1).Notably,optimal predictor combinations varied significantly across zones,with SOM spatial variability generally higher in non-Mollisol areas than in Mollisol regions.By organically integrating remote sensing zoning with feature selection,this framework effectively mitigates covariate redundancy while accounting for local heterogeneity,significantly enhancing the accuracy and stability of high-resolution SOM mapping.Furthermore,this study provides scientific basis and decision support for soil resource management and sustainable agricultural development under complex topographic conditions.展开更多
The potential of employing hyperspectral imaging(HSl)in the near-infrared(NiR)range(386.82-1,004.50 nm)for predicting the firmness of'Fuji'apples cultivated in Aksu has been evaluated.The performance of seven ...The potential of employing hyperspectral imaging(HSl)in the near-infrared(NiR)range(386.82-1,004.50 nm)for predicting the firmness of'Fuji'apples cultivated in Aksu has been evaluated.The performance of seven preprocessing algorithms and two feature selection algorithms was evaluated.The coefficient of determination(R2)and root mean square error(RMsE)of Partial Least Squares(PLS)models are contrasted using various inputs.These results confirm that the Multiplicative Scatter Correction(MSC)preprocessing algorithm was the optimal choice(R2p=0.7925,RMSEP=0.6537),and the Competitive Adaptive Reweighted Sampling(CARS)feature selection algorithm demonstrated superior performance(R2p=0.8325,RMSEP=0.6257).Based on the aforementioned findings,PLS,Multiple Linear Regression(MLR),Heterogeneous Transfer Learning(HTL),and Back Propagation Neural Network(BPNN)models were constructed for cross-validation purposes.The experimental results indicate that the CARS-BPNN model exhibits the optimal prediction performance,with an R2pvalue of 0.9350 and an RMSEP value of 0.4654.The results of the research indicated that a deep learning method combined with hyperspectral imaging technology could be utilized to non-destructively detect the firmness of'Fuji'apples,which will be beneficial and potentially applicable for post-harvest fruit firmness monitoring.This research provides a reference point for the non-destructive detection of apple in the selection of preprocessing,feature selection algorithms,and predicting firmness model.展开更多
Apple leaf disease is one of the main factors to constrain the apple production and quality.It takes a long time to detect the diseases by using the traditional diagnostic approach,thus farmers often miss the best tim...Apple leaf disease is one of the main factors to constrain the apple production and quality.It takes a long time to detect the diseases by using the traditional diagnostic approach,thus farmers often miss the best time to prevent and treat the diseases.Apple leaf disease recognition based on leaf image is an essential research topic in the field of computer vision,where the key task is to find an effective way to represent the diseased leaf images.In this research,based on image processing techniques and pattern recognition methods,an apple leaf disease recognition method was proposed.A color transformation structure for the input RGB(Red,Green and Blue)image was designed firstly and then RGB model was converted to HSI(Hue,Saturation and Intensity),YUV and gray models.The background was removed based on a specific threshold value,and then the disease spot image was segmented with region growing algorithm(RGA).Thirty-eight classifying features of color,texture and shape were extracted from each spot image.To reduce the dimensionality of the feature space and improve the accuracy of the apple leaf disease identification,the most valuable features were selected by combining genetic algorithm(GA)and correlation based feature selection(CFS).Finally,the diseases were recognized by SVM classifier.In the proposed method,the selected feature subset was globally optimum.The experimental results of more than 90%correct identification rate on the apple diseased leaf image database which contains 90 disease images for there kinds of apple leaf diseases,powdery mildew,mosaic and rust,demonstrate that the proposed method is feasible and effective.展开更多
Essential proteins are vital to the survival of a cell. There are various features related to the essentiality of proteins, such as biological and topological features. Many computational methods have been developed t...Essential proteins are vital to the survival of a cell. There are various features related to the essentiality of proteins, such as biological and topological features. Many computational methods have been developed to identify essential proteins by using these features. However, it is still a big challenge to design an effective method that is able to select suitable features and integrate them to predict essential proteins. In this work, we first collect 26 features, and use SVM-RFE to select some of them to create a feature space for predicting essential proteins, and then remove the features that share the biological meaning with other features in the feature space according to their Pearson Correlation Coefficients(PCC). The experiments are carried out on S. cerevisiae data. Six features are determined as the best subset of features. To assess the prediction performance of our method, we further compare it with some machine learning methods, such as SVM, Naive Bayes, Bayes Network, and NBTree when inputting the different number of features. The results show that those methods using the 6 features outperform that using other features, which confirms the effectiveness of our feature selection method for essential protein prediction.展开更多
In cloud computing Resource allocation is a very complex task.Handling the customer demand makes the challenges of on-demand resource allocation.Many challenges are faced by conventional methods for resource allocatio...In cloud computing Resource allocation is a very complex task.Handling the customer demand makes the challenges of on-demand resource allocation.Many challenges are faced by conventional methods for resource allocation in order tomeet the Quality of Service(QoS)requirements of users.For solving the about said problems a new method was implemented with the utility of machine learning framework of resource allocation by utilizing the cloud computing technique was taken in to an account in this research work.The accuracy in the machine learning algorithm can be improved by introducing Bat Algorithm with feature selection(BFS)in the proposed work,this further reduces the inappropriate features from the data.The similarities that were hidden can be demoralized by the Support Vector Machine(SVM)classifier which is also determine the subspace vector and then a new feature vector can be predicted by using SVM.For an unexpected circumstance SVM model can make a resource allocation decision.The efficiency of proposed SVM classifier of resource allocation can be highlighted by using a singlecell multiuser massive Multiple-Input Multiple Output(MIMO)system,with beam allocation problem as an example.The proposed resource allocation based on SVM performs efficiently than the existing conventional methods;this has been proven by analysing its results.展开更多
摘要Prediction plays a vital role in decision making. Correct prediction leads to right decision making to save the life, energy,efforts, money and time. The right decision prevents physical and material losses and it is practiced in all the fields including medical,finance, environmental studies, engineering and emerging technologies. Prediction is carried out by a model called classifier. The predictive accuracy of the classifier highly depends on the training datasets utilized for training the classifier. The irrelevant and redundant features of the training dataset reduce the accuracy of the classifier. Hence, the irrelevant and redundant features must be removed from the training dataset through the process known as feature selection. This paper proposes a feature selection algorithm namely unsupervised learning with ranking based feature selection(FSULR). It removes redundant features by clustering and eliminates irrelevant features by statistical measures to select the most significant features from the training dataset. The performance of this proposed algorithm is compared with the other seven feature selection algorithms by well known classifiers namely naive Bayes(NB),instance based(IB1) and tree based J48. Experimental results show that the proposed algorithm yields better prediction accuracy for classifiers.
基金supported by the National Key Research and Development Program of China(2020AAA0104905)in part by National Natural Science Foundation of China underGrant(62341118,62503241)+1 种基金in part by the Natural Science Foundation of Jiangsu Province of China under Grant(BK20250664)in part by Foundation of recruiting talents of HYIT under Grant(Z301B25508).
摘要Epilepsy is a chronic neurological disorder characterized by recurrent seizures,posing significant challenges to patients’quality of life.Accurate classification of seizure states is crucial for effective intervention.This paper presents a deep learning-based approach for epileptic seizure classification by integrating multi-feature analysis of electroencephalogram(EEG)signals.The proposed method begins with signal preprocessing,including denoising,segmentation,and label construction.Subsequently,a comprehensive set of temporal,spectral,and wavelet-based features—such as signal mean,power,heart rate,and wavelet coefficients—is extracted.Feature selection is then performed using the Maximal Information Coefficient(MIC)to identify the most discriminative inputs.A hybrid model combining a Transformer encoder and a Long Short-Term Memory(LSTM)network is developed to effectively capture both long-range dependencies and temporal dynamics in EEG sequences for seizure classification.Evaluated on the Bonn dataset using 5-fold cross-validation,the proposed method achieves an accuracy of 96.43%in distinguishing between epileptic patients and healthy subjects,with a sensitivity of 97.53%in detecting seizure states.It also attains a multi-class classification accuracy of 90.14%across different epileptic signal types.Ablation studies confirm that MICbased feature selection improves accuracy by over 2O%compared to using raw features without selection.The results demonstrate that the integration of multi-feature analysis with the Transformer-LSTM architecture offers an effective and reliable solution for EEG-based seizure classification.
基金supported from the Deanship of Research and Graduate Studies(DRG)at Ajman University,Ajman,UAE(Grant No.2023-IRG-ENIT-34).
摘要With the increasing dimensionality of the data,High-dimensional Feature Selection(HFS)becomes an increasingly dif-ficult task.It is not simple to find the best subset of features due to the breadth of the search space and the intricacy of the interactions between features.Many of the Feature Selection(FS)approaches now in use for these problems perform sig-nificantly less well when faced with such intricate situations involving high-dimensional search spaces.It is demonstrated that meta-heuristic algorithms can provide sub-optimal results in an acceptable amount of time.This paper presents a new binary Boosted version of the Spider Wasp Optimizer(BSWO)called Binary Boosted SWO(BBSWO),which combines a number of successful and promising strategies,in order to deal with HFS.The shortcomings of the original BSWO,including early convergence,settling into local optimums,limited exploration and exploitation,and lack of population diversity,were addressed by the proposal of this new variant of SWO.The concept of chaos optimization is introduced in BSWO,where initialization is consistently produced by utilizing the properties of sine chaos mapping.A new convergence parameter was then incorporated into BSWO to achieve a promising balance between exploration and exploitation.Multiple exploration mechanisms were then applied in conjunction with several exploitation strategies to effectively enrich the search process of BSWO within the search space.Finally,quantum-based optimization was added to enhance the diversity of the search agents in BSWO.The proposed BBSWO not only offers the most suitable subset of features located,but it also lessens the data's redundancy structure.BBSWO was evaluated using the k-Nearest Neighbor(k-NN)classifier on 23 HFS problems from the biomedical domain taken from the UCI repository.The results were compared with those of traditional BSWO and other well-known meta-heuristics-based FS.The findings indicate that,in comparison to other competing techniques,the proposed BBSWO can,on average,identify the least significant subsets of features with efficient classification accuracy of the k-NN classifier.
基金supported by the National Natural Science Foundation of China(42401460)the National Key R&D Program of China(2021YFD1500100)。
摘要Soil organic matter(SOM)is a core indicator of soil fertility and ecosystem function.However,in regions where Mollisol and non-Mollisol coexist,high-precision spatial mapping faces significant challenges due to pronounced terrain heterogeneity and redundancy in high-dimensional covariates.This study proposes a"remote sensing zoning-feature selection optimization-random forest(RSZ-FSO-RF)"framework.By integrating Landsat-8 multi-temporal imagery from 2014-2023 with topographic and climatic factors,and leveraging the Google Earth Engine(GEE)platform,it achieves highprecision remote sensing zoning of Mollisol and non-Mollisol areas(overall accuracy:92.13%,Kappa coefficient:0.70).Subsequently,local Random Forest(RF)regression models were established within each zone for SOM prediction,with predictive variables optimized using recursive feature elimination(RFE).Results demonstrate that compared to FAOzone-based modeling,the RSZ-FSO-RF framework significantly enhances prediction accuracy(R2=0.619,RMSE=6.849 g kg-1).And further feature optimization continued to enhance model performance(R2=0.627,RMSE=6.781 g kg-1).Notably,optimal predictor combinations varied significantly across zones,with SOM spatial variability generally higher in non-Mollisol areas than in Mollisol regions.By organically integrating remote sensing zoning with feature selection,this framework effectively mitigates covariate redundancy while accounting for local heterogeneity,significantly enhancing the accuracy and stability of high-resolution SOM mapping.Furthermore,this study provides scientific basis and decision support for soil resource management and sustainable agricultural development under complex topographic conditions.
基金supported by'Pioneer'and'Leading Goose'Research and Development Plan Project of Zhejiang Province(2022C04039)Major Scientific Research Achievement Transformation Project of Ningxia Hui Autonomous Region(2023CJE09060)+1 种基金Tianjin Science and Technology Program Project(22ZYCGSN00170,22ZYCGSN00470)International Exchanges Funds offered by the Royal Society(No.IEC\NSFC\233076).
摘要The potential of employing hyperspectral imaging(HSl)in the near-infrared(NiR)range(386.82-1,004.50 nm)for predicting the firmness of'Fuji'apples cultivated in Aksu has been evaluated.The performance of seven preprocessing algorithms and two feature selection algorithms was evaluated.The coefficient of determination(R2)and root mean square error(RMsE)of Partial Least Squares(PLS)models are contrasted using various inputs.These results confirm that the Multiplicative Scatter Correction(MSC)preprocessing algorithm was the optimal choice(R2p=0.7925,RMSEP=0.6537),and the Competitive Adaptive Reweighted Sampling(CARS)feature selection algorithm demonstrated superior performance(R2p=0.8325,RMSEP=0.6257).Based on the aforementioned findings,PLS,Multiple Linear Regression(MLR),Heterogeneous Transfer Learning(HTL),and Back Propagation Neural Network(BPNN)models were constructed for cross-validation purposes.The experimental results indicate that the CARS-BPNN model exhibits the optimal prediction performance,with an R2pvalue of 0.9350 and an RMSEP value of 0.4654.The results of the research indicated that a deep learning method combined with hyperspectral imaging technology could be utilized to non-destructively detect the firmness of'Fuji'apples,which will be beneficial and potentially applicable for post-harvest fruit firmness monitoring.This research provides a reference point for the non-destructive detection of apple in the selection of preprocessing,feature selection algorithms,and predicting firmness model.
基金Natural Science Foundation of China(grant Nos.61473237,61202170,and 61402331)It is also supported by the Shaanxi Provincial Natural Science Foundation Research Project(2014JM2-6096)+3 种基金Tianjin Research Program of Application Foundation and Advanced Technology(14JCYBJC42500)Tianjin science and technology correspondent project(16JCTPJC47300)the 2015 key projects of Tianjin science and technology support program(No.15ZCZDGX00200)the Fund of Tianjin Food Safety&Low Carbon Manufacturing Collaborative Innovation Center.
摘要Apple leaf disease is one of the main factors to constrain the apple production and quality.It takes a long time to detect the diseases by using the traditional diagnostic approach,thus farmers often miss the best time to prevent and treat the diseases.Apple leaf disease recognition based on leaf image is an essential research topic in the field of computer vision,where the key task is to find an effective way to represent the diseased leaf images.In this research,based on image processing techniques and pattern recognition methods,an apple leaf disease recognition method was proposed.A color transformation structure for the input RGB(Red,Green and Blue)image was designed firstly and then RGB model was converted to HSI(Hue,Saturation and Intensity),YUV and gray models.The background was removed based on a specific threshold value,and then the disease spot image was segmented with region growing algorithm(RGA).Thirty-eight classifying features of color,texture and shape were extracted from each spot image.To reduce the dimensionality of the feature space and improve the accuracy of the apple leaf disease identification,the most valuable features were selected by combining genetic algorithm(GA)and correlation based feature selection(CFS).Finally,the diseases were recognized by SVM classifier.In the proposed method,the selected feature subset was globally optimum.The experimental results of more than 90%correct identification rate on the apple diseased leaf image database which contains 90 disease images for there kinds of apple leaf diseases,powdery mildew,mosaic and rust,demonstrate that the proposed method is feasible and effective.
基金supported by the National Natural Science Foundation of China(Nos.61232001,61502166,61502214,61379108,and 61370024)Scientific Research Fund of Hunan Provincial Education Department(Nos.15CY007 and 10A076)
摘要Essential proteins are vital to the survival of a cell. There are various features related to the essentiality of proteins, such as biological and topological features. Many computational methods have been developed to identify essential proteins by using these features. However, it is still a big challenge to design an effective method that is able to select suitable features and integrate them to predict essential proteins. In this work, we first collect 26 features, and use SVM-RFE to select some of them to create a feature space for predicting essential proteins, and then remove the features that share the biological meaning with other features in the feature space according to their Pearson Correlation Coefficients(PCC). The experiments are carried out on S. cerevisiae data. Six features are determined as the best subset of features. To assess the prediction performance of our method, we further compare it with some machine learning methods, such as SVM, Naive Bayes, Bayes Network, and NBTree when inputting the different number of features. The results show that those methods using the 6 features outperform that using other features, which confirms the effectiveness of our feature selection method for essential protein prediction.
摘要In cloud computing Resource allocation is a very complex task.Handling the customer demand makes the challenges of on-demand resource allocation.Many challenges are faced by conventional methods for resource allocation in order tomeet the Quality of Service(QoS)requirements of users.For solving the about said problems a new method was implemented with the utility of machine learning framework of resource allocation by utilizing the cloud computing technique was taken in to an account in this research work.The accuracy in the machine learning algorithm can be improved by introducing Bat Algorithm with feature selection(BFS)in the proposed work,this further reduces the inappropriate features from the data.The similarities that were hidden can be demoralized by the Support Vector Machine(SVM)classifier which is also determine the subspace vector and then a new feature vector can be predicted by using SVM.For an unexpected circumstance SVM model can make a resource allocation decision.The efficiency of proposed SVM classifier of resource allocation can be highlighted by using a singlecell multiuser massive Multiple-Input Multiple Output(MIMO)system,with beam allocation problem as an example.The proposed resource allocation based on SVM performs efficiently than the existing conventional methods;this has been proven by analysing its results.