The advantages of genome selection(GS) in animal and plant breeding are self-evident.Traditional parametric models have disadvantage in better fit the increasingly large sequencing data and capture complex effects acc...The advantages of genome selection(GS) in animal and plant breeding are self-evident.Traditional parametric models have disadvantage in better fit the increasingly large sequencing data and capture complex effects accurately.Machine learning models have demonstrated remarkable potential in addressing these challenges.In this study,we introduced the concept of mixed kernel functions to explore the performance of support vector machine regression(SVR) in GS.Six single kernel functions(SVR_L,SVR_C,SVR_G,SVR_P,SVR_S,SVR_L) and four mixed kernel functions(SVR_GS,SVR_GP,SVR_LS,SVR_LP) were used to predict genome breeding values.The prediction accuracy,mean squared error(MSE) and mean absolute error(MAE) were used as evaluation indicators to compare with two traditional parametric models(GBLUP,BayesB) and two popular machine learning models(RF,KcRR).The results indicate that in most cases,the performance of the mixed kernel function model significantly outperforms that of GBLUP,BayesB and single kernel function.For instance,for T1 in the pig dataset,the predictive accuracy of SVR_GS is improved by 10% compared to GBLUP,and by approximately 4.4 and 18.6% compared to SVR_G and SVR_S respectively.For E1 in the wheat dataset,SVR_GS achieves 13.3% higher prediction accuracy than GBLUP.Among single kernel functions,the Laplacian and Gaussian kernel functions yield similar results,with the Gaussian kernel function performing better.The mixed kernel function notably reduces the MSE and MAE when compared to all single kernel functions.Furthermore,regarding runtime,SVR_GS and SVR_GP mixed kernel functions run approximately three times faster than GBLUP in the pig dataset,with only a slight increase in runtime compared to the single kernel function model.In summary,the mixed kernel function model of SVR demonstrates speed and accuracy competitiveness,and the model such as SVR_GS has important application potential for GS.展开更多
A proper non-landslide sample selection strategy can improve landslide susceptibility prediction(LSP)accuracy.However,there may be uncertainties regarding the compatibility between different selection strategies and m...A proper non-landslide sample selection strategy can improve landslide susceptibility prediction(LSP)accuracy.However,there may be uncertainties regarding the compatibility between different selection strategies and machine learning models,as well as in the extent of LSP performance enhancement after their coupling.To overcome these uncertainties,this study takes Wuning county of China as a case area,collecting 24 conditioning factors and 379 landslides data.Four non-landslide sample selection strategies,namely random selection,low-slope,buffer zone,and semi-supervised strategies,are then combined with landslide samples in a 1:1 ratio to serve as input variables for constructing LSP models using support vector machine(SVM),logistic regression(LR),random forest(RF)and extreme gradient boosting(XGBoost).Finally,the uncertainty of semi-supervised machine learning coupled models with a 1:2 ratio of landslide to non-landslide samples is analyzed and compared.The results show that:(1)The semi-supervised and low-slope strategies demonstrate higher prediction accuracy compared to the buffer zone and random selection strategies.Moreover,the RF coupled models are the most reliable,followed by the XGBoost,SVM,and LR coupled models;(2)Compared to a 1:1 ratio,a 1:2 ratio of landslide to non-landslide samples significantly improves prediction accuracy,suggesting that appropriately increasing the proportion of non-landslide samples helps to mitigate overfitting and enhance the identification of landslide samples;and(3)LSP is more sensitive to non-landslide sample selection strategies than to the choice of machine learning models.In conclusion,prioritizing reliable non-landslide samples is crucial for improving accuracy of LSP.展开更多
Affinity selection mass spectrometry(AS-MS)has emerged as a powerful label-free technique for identifying and characterizing ligand-target interactions.This review explores the diverse applications of AS-MS in drug di...Affinity selection mass spectrometry(AS-MS)has emerged as a powerful label-free technique for identifying and characterizing ligand-target interactions.This review explores the diverse applications of AS-MS in drug discovery,including its role in selective screening,binding site characterization,and quantitative affinity determination.We discuss the use of AS-MS for determining equilibrium dissociation constants(KD)and competitive binding parameters(affinity competition experiment 50%(ACE50)),highlighting its ability to rank ligand affinities efficiently.The review also examines AS-MS applications in fragment-based drug discovery(FBDD),screening for molecular glues,and investigating interactions with membrane proteins.Moreover,we address key technical challenges,including competitive binding effects,protein stability,and ligand dissociation kinetics,along with recent advancements in automation and artificial intelligence(AI)integration.Rather than providing a comprehensive literature review,this work aims to broaden the applicability of AS-MS assays and encourage researchers to explore its use in underutilized contexts.By providing rapid and high-sensitivity affinity measurements,AS-MS continues to expand its role in drug discovery and structural biology,complementing conventional biophysical techniques.展开更多
Advances in data acquisition and accumulation on a massive scale are fueling“the curse of dimensionality”which may deteriorate the generalization performance of machine learning models.Such a dilemma gives birth to ...Advances in data acquisition and accumulation on a massive scale are fueling“the curse of dimensionality”which may deteriorate the generalization performance of machine learning models.Such a dilemma gives birth to the technique of feature selection excelling in the presence of high-dimensional data.As a specific method based on rough set theory,rough feature selection(RFS)has been widely concerned and fruitfully applied.In this survey,we provide a comprehensive review of RFS algorithms that have proliferated in recent years.Firstly,we briefly introduce some typical rough set models especially neighborhood rough set and fuzzy rough set,as well as representative rough feature evaluation criteria.We then systematically discuss several emerging topics of RFS including accelerated,ensemble,incremental,label ambiguous,weakly-supervised,and multi-granularity RFS.Additionally,we illuminate the regular performance validation scheme of RFS and conduct a number of experiments to present benchmarking results of state-of-the-art RFS algorithms.Finally,we summarize the pros and cons of existing research efforts and outline the open challenges and opportunities of class imbalance,multi-modal scenario,causality inference,and highlevel representation for RFS.By providing in-depth knowledge of RFS,we anticipate this survey will:1)serve as a guidebook for newcomers intending to delve into RFS and a stepping-stone for researchers and practitioners to solve domain-specific problems;2)gain insights into the state-of-the-art published findings,triggering a series of breakthroughs in RFS;3)underscore some challenges ahead of RFS,directing future efforts toward punctuating advances beyond questions currently pursued.展开更多
As the core of cathode materials,sensitive metals play important roles in the optimization of acetate production from carbon dioxide(CO2)in microbial electrochemical system(MES).In this work,iron(Fe),copper(Cu),and...As the core of cathode materials,sensitive metals play important roles in the optimization of acetate production from carbon dioxide(CO2)in microbial electrochemical system(MES).In this work,iron(Fe),copper(Cu),and nickel(Ni)as sensitive metal cathode materials were evaluated for CO2 conversion in MES.The MES with Feelectrode as a promising electrode material demonstrated a superior CO2 reduction performance with a maximum acetate accumulation of 417.9±39.2 mg/L,which was 1.5 and 1.7 folds higher than that in the Ni-electrode and Cu-electrode groups,respectively.Furthermore,an outstanding electron recovery efficiency of 67.7%was shown in the Fe-electrode group.The electron transfer between electrode-suspended sludge was systematically cross-evaluated by the electrochemical behavior and extracellular polymeric substances.The Fe-electrode group had the highest electron transfer rate with 0.194 s-1(kapp),which was 17.6 and 21.5 times higher than that of the Cu-and Ni-electrode groups,respectively.Fe-electrode was beneficial for reducing electrochemical impedance between the electrode and suspended sludge.Additionally,redox substances in extracellular polymeric substances of the Fe-electrode group were increased,implying more favorable electron transport dynamics.Simultaneously,enrichments of functional bacteria Acetoanerobium and increased key enzymes involved in the carbonyl pathway of the Fe-electrode group were observed,which also promoted CO2 conversion in MES.This study provides a perspective on evaluating the promising sensitive metal electrode material for the process of CO2 valorization in MES and offers a reference for the subsequent electrode modification.展开更多
Plant height(PH)and aboveground biomass(AGB)are critical agronomic traits that determine the yield potential of maize(Zea mays L.).However,the application of genomic selection(GS)and genome-wide association studies(GW...Plant height(PH)and aboveground biomass(AGB)are critical agronomic traits that determine the yield potential of maize(Zea mays L.).However,the application of genomic selection(GS)and genome-wide association studies(GWAS)in maize breeding is often hindered by the limitations of phenotypic data collection,which is typically characterized by low throughput and inadequate accuracy.To address this challenge,we employed an unmanned aerial vehicle(UAV)equipped with LiDAR and RGB cameras for high-throughput assessment of PH and AGB in a panel of 817 maize hybrids derived from 364 inbred lines over two growing seasons.Our results demonstrated that the integration of UAV-derived LiDAR point clouds with crop surface models(CSMs)enabled robust estimation of PH across multiple years(R2>0.90).Furthermore,a three-dimensional AGB estimation model was developed using UAVderived PH and canopy coverage(CC),achieving high estimation accuracy(R2>0.83).Subsequently,the UAV-derived PH and AGB were utilized for GS and GWAS analyses.Replicated 10-fold crossvalidation showed that the mean predictability was 0.504 for PH and 0.402 for AGB across eight commonly used GS models.Moreover,of the 66,066 potential crosses derived from the 364 inbred lines,the top 200 crosses selected for AGB showed up to twice the AGB of the bottom 200 crosses.Field validation demonstrated that the mean ear weight(EW)in the AGB top group was 39.0% higher than that in the bottom group.A total of 16 and 11 significant SNPs were identified by at least two GWAS methods for PH and AGB,respectively.Based on these SNPs,81 candidate genes were functionally annotated,six of which were simultaneously associated with both traits.The candidate gene association analysis suggested that variations in the promoter region of ZmFLA9 may affect both traits.Overall,our study highlights the potential of UAV-based high-throughput phenotyping to accelerate maize genomic breeding by enabling rapid,precise,and large-scale trait assessment.展开更多
In Wireless Sensor Networks(WSNs),survivability is a crucial issue that is greatly impacted by energy efficiency.Solutions that satisfy application objectives while extending network life are needed to address severe ...In Wireless Sensor Networks(WSNs),survivability is a crucial issue that is greatly impacted by energy efficiency.Solutions that satisfy application objectives while extending network life are needed to address severe energy constraints inWSNs.This paper presents an Adaptive Enhanced GreyWolf Optimizer(AEGWO)for energy-efficient cluster head(CH)selection that mitigates the exploration–exploitation imbalance,preserves population diversity,and avoids premature convergence inherent in baseline GWO.The AEGWO combines adaptive control of the parameter of the search pressure to accelerate convergence without stagnation,a hybrid velocity-momentum update based on the dynamics of PSO,and an intelligent mutation operator to maintain the diversity of the population.The search is guided by a multi-objective fitness,which aims at maximizing the residual energy,equal distribution of CH,minimizing the intra-cluster distance,desirable proximity to sinks,and enhancing the coverage.Simulations on 100 nodes homogeneousWSN Tested the proposed AEGWO under the same conditions with LEACH,GWO,IGWO,PSO,WOA,and GA,AEGWO significantly increases stability and lifetime compared to LEACHand other tested algorithms;it has the best first,half,and last node dead,and higher residual energy and smaller communication overhead.The findings prove that AEGWO provides sustainable energy management and better lifetime extension,which makes it a robust,flexible clustering protocol of large-scaleWSNs.展开更多
The selection of a suitable navigation area is pivotal in aircraft scene matching guidance technology.This study addresses the challenge of identifying suitable reference image ranges for precise scene matching,which ...The selection of a suitable navigation area is pivotal in aircraft scene matching guidance technology.This study addresses the challenge of identifying suitable reference image ranges for precise scene matching,which is crucial for enhancing aircraft positioning accuracy.Traditional methods for image matchability analysis are often limited by their reliance on manual feature parameter design and threshold-based filtering,resulting in suboptimal accuracy and efficiency.This paper proposes a novel network architecture for selecting suitable navigation areas using image Matching Level Segmentation(MLSNet).The approach involves two key innovations:a method for generating segmentation labels that quantify matchability levels and an end-to-end network architecture for rapid and precise prediction of reference image matchability segmentation maps.The network includes two core modules:the saliency analysis module uses multi-layer convolutional networks to accurately detect image saliency features across various levels and scales;the multidimensional attention module utilizes attention mechanisms to focus on feature channels and spatial neighborhood scenes to assess the image’s matchability.Our method was rigorously tested on an extensive collection of remote sensing images,where it was benchmarked against a range of both traditional and cutting-edge deep learning methods.The findings indicate that MLSNet is significantly superior to traditional methods in accuracy and efficiency of matchability analysis,and is also relatively ahead of state-of-the-art deep learning models.展开更多
Existing feature selection methods for intrusion detection systems in the Industrial Internet of Things often suffer from local optimality and high computational complexity.These challenges hinder traditional IDS from...Existing feature selection methods for intrusion detection systems in the Industrial Internet of Things often suffer from local optimality and high computational complexity.These challenges hinder traditional IDS from effectively extracting features while maintaining detection accuracy.This paper proposes an industrial Internet ofThings intrusion detection feature selection algorithm based on an improved whale optimization algorithm(GSLDWOA).The aim is to address the problems that feature selection algorithms under high-dimensional data are prone to,such as local optimality,long detection time,and reduced accuracy.First,the initial population’s diversity is increased using the Gaussian Mutation mechanism.Then,Non-linear Shrinking Factor balances global exploration and local development,avoiding premature convergence.Lastly,Variable-step Levy Flight operator and Dynamic Differential Evolution strategy are introduced to improve the algorithm’s search efficiency and convergence accuracy in highdimensional feature space.Experiments on the NSL-KDD and WUSTL-IIoT-2021 datasets demonstrate that the feature subset selected by GSLDWOA significantly improves detection performance.Compared to the traditional WOA algorithm,the detection rate and F1-score increased by 3.68%and 4.12%.On the WUSTL-IIoT-2021 dataset,accuracy,recall,and F1-score all exceed 99.9%.展开更多
Feature selection grounded in neighborhood rough sets has attracted sustained research attention owing to its principled treatment of classification uncertainty.However,existing forward greedy algorithms typically eva...Feature selection grounded in neighborhood rough sets has attracted sustained research attention owing to its principled treatment of classification uncertainty.However,existing forward greedy algorithms typically evaluate uncertainty over the entire object universe at each iteration,resulting in prohibitive computational complexity on large-scale datasets.To address this inefficiency,we introduce a new uncertainty index built upon Boundary Object Sets(BOS).BOS are defined as objects whose neighborhood granules intersect with multiple decision classes,thereby capturing intrinsic classification ambiguity.The proposed measure quantifies the proportion of these boundary objects relative to the total universe size.Grounded in this measure,we develop the Boundary Object Set Feature Selection(BOSFS)algorithm.BOSFS employs a double-contraction strategy that simultaneously reduces the candidate attribute pool and eliminates consistent objects from the active working set.Consequently,the algorithm restricts computationally expensive distance calculations to the monotonically shrinking BOS.Experiments on ten benchmark datasets,evaluated against six competing algorithms under three classifiers,confirm that BOSFS achieves the highest classification accuracy in 15 of 30 test cases while consuming only 25%of the runtime of the second-fastest competitor.展开更多
Feature selection serves as a critical preprocessing step inmachine learning,focusing on identifying and preserving the most relevant features to improve the efficiency and performance of classification algorithms.Par...Feature selection serves as a critical preprocessing step inmachine learning,focusing on identifying and preserving the most relevant features to improve the efficiency and performance of classification algorithms.Particle Swarm Optimization has demonstrated significant potential in addressing feature selection challenges.However,there are inherent limitations in Particle Swarm Optimization,such as the delicate balance between exploration and exploitation,susceptibility to local optima,and suboptimal convergence rates,hinder its performance.To tackle these issues,this study introduces a novel Leveraged Opposition-Based Learning method within Fitness Landscape Particle Swarm Optimization,tailored for wrapper-based feature selection.The proposed approach integrates:(1)a fitness-landscape adaptive strategy to dynamically balance exploration and exploitation,(2)the lever principle within Opposition-Based Learning to improve search efficiency,and(3)a Local Selection and Re-optimization mechanism combined with random perturbation to expedite convergence and enhance the quality of the optimal feature subset.The effectiveness of is rigorously evaluated on 24 benchmark datasets and compared against 13 advancedmetaheuristic algorithms.Experimental results demonstrate that the proposed method outperforms the compared algorithms in classification accuracy on over half of the datasets,whilst also significantly reducing the number of selected features.These findings demonstrate its effectiveness and robustness in feature selection tasks.展开更多
This paper addresses an optimal sensor selection problem under the framework of linear quadratic regulation.Unlike prior work on optimal sensor scheduling,we assume that the sensor noise covariance matrices are compar...This paper addresses an optimal sensor selection problem under the framework of linear quadratic regulation.Unlike prior work on optimal sensor scheduling,we assume that the sensor noise covariance matrices are comparable but unknown.Then,the optimal sensor selection problem is formulated as finding an optimal policy of selecting a sensor from a set of sensors to minimize the expected quadratic performance of a linear system given the number of trials.An action value method from reinforcement learning is adopted for estimating the values of selections and making selection decisions based on the estimates.Several ways of balancing exploration and exploitation are presented and compared for efficacy.Numerical simulations are conducted to demonstrate the effectiveness of the proposed algorithms.展开更多
With the advancement of brain–computer interfaces(BCI),motor imagery(MI)electroencephalogram(EEG)decoding can greatly benefit from spatial filtering features derived from common spatial patterns(CSP).However,CSP-base...With the advancement of brain–computer interfaces(BCI),motor imagery(MI)electroencephalogram(EEG)decoding can greatly benefit from spatial filtering features derived from common spatial patterns(CSP).However,CSP-based features often exhibit high redundancy and intersubject variability.These limitations make the feature selection methods based on sparse learning difficult to effectively balance the heterogeneous contributions of different temporal and spatial components.Moreover,these models tend to prioritise features with larger coefficients,potentially overlooking intrinsic feature importance and compromising the quality of the selected feature subset.To address these issues,we propose an Adaptive Sparse Group Lasso(ASGL)method for structured feature selection,designed to enhance discriminative CSP features whilst suppressing irrelevant components.The proposed method partitions EEG signals into consecutive segments using a sliding window,treating each as a separate feature group.Benefiting from this,the importance of features at both the group level and the within-group level can be effectively quantified through mutual information and copula mutual information,thereby assigning adaptive weights for selective penalisation within the model.This weight construction strategy preserves important features from relevant time intervals and frequency bands.The resulting optimization problem is solved efficiently via the alternating direction method of multipliers(ADMM).Evaluations on simulated and real-world datasets demonstrate that the proposed ASGL outperforms existing methods.展开更多
In the quest to enhance energy efficiency and reduce environmental impact in the transportation sector,the recovery of waste heat from diesel engines has become a critical area of focus.This study provided an exhausti...In the quest to enhance energy efficiency and reduce environmental impact in the transportation sector,the recovery of waste heat from diesel engines has become a critical area of focus.This study provided an exhaustive thermodynamic analysis optimizing Organic Rankine Cycle(ORC)systems forwaste heat recovery fromdiesel engines.Thestudy assessed the performance of five candidateworking fluids—R11,R123,R113,R245fa,and R141b—under a range of operating conditions,specifically varying overheat temperatures and evaporation pressures.The results indicated that the choice of working fluid substantially influences the system’s exergetic efficiency,net output power,and thermal efficiency.R245fa showed an outstanding net output power of 30.39 kW at high overheat conditions,outperforming R11,which is significant for high-temperature waste heat recovery.At lower temperatures,R11 and R113 demonstrated higher exergetic efficiencies,with R11 reaching a peak exergetic efficiency of 7.4%at an evaporation pressure of 10 bar and an overheat of 10℃.The study also revealed that controlling the overheat and optimizing the evaporation pressure are crucial for enhancing the net output power of the ORC system.Specifically,at an evaporation pressure of 30 bar and an overheat of 0℃,R113 exhibited the lowest exergetic destruction of 544.5 kJ/kg,making it a suitable choice for minimizing irreversible losses.These findings are instrumental for understanding the performance of ORC systems in waste heat recovery applications and offer valuable insights for the design and operation of more efficient and environmentally friendly diesel engine systems.展开更多
Multi-label feature selection(MFS)is a crucial dimensionality reduction technique aimed at identifying informative features associated with multiple labels.However,traditional centralized methods face significant chal...Multi-label feature selection(MFS)is a crucial dimensionality reduction technique aimed at identifying informative features associated with multiple labels.However,traditional centralized methods face significant challenges in privacy-sensitive and distributed settings,often neglecting label dependencies and suffering from low computational efficiency.To address these issues,we introduce a novel framework,Fed-MFSDHBCPSO—federated MFS via dual-layer hybrid breeding cooperative particle swarm optimization algorithm with manifold and sparsity regularization(DHBCPSO-MSR).Leveraging the federated learning paradigm,Fed-MFSDHBCPSO allows clients to perform local feature selection(FS)using DHBCPSO-MSR.Locally selected feature subsets are encrypted with differential privacy(DP)and transmitted to a central server,where they are securely aggregated and refined through secure multi-party computation(SMPC)until global convergence is achieved.Within each client,DHBCPSO-MSR employs a dual-layer FS strategy.The inner layer constructs sample and label similarity graphs,generates Laplacian matrices to capture the manifold structure between samples and labels,and applies L2,1-norm regularization to sparsify the feature subset,yielding an optimized feature weight matrix.The outer layer uses a hybrid breeding cooperative particle swarm optimization algorithm to further refine the feature weight matrix and identify the optimal feature subset.The updated weight matrix is then fed back to the inner layer for further optimization.Comprehensive experiments on multiple real-world multi-label datasets demonstrate that Fed-MFSDHBCPSO consistently outperforms both centralized and federated baseline methods across several key evaluation metrics.展开更多
With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the mod...With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the model in identifying non-standard semantics due to its semantic complexity and high-risk features.However,existing fine-tuning methods rely heavily on static data selection strategies,making it difficult to adapt to the dynamic evolution of model capabilities,resulting in low training efficiency.This article proposes ADS(Adaptive Dataset Selection),an adaptive framework for selecting data in anomaly text detection.ADS performs model-aware data selection prior to fine-tuning,adapting the initial state of pre-trained language models by selecting samples that are most informative for the target anomaly detection task.Empirical results on mainstream large language model architectures show that ADS significantly compresses data size while still outperforming existing static strategies and mainstream compression methods.When using only 1000 fine-tuning samples,ADS achieves a 92%F1 score,with an accuracy improvement of over 22%compared to the baseline,demonstrating excellent performance.This study proposes an efficient data selection mechanism from the perspective of model capability and dynamic adaptation of data,providing theoretical support and a practical path for fine-tuning large models in low-resource scenarios.展开更多
Sleep-dependent memory consolidation relies on coordinated interactions between hippocampal replay,cortical oscillations,and brain state transitions that unfold across the sleep onset period and early non-rapid eye mo...Sleep-dependent memory consolidation relies on coordinated interactions between hippocampal replay,cortical oscillations,and brain state transitions that unfold across the sleep onset period and early non-rapid eye movement(NREM)sleep[1].While substantial progress has been made in identifying the neuronal substrates of replay and its role in systems consolidation[2,3],less attention has been paid to how global physiological states shape the conditions under which distinct replay patterns are expressed and effectively coupled to downstream consolidation processes.In particular,autonomic regulation during the wake-to-sleep transition has emerged as a key modulator of NREM sleep depth,oscillatory synchronization,and hippocampal cortical communication[4],yet its potential influence on the structure of hippocampal replay remains largely unexplored.展开更多
High-dimensional data causes difficulties in machine learning due to high time consumption and large memory requirements.In particular,in amulti-label environment,higher complexity is required asmuch as the number of ...High-dimensional data causes difficulties in machine learning due to high time consumption and large memory requirements.In particular,in amulti-label environment,higher complexity is required asmuch as the number of labels.Moreover,an optimization problem that fully considers all dependencies between features and labels is difficult to solve.In this study,we propose a novel regression-basedmulti-label feature selectionmethod that integrates mutual information to better exploit the underlying data structure.By incorporating mutual information into the regression formulation,the model captures not only linear relationships but also complex non-linear dependencies.The proposed objective function simultaneously considers three types of relationships:(1)feature redundancy,(2)featurelabel relevance,and(3)inter-label dependency.These three quantities are computed usingmutual information,allowing the proposed formulation to capture nonlinear dependencies among variables.These three types of relationships are key factors in multi-label feature selection,and our method expresses them within a unified formulation,enabling efficient optimization while simultaneously accounting for all of them.To efficiently solve the proposed optimization problem under non-negativity constraints,we develop a gradient-based optimization algorithm with fast convergence.Theexperimental results on sevenmulti-label datasets show that the proposed method outperforms existingmulti-label feature selection techniques.展开更多
Emerging and powerful genome editing tools,particularly CRISPR/Cas9,are facilitating functional genomics research and accelerating crop improvement(Jiang et al.2021;Cao et al.2023;Chen C et al.2023;Liu et al.2023a).Ho...Emerging and powerful genome editing tools,particularly CRISPR/Cas9,are facilitating functional genomics research and accelerating crop improvement(Jiang et al.2021;Cao et al.2023;Chen C et al.2023;Liu et al.2023a).However,the detection and screening of transgenic lines remain major bottlenecks,being time-consuming,labor-intensive,and inefficient during transformation and subsequent mutation identification.A simple and efficient visual marker system plays a critical role in addressing these challenges.Recent studies demonstrated that the GmW1 and RUBY reporter systems were used to obtain visual transgenic soybean(Glycine max) plants(Chen L et al.2023;Chen et al.2024).展开更多
We present Hi4GS,a hybrid feature selection(HFS)algorithm for selecting SNP subsets from highdimensional genotypes to improve the prediction of genomic estimated breeding value(GEBV)under genomic selection(GS).Hi4GS c...We present Hi4GS,a hybrid feature selection(HFS)algorithm for selecting SNP subsets from highdimensional genotypes to improve the prediction of genomic estimated breeding value(GEBV)under genomic selection(GS).Hi4GS combines feature importance weighting with quantity determining to construct a fused feature set from which it extracts an optimal feature subset for subsequent GS.In a study of wheat using four datasets covering 11 yield traits via large-scale GS models,the SNPs selected by Hi4GS increased the average predictive accuracy by over 82%than using all SNPs.Hi4GS was used to identify SNPs potentially affecting wheat yield,and SHAP-based interpretability was applied to explain the contributions of these SNPs and their potential interactions.Hi4GS can be used for assisting in improving the prediction accuracy of GS,wheat and other plants'yield-associated SNPs identification,and target information for breeding chip development.The free R package Hi4GS is available at http://gffzz188fe103f8f1460as0ucunx6b90o66kvw.ffgz.tsg.suse.edu.cn/shgs19/Hi4GS.展开更多
基金supported by the China Agriculture Research System of MOF and MARAthe National Natural Science Foundation of China (31872337 and 31501919)the Agricultural Science and Technology Innovation Project,China (ASTIP-IAS02)。
摘要The advantages of genome selection(GS) in animal and plant breeding are self-evident.Traditional parametric models have disadvantage in better fit the increasingly large sequencing data and capture complex effects accurately.Machine learning models have demonstrated remarkable potential in addressing these challenges.In this study,we introduced the concept of mixed kernel functions to explore the performance of support vector machine regression(SVR) in GS.Six single kernel functions(SVR_L,SVR_C,SVR_G,SVR_P,SVR_S,SVR_L) and four mixed kernel functions(SVR_GS,SVR_GP,SVR_LS,SVR_LP) were used to predict genome breeding values.The prediction accuracy,mean squared error(MSE) and mean absolute error(MAE) were used as evaluation indicators to compare with two traditional parametric models(GBLUP,BayesB) and two popular machine learning models(RF,KcRR).The results indicate that in most cases,the performance of the mixed kernel function model significantly outperforms that of GBLUP,BayesB and single kernel function.For instance,for T1 in the pig dataset,the predictive accuracy of SVR_GS is improved by 10% compared to GBLUP,and by approximately 4.4 and 18.6% compared to SVR_G and SVR_S respectively.For E1 in the wheat dataset,SVR_GS achieves 13.3% higher prediction accuracy than GBLUP.Among single kernel functions,the Laplacian and Gaussian kernel functions yield similar results,with the Gaussian kernel function performing better.The mixed kernel function notably reduces the MSE and MAE when compared to all single kernel functions.Furthermore,regarding runtime,SVR_GS and SVR_GP mixed kernel functions run approximately three times faster than GBLUP in the pig dataset,with only a slight increase in runtime compared to the single kernel function model.In summary,the mixed kernel function model of SVR demonstrates speed and accuracy competitiveness,and the model such as SVR_GS has important application potential for GS.
基金financially supported by the National Natural Science Foundation of China(Grant Nos.42202278,42407241)Natural Science Foundation of Jiangxi Province(Grant No.20242BAB20238).
摘要A proper non-landslide sample selection strategy can improve landslide susceptibility prediction(LSP)accuracy.However,there may be uncertainties regarding the compatibility between different selection strategies and machine learning models,as well as in the extent of LSP performance enhancement after their coupling.To overcome these uncertainties,this study takes Wuning county of China as a case area,collecting 24 conditioning factors and 379 landslides data.Four non-landslide sample selection strategies,namely random selection,low-slope,buffer zone,and semi-supervised strategies,are then combined with landslide samples in a 1:1 ratio to serve as input variables for constructing LSP models using support vector machine(SVM),logistic regression(LR),random forest(RF)and extreme gradient boosting(XGBoost).Finally,the uncertainty of semi-supervised machine learning coupled models with a 1:2 ratio of landslide to non-landslide samples is analyzed and compared.The results show that:(1)The semi-supervised and low-slope strategies demonstrate higher prediction accuracy compared to the buffer zone and random selection strategies.Moreover,the RF coupled models are the most reliable,followed by the XGBoost,SVM,and LR coupled models;(2)Compared to a 1:1 ratio,a 1:2 ratio of landslide to non-landslide samples significantly improves prediction accuracy,suggesting that appropriately increasing the proportion of non-landslide samples helps to mitigate overfitting and enhance the identification of landslide samples;and(3)LSP is more sensitive to non-landslide sample selection strategies than to the choice of machine learning models.In conclusion,prioritizing reliable non-landslide samples is crucial for improving accuracy of LSP.
基金Support of the State of Rio de Janeiro(FAPERJ),Brazil(Grant Nos.:E-26/210.017/2024,E-200.172/2023,E-26/200.165/2024,E-26/200.164/2024,and E-26/210.547/2025)the Coordination for the Improvement of Higher Education Personnel(CAPES),Brazil(Finance Code 001)National Council for Scientific and Technological Development(CNPq),Brazil(Grant Nos.:307108/2021-0 and 302464/2022-0)for their support.
摘要Affinity selection mass spectrometry(AS-MS)has emerged as a powerful label-free technique for identifying and characterizing ligand-target interactions.This review explores the diverse applications of AS-MS in drug discovery,including its role in selective screening,binding site characterization,and quantitative affinity determination.We discuss the use of AS-MS for determining equilibrium dissociation constants(KD)and competitive binding parameters(affinity competition experiment 50%(ACE50)),highlighting its ability to rank ligand affinities efficiently.The review also examines AS-MS applications in fragment-based drug discovery(FBDD),screening for molecular glues,and investigating interactions with membrane proteins.Moreover,we address key technical challenges,including competitive binding effects,protein stability,and ligand dissociation kinetics,along with recent advancements in automation and artificial intelligence(AI)integration.Rather than providing a comprehensive literature review,this work aims to broaden the applicability of AS-MS assays and encourage researchers to explore its use in underutilized contexts.By providing rapid and high-sensitivity affinity measurements,AS-MS continues to expand its role in drug discovery and structural biology,complementing conventional biophysical techniques.
基金supported by the National Natural Science Foundation of China(62506145,62576178,U2433216)Natural Science Foundation of Jiangsu Higher Education Institutions(25KJB520008)。
摘要Advances in data acquisition and accumulation on a massive scale are fueling“the curse of dimensionality”which may deteriorate the generalization performance of machine learning models.Such a dilemma gives birth to the technique of feature selection excelling in the presence of high-dimensional data.As a specific method based on rough set theory,rough feature selection(RFS)has been widely concerned and fruitfully applied.In this survey,we provide a comprehensive review of RFS algorithms that have proliferated in recent years.Firstly,we briefly introduce some typical rough set models especially neighborhood rough set and fuzzy rough set,as well as representative rough feature evaluation criteria.We then systematically discuss several emerging topics of RFS including accelerated,ensemble,incremental,label ambiguous,weakly-supervised,and multi-granularity RFS.Additionally,we illuminate the regular performance validation scheme of RFS and conduct a number of experiments to present benchmarking results of state-of-the-art RFS algorithms.Finally,we summarize the pros and cons of existing research efforts and outline the open challenges and opportunities of class imbalance,multi-modal scenario,causality inference,and highlevel representation for RFS.By providing in-depth knowledge of RFS,we anticipate this survey will:1)serve as a guidebook for newcomers intending to delve into RFS and a stepping-stone for researchers and practitioners to solve domain-specific problems;2)gain insights into the state-of-the-art published findings,triggering a series of breakthroughs in RFS;3)underscore some challenges ahead of RFS,directing future efforts toward punctuating advances beyond questions currently pursued.
基金supported by the Science and Technology Commission of Shanghai Municipality Foundation(No.22230710500)the Interdisciplinary joint research project of Tongji University(No.2023-3-YB-07).
摘要As the core of cathode materials,sensitive metals play important roles in the optimization of acetate production from carbon dioxide(CO2)in microbial electrochemical system(MES).In this work,iron(Fe),copper(Cu),and nickel(Ni)as sensitive metal cathode materials were evaluated for CO2 conversion in MES.The MES with Feelectrode as a promising electrode material demonstrated a superior CO2 reduction performance with a maximum acetate accumulation of 417.9±39.2 mg/L,which was 1.5 and 1.7 folds higher than that in the Ni-electrode and Cu-electrode groups,respectively.Furthermore,an outstanding electron recovery efficiency of 67.7%was shown in the Fe-electrode group.The electron transfer between electrode-suspended sludge was systematically cross-evaluated by the electrochemical behavior and extracellular polymeric substances.The Fe-electrode group had the highest electron transfer rate with 0.194 s-1(kapp),which was 17.6 and 21.5 times higher than that of the Cu-and Ni-electrode groups,respectively.Fe-electrode was beneficial for reducing electrochemical impedance between the electrode and suspended sludge.Additionally,redox substances in extracellular polymeric substances of the Fe-electrode group were increased,implying more favorable electron transport dynamics.Simultaneously,enrichments of functional bacteria Acetoanerobium and increased key enzymes involved in the carbonyl pathway of the Fe-electrode group were observed,which also promoted CO2 conversion in MES.This study provides a perspective on evaluating the promising sensitive metal electrode material for the process of CO2 valorization in MES and offers a reference for the subsequent electrode modification.
基金supported by grants from the National Key Research and Development Program of China(2023YFD1202200)the Key Research and Development Program of Jiangsu Province(BE2022343)+3 种基金the National Natural Science Foundation of China(32561143291,32261143462)the State Key Laboratory of Crop Gene Resources and Breeding(CGRB-2026-04)the Seed Industry Revitalization Project of Jiangsu Province(JBGS[2021]009)the Priority Academic Program Development of Jiangsu Higher Education Institutions(PAPD)。
摘要Plant height(PH)and aboveground biomass(AGB)are critical agronomic traits that determine the yield potential of maize(Zea mays L.).However,the application of genomic selection(GS)and genome-wide association studies(GWAS)in maize breeding is often hindered by the limitations of phenotypic data collection,which is typically characterized by low throughput and inadequate accuracy.To address this challenge,we employed an unmanned aerial vehicle(UAV)equipped with LiDAR and RGB cameras for high-throughput assessment of PH and AGB in a panel of 817 maize hybrids derived from 364 inbred lines over two growing seasons.Our results demonstrated that the integration of UAV-derived LiDAR point clouds with crop surface models(CSMs)enabled robust estimation of PH across multiple years(R2>0.90).Furthermore,a three-dimensional AGB estimation model was developed using UAVderived PH and canopy coverage(CC),achieving high estimation accuracy(R2>0.83).Subsequently,the UAV-derived PH and AGB were utilized for GS and GWAS analyses.Replicated 10-fold crossvalidation showed that the mean predictability was 0.504 for PH and 0.402 for AGB across eight commonly used GS models.Moreover,of the 66,066 potential crosses derived from the 364 inbred lines,the top 200 crosses selected for AGB showed up to twice the AGB of the bottom 200 crosses.Field validation demonstrated that the mean ear weight(EW)in the AGB top group was 39.0% higher than that in the bottom group.A total of 16 and 11 significant SNPs were identified by at least two GWAS methods for PH and AGB,respectively.Based on these SNPs,81 candidate genes were functionally annotated,six of which were simultaneously associated with both traits.The candidate gene association analysis suggested that variations in the promoter region of ZmFLA9 may affect both traits.Overall,our study highlights the potential of UAV-based high-throughput phenotyping to accelerate maize genomic breeding by enabling rapid,precise,and large-scale trait assessment.
基金The Open Access publication fee for this article was fully covered by Abu Dhabi University.
摘要In Wireless Sensor Networks(WSNs),survivability is a crucial issue that is greatly impacted by energy efficiency.Solutions that satisfy application objectives while extending network life are needed to address severe energy constraints inWSNs.This paper presents an Adaptive Enhanced GreyWolf Optimizer(AEGWO)for energy-efficient cluster head(CH)selection that mitigates the exploration–exploitation imbalance,preserves population diversity,and avoids premature convergence inherent in baseline GWO.The AEGWO combines adaptive control of the parameter of the search pressure to accelerate convergence without stagnation,a hybrid velocity-momentum update based on the dynamics of PSO,and an intelligent mutation operator to maintain the diversity of the population.The search is guided by a multi-objective fitness,which aims at maximizing the residual energy,equal distribution of CH,minimizing the intra-cluster distance,desirable proximity to sinks,and enhancing the coverage.Simulations on 100 nodes homogeneousWSN Tested the proposed AEGWO under the same conditions with LEACH,GWO,IGWO,PSO,WOA,and GA,AEGWO significantly increases stability and lifetime compared to LEACHand other tested algorithms;it has the best first,half,and last node dead,and higher residual energy and smaller communication overhead.The findings prove that AEGWO provides sustainable energy management and better lifetime extension,which makes it a robust,flexible clustering protocol of large-scaleWSNs.
基金supported in part by the National Natural Science Foundation of China(No.42271446)in part by the Tianjin Key Laboratory of Rail Transit Navigation Positioning and Spatio-Temporary Big Data Technology,China(No.TKL2024B13)in part by the Science and Technology Program of Tianjin,China(No.24YFYSHZ00080)。
摘要The selection of a suitable navigation area is pivotal in aircraft scene matching guidance technology.This study addresses the challenge of identifying suitable reference image ranges for precise scene matching,which is crucial for enhancing aircraft positioning accuracy.Traditional methods for image matchability analysis are often limited by their reliance on manual feature parameter design and threshold-based filtering,resulting in suboptimal accuracy and efficiency.This paper proposes a novel network architecture for selecting suitable navigation areas using image Matching Level Segmentation(MLSNet).The approach involves two key innovations:a method for generating segmentation labels that quantify matchability levels and an end-to-end network architecture for rapid and precise prediction of reference image matchability segmentation maps.The network includes two core modules:the saliency analysis module uses multi-layer convolutional networks to accurately detect image saliency features across various levels and scales;the multidimensional attention module utilizes attention mechanisms to focus on feature channels and spatial neighborhood scenes to assess the image’s matchability.Our method was rigorously tested on an extensive collection of remote sensing images,where it was benchmarked against a range of both traditional and cutting-edge deep learning methods.The findings indicate that MLSNet is significantly superior to traditional methods in accuracy and efficiency of matchability analysis,and is also relatively ahead of state-of-the-art deep learning models.
基金supported by the Major Science and Technology Programs in Henan Province(No.241100210100)Henan Provincial Science and Technology Research Project(No.252102211085,No.252102211105)+3 种基金Endogenous Security Cloud Network Convergence R&D Center(No.602431011PQ1)The Special Project for Research and Development in Key Areas of Guangdong Province(No.2021ZDZX1098)The Stabilization Support Program of Science,Technology and Innovation Commission of Shenzhen Municipality(No.20231128083944001)The Key scientific research projects of Henan higher education institutions(No.24A520042).
摘要Existing feature selection methods for intrusion detection systems in the Industrial Internet of Things often suffer from local optimality and high computational complexity.These challenges hinder traditional IDS from effectively extracting features while maintaining detection accuracy.This paper proposes an industrial Internet ofThings intrusion detection feature selection algorithm based on an improved whale optimization algorithm(GSLDWOA).The aim is to address the problems that feature selection algorithms under high-dimensional data are prone to,such as local optimality,long detection time,and reduced accuracy.First,the initial population’s diversity is increased using the Gaussian Mutation mechanism.Then,Non-linear Shrinking Factor balances global exploration and local development,avoiding premature convergence.Lastly,Variable-step Levy Flight operator and Dynamic Differential Evolution strategy are introduced to improve the algorithm’s search efficiency and convergence accuracy in highdimensional feature space.Experiments on the NSL-KDD and WUSTL-IIoT-2021 datasets demonstrate that the feature subset selected by GSLDWOA significantly improves detection performance.Compared to the traditional WOA algorithm,the detection rate and F1-score increased by 3.68%and 4.12%.On the WUSTL-IIoT-2021 dataset,accuracy,recall,and F1-score all exceed 99.9%.
基金supported by the Anhui Provincial Department of Education University Research Project(2024AH051375,2024AH051368)industry-sponsored research project from Nanjing Wenstone Information Technology Co.,Ltd.(2025HX0249).
摘要Feature selection grounded in neighborhood rough sets has attracted sustained research attention owing to its principled treatment of classification uncertainty.However,existing forward greedy algorithms typically evaluate uncertainty over the entire object universe at each iteration,resulting in prohibitive computational complexity on large-scale datasets.To address this inefficiency,we introduce a new uncertainty index built upon Boundary Object Sets(BOS).BOS are defined as objects whose neighborhood granules intersect with multiple decision classes,thereby capturing intrinsic classification ambiguity.The proposed measure quantifies the proportion of these boundary objects relative to the total universe size.Grounded in this measure,we develop the Boundary Object Set Feature Selection(BOSFS)algorithm.BOSFS employs a double-contraction strategy that simultaneously reduces the candidate attribute pool and eliminates consistent objects from the active working set.Consequently,the algorithm restricts computationally expensive distance calculations to the monotonically shrinking BOS.Experiments on ten benchmark datasets,evaluated against six competing algorithms under three classifiers,confirm that BOSFS achieves the highest classification accuracy in 15 of 30 test cases while consuming only 25%of the runtime of the second-fastest competitor.
基金supported by National Natural Science Foundation of China(62106092)Natural Science Foundation of Fujian Province(2024J01822,2024J01820,2022J01916)Natural Science Foundation of Zhangzhou City(ZZ2024J28).
摘要Feature selection serves as a critical preprocessing step inmachine learning,focusing on identifying and preserving the most relevant features to improve the efficiency and performance of classification algorithms.Particle Swarm Optimization has demonstrated significant potential in addressing feature selection challenges.However,there are inherent limitations in Particle Swarm Optimization,such as the delicate balance between exploration and exploitation,susceptibility to local optima,and suboptimal convergence rates,hinder its performance.To tackle these issues,this study introduces a novel Leveraged Opposition-Based Learning method within Fitness Landscape Particle Swarm Optimization,tailored for wrapper-based feature selection.The proposed approach integrates:(1)a fitness-landscape adaptive strategy to dynamically balance exploration and exploitation,(2)the lever principle within Opposition-Based Learning to improve search efficiency,and(3)a Local Selection and Re-optimization mechanism combined with random perturbation to expedite convergence and enhance the quality of the optimal feature subset.The effectiveness of is rigorously evaluated on 24 benchmark datasets and compared against 13 advancedmetaheuristic algorithms.Experimental results demonstrate that the proposed method outperforms the compared algorithms in classification accuracy on over half of the datasets,whilst also significantly reducing the number of selected features.These findings demonstrate its effectiveness and robustness in feature selection tasks.
基金supported in part by the National Natural Science Foundation of China(62073158)the Key Science and Technology Research Project of the Education Department of Liaoning Province(LJ222410148037)the“Xingliao Talent Program”of Liaoning Province(XLYC2402025,XLYC2203160)。
摘要This paper addresses an optimal sensor selection problem under the framework of linear quadratic regulation.Unlike prior work on optimal sensor scheduling,we assume that the sensor noise covariance matrices are comparable but unknown.Then,the optimal sensor selection problem is formulated as finding an optimal policy of selecting a sensor from a set of sensors to minimize the expected quadratic performance of a linear system given the number of trials.An action value method from reinforcement learning is adopted for estimating the values of selections and making selection decisions based on the estimates.Several ways of balancing exploration and exploitation are presented and compared for efficacy.Numerical simulations are conducted to demonstrate the effectiveness of the proposed algorithms.
基金supported by grants from the Henan Province Science Foundation of Excellent Young Scholars(Grant 242300421171)the National Natural Science Foundation of China(Grants 62106066,62506109 and 62576128)the Zhejiang Provincial Natural Science Foundation of China under(Grant LMS26F020035).
摘要With the advancement of brain–computer interfaces(BCI),motor imagery(MI)electroencephalogram(EEG)decoding can greatly benefit from spatial filtering features derived from common spatial patterns(CSP).However,CSP-based features often exhibit high redundancy and intersubject variability.These limitations make the feature selection methods based on sparse learning difficult to effectively balance the heterogeneous contributions of different temporal and spatial components.Moreover,these models tend to prioritise features with larger coefficients,potentially overlooking intrinsic feature importance and compromising the quality of the selected feature subset.To address these issues,we propose an Adaptive Sparse Group Lasso(ASGL)method for structured feature selection,designed to enhance discriminative CSP features whilst suppressing irrelevant components.The proposed method partitions EEG signals into consecutive segments using a sliding window,treating each as a separate feature group.Benefiting from this,the importance of features at both the group level and the within-group level can be effectively quantified through mutual information and copula mutual information,thereby assigning adaptive weights for selective penalisation within the model.This weight construction strategy preserves important features from relevant time intervals and frequency bands.The resulting optimization problem is solved efficiently via the alternating direction method of multipliers(ADMM).Evaluations on simulated and real-world datasets demonstrate that the proposed ASGL outperforms existing methods.
基金funded by the Huaiyin Institute of Technology—Institute of Smart Energy.
摘要In the quest to enhance energy efficiency and reduce environmental impact in the transportation sector,the recovery of waste heat from diesel engines has become a critical area of focus.This study provided an exhaustive thermodynamic analysis optimizing Organic Rankine Cycle(ORC)systems forwaste heat recovery fromdiesel engines.Thestudy assessed the performance of five candidateworking fluids—R11,R123,R113,R245fa,and R141b—under a range of operating conditions,specifically varying overheat temperatures and evaporation pressures.The results indicated that the choice of working fluid substantially influences the system’s exergetic efficiency,net output power,and thermal efficiency.R245fa showed an outstanding net output power of 30.39 kW at high overheat conditions,outperforming R11,which is significant for high-temperature waste heat recovery.At lower temperatures,R11 and R113 demonstrated higher exergetic efficiencies,with R11 reaching a peak exergetic efficiency of 7.4%at an evaporation pressure of 10 bar and an overheat of 10℃.The study also revealed that controlling the overheat and optimizing the evaporation pressure are crucial for enhancing the net output power of the ORC system.Specifically,at an evaporation pressure of 30 bar and an overheat of 0℃,R113 exhibited the lowest exergetic destruction of 544.5 kJ/kg,making it a suitable choice for minimizing irreversible losses.These findings are instrumental for understanding the performance of ORC systems in waste heat recovery applications and offer valuable insights for the design and operation of more efficient and environmentally friendly diesel engine systems.
摘要Multi-label feature selection(MFS)is a crucial dimensionality reduction technique aimed at identifying informative features associated with multiple labels.However,traditional centralized methods face significant challenges in privacy-sensitive and distributed settings,often neglecting label dependencies and suffering from low computational efficiency.To address these issues,we introduce a novel framework,Fed-MFSDHBCPSO—federated MFS via dual-layer hybrid breeding cooperative particle swarm optimization algorithm with manifold and sparsity regularization(DHBCPSO-MSR).Leveraging the federated learning paradigm,Fed-MFSDHBCPSO allows clients to perform local feature selection(FS)using DHBCPSO-MSR.Locally selected feature subsets are encrypted with differential privacy(DP)and transmitted to a central server,where they are securely aggregated and refined through secure multi-party computation(SMPC)until global convergence is achieved.Within each client,DHBCPSO-MSR employs a dual-layer FS strategy.The inner layer constructs sample and label similarity graphs,generates Laplacian matrices to capture the manifold structure between samples and labels,and applies L2,1-norm regularization to sparsify the feature subset,yielding an optimized feature weight matrix.The outer layer uses a hybrid breeding cooperative particle swarm optimization algorithm to further refine the feature weight matrix and identify the optimal feature subset.The updated weight matrix is then fed back to the inner layer for further optimization.Comprehensive experiments on multiple real-world multi-label datasets demonstrate that Fed-MFSDHBCPSO consistently outperforms both centralized and federated baseline methods across several key evaluation metrics.
摘要With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the model in identifying non-standard semantics due to its semantic complexity and high-risk features.However,existing fine-tuning methods rely heavily on static data selection strategies,making it difficult to adapt to the dynamic evolution of model capabilities,resulting in low training efficiency.This article proposes ADS(Adaptive Dataset Selection),an adaptive framework for selecting data in anomaly text detection.ADS performs model-aware data selection prior to fine-tuning,adapting the initial state of pre-trained language models by selecting samples that are most informative for the target anomaly detection task.Empirical results on mainstream large language model architectures show that ADS significantly compresses data size while still outperforming existing static strategies and mainstream compression methods.When using only 1000 fine-tuning samples,ADS achieves a 92%F1 score,with an accuracy improvement of over 22%compared to the baseline,demonstrating excellent performance.This study proposes an efficient data selection mechanism from the perspective of model capability and dynamic adaptation of data,providing theoretical support and a practical path for fine-tuning large models in low-resource scenarios.
基金The Article Processing Charge(APC)for the publication of this research was funded by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior-Brasil(CAPES)(ROR identifier:00x0ma614).
摘要Sleep-dependent memory consolidation relies on coordinated interactions between hippocampal replay,cortical oscillations,and brain state transitions that unfold across the sleep onset period and early non-rapid eye movement(NREM)sleep[1].While substantial progress has been made in identifying the neuronal substrates of replay and its role in systems consolidation[2,3],less attention has been paid to how global physiological states shape the conditions under which distinct replay patterns are expressed and effectively coupled to downstream consolidation processes.In particular,autonomic regulation during the wake-to-sleep transition has emerged as a key modulator of NREM sleep depth,oscillatory synchronization,and hippocampal cortical communication[4],yet its potential influence on the structure of hippocampal replay remains largely unexplored.
基金supported by Basic Science Research Program through the National Research Foundation of Korea(NRF)funded by the Ministry of Education(RS-2020-NR049579).
摘要High-dimensional data causes difficulties in machine learning due to high time consumption and large memory requirements.In particular,in amulti-label environment,higher complexity is required asmuch as the number of labels.Moreover,an optimization problem that fully considers all dependencies between features and labels is difficult to solve.In this study,we propose a novel regression-basedmulti-label feature selectionmethod that integrates mutual information to better exploit the underlying data structure.By incorporating mutual information into the regression formulation,the model captures not only linear relationships but also complex non-linear dependencies.The proposed objective function simultaneously considers three types of relationships:(1)feature redundancy,(2)featurelabel relevance,and(3)inter-label dependency.These three quantities are computed usingmutual information,allowing the proposed formulation to capture nonlinear dependencies among variables.These three types of relationships are key factors in multi-label feature selection,and our method expresses them within a unified formulation,enabling efficient optimization while simultaneously accounting for all of them.To efficiently solve the proposed optimization problem under non-negativity constraints,we develop a gradient-based optimization algorithm with fast convergence.Theexperimental results on sevenmulti-label datasets show that the proposed method outperforms existingmulti-label feature selection techniques.
基金supported by the Jilin Science and Technology Development Program,China (20240602032RC)the Jilin Agricultural Science and Technology Innovation Project,China (CXGC2024ZD001)+1 种基金the Jilin Agricultural Science and Technology Innovation Project,China (CXGC2024ZY012)the Jilin Province Development and Reform Commission-Project for Improving the Independent Innovation Capacity of Major Grain Crops,China (2024C002)。
摘要Emerging and powerful genome editing tools,particularly CRISPR/Cas9,are facilitating functional genomics research and accelerating crop improvement(Jiang et al.2021;Cao et al.2023;Chen C et al.2023;Liu et al.2023a).However,the detection and screening of transgenic lines remain major bottlenecks,being time-consuming,labor-intensive,and inefficient during transformation and subsequent mutation identification.A simple and efficient visual marker system plays a critical role in addressing these challenges.Recent studies demonstrated that the GmW1 and RUBY reporter systems were used to obtain visual transgenic soybean(Glycine max) plants(Chen L et al.2023;Chen et al.2024).
基金sponsored by Shandong Provincial Key R&D Plan(2024LZGCQY012,2022LZG002)the National Natural Science Foundation of China(32472134)+1 种基金the Taishan Scholar Young Expert(20230119)Yantai Science and Technology Plan Project(2023ZDCX023)。
摘要We present Hi4GS,a hybrid feature selection(HFS)algorithm for selecting SNP subsets from highdimensional genotypes to improve the prediction of genomic estimated breeding value(GEBV)under genomic selection(GS).Hi4GS combines feature importance weighting with quantity determining to construct a fused feature set from which it extracts an optimal feature subset for subsequent GS.In a study of wheat using four datasets covering 11 yield traits via large-scale GS models,the SNPs selected by Hi4GS increased the average predictive accuracy by over 82%than using all SNPs.Hi4GS was used to identify SNPs potentially affecting wheat yield,and SHAP-based interpretability was applied to explain the contributions of these SNPs and their potential interactions.Hi4GS can be used for assisting in improving the prediction accuracy of GS,wheat and other plants'yield-associated SNPs identification,and target information for breeding chip development.The free R package Hi4GS is available at http://gffzz188fe103f8f1460as0ucunx6b90o66kvw.ffgz.tsg.suse.edu.cn/shgs19/Hi4GS.