Reconfigurable intelligent surface(RIS)have been cast as a promising alternative to alleviate blockage vulnerability and enhance coverage capability for terahertz(THz)communications.Owing to large-scale array elements...Reconfigurable intelligent surface(RIS)have been cast as a promising alternative to alleviate blockage vulnerability and enhance coverage capability for terahertz(THz)communications.Owing to large-scale array elements at transceivers and RIS,the codebook based beamforming can be utilized in a computationally efficient manner.However,the codeword selection for analog beamforming is an intractable combinatorial optimization(CO)problem.To this end,by taking the CO problem as a classification problem,a multi-task learning based analog beam selection(MTL-ABS)framework is developed to implement cooperative beam selection concurrently at transceivers and RIS.In addition,residual network and self-attention mechanism are used to combat the network degradation and mine intrinsic THz channel features.Finally,the network convergence is analyzed from a blockwise perspective,and numerical results demonstrate that the MTL-ABS framework greatly decreases the beam selection overhead and achieves near optimal sum-rate compared with heuristic search based counterparts.展开更多
Underground engineering projects such as deep tunnel excavation often encounter rockburst disasters accompanied by numerous microseismic events.Rapid interpretation of microseismic signals is crucial for the timely id...Underground engineering projects such as deep tunnel excavation often encounter rockburst disasters accompanied by numerous microseismic events.Rapid interpretation of microseismic signals is crucial for the timely identification of rockbursts.However,conventional processing encompasses multi-step workflows,including classification,denoising,picking,locating,and computational analysis,coupled with manual intervention,which collectively compromise the reliability of early warnings.To address these challenges,this study innovatively proposes the“microseismic stethoscope"-a multi-task machine learning and deep learning model designed for the automated processing of massive microseismic signals.This model efficiently extracts three key parameters that are necessary for recognizing rockburst disasters:rupture location,microseismic energy,and moment magnitude.Specifically,the model extracts raw waveform features from three dedicated sub-networks:a classifier for source zone classification,and two regressors for microseismic energy and moment magnitude estimation.This model demonstrates superior efficiency compared to traditional processing and semi-automated processing,reducing per-event processing time from 0.71 s to 0.49 s to merely 0.036 s.It concurrently achieves 98%accuracy in source zone classification,with microseismic energy and moment magnitude estimation errors of 0.13 and 0.05,respectively.This model has been well applied and validated in the Daxiagu Tunnel case in Sichuan,China.The application results indicate that the model is as accurate as traditional methods in determining source parameters,and thus can be used to identify potential geomechanical processes of rockburst disasters.By enhancing the signal processing reliability of microseismic events,the proposed model in this study presents a significant advancement in the identification of rockburst disasters.展开更多
Convolutional neural networks(CNNs)have shown remarkable success across numerous tasks such as image classification,yet the theoretical understanding of their convergence remains underdeveloped compared to their empir...Convolutional neural networks(CNNs)have shown remarkable success across numerous tasks such as image classification,yet the theoretical understanding of their convergence remains underdeveloped compared to their empirical achievements.In this paper,the first filter learning framework with convergence-guaranteed learning laws for end-to-end learning of deep CNNs is proposed.Novel update laws with convergence analysis are formulated based on the mathematical representation of each layer in convolutional neural networks.The proposed learning laws enable concurrent updates of weights across all layers of the deep convolutional neural network and the analysis shows that the training errors converge to certain bounds which are dependent on the approximation errors.Case studies are conducted on benchmark datasets and the results show that the proposed concurrent filter learning framework guarantees the convergence and offers more consistent and reliable results during training with a trade-off in performance compared to stochastic gradient descent methods.This framework represents a significant step towards enhancing the reliability and effectiveness of deep convolutional neural network by developing a theoretical analysis which allows practical implementation of the learning laws with automatic tuning of the learning rate to guarantee the convergence during training.展开更多
The edge deployment of artificial intelligence has driven the exploitation of compact,energy-efficient information processing systems that integrate sensing,memory,and multi-task processing functions.However,conventio...The edge deployment of artificial intelligence has driven the exploitation of compact,energy-efficient information processing systems that integrate sensing,memory,and multi-task processing functions.However,conventional vision systems suffer from significant energyime overhead,extra hardware costs,and an unaffordable algorithm.Herein,we demonstrate an in-sensor computing system employing reconfigurable optoelectronic transistors(ROETs)for multi-task learning.These transistors exhibit reconfigurable volatile and nonvolatile characteristics under both optical and electrical stimuli.Capitalizing on this reconfigurability,we establish an in-sensor reservoir computing(RC)system operating in multi-signal modes:volatile dynamics function as the reservoir,whereas nonvolatile properties configure the readout layer.The abundant optoelectronic reservoir states display exceptional feature separability and prolonged stability in the ambient atmosphere.Such a reliable RC system successfully achieves multi-task processing of images.Notably,under the optoelectronic coordination mode,it effectively alleviates feature degradation while sustaining consistently high recognition accuracy.Furthermore,the system exhibits remarkable dynamic information processing capabilities,achieving recognition accuracies of 89.02%for dynamic gestures and 96.04%for moving vehicles recognition,respectively.Supplemental functionalities,including light adaptation and image sharpening,are also implemented.This work presents a configurable multimodal platform featuring a flexible in-sensor reservoir computing architecture,providing a potential solution for efficient multi-task processing.展开更多
Real-time and accurate dynamic wake information is essential for wind resource assessment and the optimization of wind farm operations.To further understand the wake characteristics of wind turbines,we propose a hiera...Real-time and accurate dynamic wake information is essential for wind resource assessment and the optimization of wind farm operations.To further understand the wake characteristics of wind turbines,we propose a hierarchical learning approach integrated with a deep neural network-based prediction method.The integrated framework combines physical and mathematical models,enabling 3D spatiotemporal wind field predictions with minimal measured data requirements.Evaluation and validation results demonstrate that the proposed method achieves accurate ultra-short-term wake predictions across the entire domain with minimal prediction errors.Compared with conventional methods,the proposed hierarchical learning framework markedly lowers the training-data requirements of physics-informed neural networks for large-scale flow-field prediction while maintaining high accuracy.In addition,it demonstrates superior performance in both local and global wake forecasts,offering practical insights for efficient turbine operation and wake analysis.展开更多
Efficient and accurate anomaly detection in a network is of great significance for maintaining network and device security.Most anomaly detection methods assume that different anomalous network data distributions are ...Efficient and accurate anomaly detection in a network is of great significance for maintaining network and device security.Most anomaly detection methods assume that different anomalous network data distributions are the same or similar and ignore data privacy preservation.In this paper,a novel Federated Learning(FL)is proposed that it can quickly detect different types of anomalies in Non-Independent and Identically Distributed(Non-IID)data.First,we design a multi-domain machine learning model for multi-domain data,named Aegean,which consists of two modules:an ensemble AutoEncoder(AE)and a Generative Adversarial Network(GAN).Second,because data from different domains are non-IID,we model the anomaly detection problem as a dual problem,which can be recast as a robust optimization problem.The robust optimization problem is non-convex and therefore difficult to solve.As a remedy,we formulate and solve a dual problem by taking the Lagrangian dual function of the original problem.Experiments demonstrate that Aegean significantly outperforms the current state-of-the-art methods,with a 16%F1 score improvement over that of a One-Class Support Vector Machine(OCSVM).The designed FL significantly reduces the communication overhead of FedAvg without sacrificing anomaly detection performance.展开更多
Landslides are one of the most significant types of geological disasters worldwide.However,current landslide disaster recognition models generally face issues such as low training efficiency,reliance on large amounts ...Landslides are one of the most significant types of geological disasters worldwide.However,current landslide disaster recognition models generally face issues such as low training efficiency,reliance on large amounts of samples,and vague feature extraction.To address these problems,this paper proposes an enhanced intelligent recognition model based on an improved residual network.Taking remote sensing images from landslide-prone areas in Bijie City,Guizhou Province,China as the research subject,we implemented data augmentation techniques to expand the datasets,subsequently partitioning it into training and validation subsets at a73:27 ratio.During model training,a transfer learning strategy was employed to improve training efficiency through the utilization of ImageNet pre-trained weights.The Res Net50 network architecture was enhanced through the integration of channel attention and spatial attention mechanisms,enabling more effective extraction of critical landslide features.Comparative analysis indicates that after introducing the attention mechanisms,the enhanced model achieved an increase of 1.49% in accuracy,0.45% in precision,1.52% in F1 score,and 0.0297 in Kappa coefficient on the validation set.Notably,recall improved by 2.59%,indicating that the improved model has enhanced landslide disaster recognition capability and overall performance.This study successfully coupled the transfer learning strategy with a dual-attention mechanism into the ResNet50 architecture,allowing the rapid construction of an efficient recognition model under limited sample conditions,significantly improving the comprehensive performance and generalisation ability of the landslide recognition model.展开更多
This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to cap...This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to capture complex spatiotemporal relationships in healthcare demand and resource allocation patterns,combined with centralised training and a decentralised execution approach.The system is modelled as a Markov game and solved using a deep reinforcement learning algorithm.Extensive simulations demonstrate that HealthNet outperforms eight state-of-the-art baseline methods across multiple network configurations and evaluation metrics.In a 4×4 grid network,HealthNet reduces average waiting times by 47.6%compared to model predictive control and 22.1%compared to the best-performing baseline.Traffic congestion rates are reduced to 16.7%compared to 42.3%for the worst baseline and 23.1%for the best baseline.Under irregular network topologies with stochastic disruptions,including demand surges and vehicle unavailability,HealthNet maintains superior performance with 42.1%lower average waiting time and 51.1%improvement in peak response times compared to competing approaches.These findings indicate that HealthNet can enhance both efficiency and resilience in healthcare transportation systems,potentially improving patient outcomes in complex urban environments.展开更多
Shield tunneling is the method most commonly used for underground projects.Segment typesetting,which involves determining the optimal assembly point for each segment ring and sequentially assembling them into a comple...Shield tunneling is the method most commonly used for underground projects.Segment typesetting,which involves determining the optimal assembly point for each segment ring and sequentially assembling them into a complete tunnel,is a critical step in shield tunneling.Currently,this typesetting process relies heavily on the operator’s experience at construction sites,which does not guarantee quality.Furthermore,research focuses mainly on the commonly used 16-point segment typesetting,largely ignoring other segment types.In addition,the reliability of these studies in site applications remains unsatisfactory.To address these issues,we propose an intelligent method for segment typesetting using an artificial neural network(ANN)and transfer learning.Due to insufficient historical data for ANN training,a dataset creation method was devised based on the Monte Carlo method and manual annotation.An ANN model was then developed to typeset 16-point segments,with its hyperparameters optimized through Bayesian optimization.Subsequently,the trained model was adapted to other segment types via transfer learning,using 10-point segments as a case study.Based on the test set established in this study,our proposed method showed superior performance compared with several commonly used machine learning methods and a representative and well-validated segment typesetting method.It was also validated using real data collected from construction sites,achieving an accuracy of 93.75%for 16-point segments and 91.43%for 10-point segments,both of which significantly surpass the results from manual typesetting on-site.The proposed method achieves accurate,rapid,and intelligent segment typesetting,which is adaptable to various segment types.展开更多
Recently the depth estimation methods based on deep learning(DL)retain challenging to estimate a high-precision depth map in fringe projection structured light three-dimensional(3D)measurement with limited information...Recently the depth estimation methods based on deep learning(DL)retain challenging to estimate a high-precision depth map in fringe projection structured light three-dimensional(3D)measurement with limited information from a single-frame fringe pattern.In this letter,we proposed a FDSUNet++convolutional neural network(CNN),which consists of a UNet++base model,an improved squeeze-andexcitation(ISE)block,a Fourier transform(FT)data preprocessing block,and a discrete wavelet transform(DWT)block.The proposed ISE block can improve the ability of feature extraction and the designed FT data preprocessing block preserves the key features of the fringe pattern by FT.The introduced DWT block reduces the complexity and training cost of the model.By integrating these three blocks into the UNet++,it can better achieve depth estimation.Experimental results from two structured light datasets demonstrate that the proposed FDSUNet++outperforms the state-of-the-art networks,achieving the best performance in both qualitative and quantitative evaluation.展开更多
Wi-Fi technology has evolved significantly since its introduction in 1997,advancing to Wi-Fi 6 as the latest standard,with Wi-Fi 7 currently under development.Despite these advancements,integrating machine learning in...Wi-Fi technology has evolved significantly since its introduction in 1997,advancing to Wi-Fi 6 as the latest standard,with Wi-Fi 7 currently under development.Despite these advancements,integrating machine learning into Wi-Fi networks remains challenging,especially in decentralized environments with multiple access points(mAPs).This paper is a short review that summarizes the potential applications of federated reinforcement learning(FRL)across eight key areas of Wi-Fi functionality,including channel access,link adaptation,beamforming,multi-user transmissions,channel bonding,multi-link operation,spatial reuse,and multi-basic servic set(multi-BSS)coordination.FRL is highlighted as a promising framework for enabling decentralized training and decision-making while preserving data privacy.To illustrate its role in practice,we present a case study on link activation in a multi-link operation(MLO)environment with multiple APs.Through theoretical discussion and simulation results,the study demonstrates how FRL can improve performance and reliability,paving the way for more adaptive and collaborative Wi-Fi networks in the era of Wi-Fi 7 and beyond.展开更多
The rapid proliferation of Internet of Things(IoT)devices has increased the importance of network intrusion detection systems(NIDS)for protecting modern networks.However,many machine learning and deep learning based N...The rapid proliferation of Internet of Things(IoT)devices has increased the importance of network intrusion detection systems(NIDS)for protecting modern networks.However,many machine learning and deep learning based NIDS rely on large volumes of labeled attack data,which is often impractical to obtain for newly emerging or rare attacks.This paper presents a benchmark-style systematic evaluation of meta-learning-based FewShot Learning(FSL)classifiers for detecting previously unseen intrusions with limited labeled data.We investigate three representative FSL models,namely Prototypical Networks,Relation Networks,and MetaOptNet,and further examine two decision-level ensemble strategies based on majority voting and probability averaging.Experiments are conducted on the UQ-IoT-IDS-2021 dataset under four attack distributions and five K-shot settings(K∈{3,5,7,10,15}),with performance reported separately for seen and unseen attacks using class-balanced macro-averaged F1 score.The experimental results show that MetaOptNet consistently provides the strongest and most stable performance on unseen attacks.Furthermore,investigating decision-level ensembling reveals that probability averaging frequently matches or marginally outperforms the best base model by successfully mitigating weaker predictions,whereas majority voting demonstrates slight performance degradation due to strict consensus requirements.We also highlight prominent error modes,such as benign-attack ambiguity and cross-device attack correlations,alongside an analysis of model complexity and inference latency.These findings highlight the potential of few-shot learning for data-scarce intrusion detection and provide insights into model selection and architectural trade-offs under varying support set sizes and attack compositions.展开更多
Purpose-Rail freight is widely recognized for its economic and environmental advantages,yet it remains weakly integrated into firms’supply chains,particularly in emerging economies.This study aims to investigate the ...Purpose-Rail freight is widely recognized for its economic and environmental advantages,yet it remains weakly integrated into firms’supply chains,particularly in emerging economies.This study aims to investigate the conditions under which rail freight can be effectively integrated into multi-actor supply chains,with specific attention to the role of organizational coordination,information quality and artificial intelligence(AI)in shaping logistics integration outcomes.Design/methodology/approach-The study draws on a quantitative survey of 3,185 stakeholders involved in rail-based and multimodal supply chains in Morocco.The data are analyzed using a combination of machine learning,deep learning and artificial neural network models.These methods are used not only to identify the main determinants of rail freight integration but also to capture non-linear relationships,interaction effects and potential integration trajectories that cannot be addressed through conventional linear models.Findings-The results show that rail freight integration depends primarily on organizational and informational mechanisms rather than on infrastructure alone.Inter-organizational coordination and logistics information quality emerge as the most influential factors.AI contributes positively to rail freight integration,but its effect is conditional:AI tools significantly enhance integration only when adequate levels of coordination and information sharing are already in place.Scenario simulations further reveal that the strongest integration gains arise from the combined improvement of organizational practices and AI adoption.Originality/value-This research contributes to the literature by shifting the focus from infrastructure-centered explanations toward a systemic understanding of rail freight integration.It is among the first studies to empirically combine machine learning,deep learning and artificial neural networks to analyze logistics integration in an emerging-economy context and to show that AI functions as a complementary and amplifying mechanism rather than a standalone solution.展开更多
Surface properties of crystals are critical in many fields,including electrochemistry and photoelectronics,the efficient prediction of which can expedite the design and optimization of catalysts,batteries,alloys etc.H...Surface properties of crystals are critical in many fields,including electrochemistry and photoelectronics,the efficient prediction of which can expedite the design and optimization of catalysts,batteries,alloys etc.However,we are still far from realizing this vision due to the rarity of surface property-related databases,especially for multicomponent compounds,due to the large sample spaces and limited computing resources.In this work,we present a surface emphasized multi-task crystal graph convolutional neural network(SEM-CGCNN)to predict multiple surface properties simultaneously from crystal structures.The model is evaluated on a dataset of 3526 surface energies and work functions of binary magnesium intermetallics obtained through first-principles calculations,and obvious improvements are observed both in efficiency and accuracy over the original CGCNN model.By transferring the pre-trained model to the datasets of pure metals and other intermetallics,the fine-tuned SEM-CGCNN outperforms learning from scratch and can be further applied to other surface properties and materials systems.This study could be a paradigm for the end-to-end mapping of atomic structures to anisotropic surface properties of crystals,which provides an efficient framework to understand and screen materials with desired surface characteristics.展开更多
The increasing adoption of 5G cellular networks has introduced significant challenges for network operators.The main challenge lies in the management of seamless handoff(HO),which occurs owing to the rapid expansion o...The increasing adoption of 5G cellular networks has introduced significant challenges for network operators.The main challenge lies in the management of seamless handoff(HO),which occurs owing to the rapid expansion of equipment,data,and network complexity.To address this challenge,a model named optimal HO management deep learning neural network(OHMDLNN)is proposed.The model is trained on network activity data,and it uses KPIs(key performance indicators)and system-level parameters to make HO decisions.As demonstrated in the article,OHMDLNN is successful in analyzing the effect and interdependence of KPIs from both the network and user equipment(UE)perspectives.Moreover,the model is evaluated for accuracy(the percentage of correct decisions made by a model on a dataset)in comparison with existing neural network-based HO decision models.These include temporal convolution networks(TCN),recurrent neural networks(RNN),long short-term memory(LSTM),gated recurrent units(GRU),and convolutional neural networks(CNN).The dataset used to evaluate the performance of the model consisted of 65,000 records.The model demonstrates superior performance,with an average improvement on accuracy of 8 percent over TCN,18 percent over RNN,6 percent over LSTM,14 percent over GRU and 4 percent over CNN.Along with accuracy,the model is also tested on important performance indicators,including the packet loss rate,the success rate,latency,and throughput at the time of handover.These results affirm its efficiency in the HO decision-making process.Future research will consider the use of advanced deep learning architectures and simplify the process of integrating system-level inputs to optimize system performance during HO events.展开更多
Communication infrastructure is often among the first casualties in natural or human-induced disasters,severely impairing the coordination and efficiency of rescue operations.Rapid deployment of Unmanned Aerial Vehicl...Communication infrastructure is often among the first casualties in natural or human-induced disasters,severely impairing the coordination and efficiency of rescue operations.Rapid deployment of Unmanned Aerial Vehicles(UAVs)and satellite systems has thus become essential for establishing robust communication links to support rescue-critical tasks.However,existing emergency communication networks rely heavily on domain expertise for topology design,thereby suffering from issues such as inefficient resource allocation and network congestion,among others.To address these challenges,we present TopoLLM,a framework that leverages Large Language Models(LLMs)for tool-driven optimization of emergency network topologies.This framework effectively combines the reasoning capabilities of the LLM with TopoTool,a domain-specific optimization toolkit engineered for high-precision and load-balanced network planning in disaster scenarios.Guided by an adaptive toolselection mechanism,TopoLLM autonomously generates resilient topologies and allocates resources intelligently,reducing the need for extensive human interventions.Experimental evaluations on simulated disaster scenarios verify that TopoLLM can rapidly generate high-accuracy and robust topologies,achieving notable performance improvements compared with existing approaches.展开更多
Dear Editor,This letter proposes a fully distributed multi-agent reinforcement learning(DMARL)algorithm for the coordinated optimization and scheduling of source-load-storage in networked microgrids.To accommodate the...Dear Editor,This letter proposes a fully distributed multi-agent reinforcement learning(DMARL)algorithm for the coordinated optimization and scheduling of source-load-storage in networked microgrids.To accommodate the rapid development of networked microgrids,we have designed a DMARL algorithm with an event-triggered mechanism(ETM).Unlike centralized approaches,DMARL empowers individual agents to cooperatively learn and optimize based only on local observations,while reducing the communication burden.Simulation results validate the performance of the proposed algorithm,demonstrating its effectiveness in efficient microgrid resource management.展开更多
Fe-based amorphous alloys are promising soft magnetic materials for developing next-generation devices with high frequency and efficiency.However,optimizing Fe-based alloys with ultra-high saturation magnetic flux den...Fe-based amorphous alloys are promising soft magnetic materials for developing next-generation devices with high frequency and efficiency.However,optimizing Fe-based alloys with ultra-high saturation magnetic flux density(Bs),ultralow coercivity(Hc),and good glass-forming ability remains a notorious challenge owing to the vast composition space and complex trade-offs among these properties.Thus,conventional design methods face great challenges.Here,we develop a generative multi-task deep learning(GMTDL)approach to achieve simultaneous optimization of compositions and tradeoff properties.The GMTDL can sufficiently exploit and share knowledge from datasets across different tasks,despite the limitations and imbalances of these datasets.Therefore,it exhibits superior performance in predicting alloys with multiple targeted properties,outperforming previous machine learning-based design strategies.Moreover,the GMTDL can also tailor compositions,providing an efficient way to regulate properties and generate desired candidates for further experimental processing.The validity and reliability of GMTDL are rigorously tested by benchmarking against Fe-based alloys reported very recently.Moreover,some new alloys with ultra-high Bs and ultra-low Hc are predicted.The optimal content windows of key elements and their synergistic effects are also unraveled,providing practical guidance.Thus,our study establishes an effective and reliable paradigm for simultaneous prediction and optimization of high-performance materials with multiple properties.展开更多
With the growing complexity and decentralization of network systems,the attack surface has expanded,which has led to greater concerns over network threats.In this context,artificial intelligence(AI)-based network intr...With the growing complexity and decentralization of network systems,the attack surface has expanded,which has led to greater concerns over network threats.In this context,artificial intelligence(AI)-based network intrusion detection systems(NIDS)have been extensively studied,and recent efforts have shifted toward integrating distributed learning to enable intelligent and scalable detection mechanisms.However,most existing works focus on individual distributed learning frameworks,and there is a lack of systematic evaluations that compare different algorithms under consistent conditions.In this paper,we present a comprehensive evaluation of representative distributed learning frameworks—Federated Learning(FL),Split Learning(SL),hybrid collaborative learning(SFL),and fully distributed learning—in the context of AI-driven NIDS.Using recent benchmark intrusion detection datasets,a unified model backbone,and controlled distributed scenarios,we assess these frameworks across multiple criteria,including detection performance,communication cost,computational efficiency,and convergence behavior.Our findings highlight distinct trade-offs among the distributed learning frameworks,demonstrating that the optimal choice depends strongly on systemconstraints such as bandwidth availability,node resources,and data distribution.This work provides the first holistic analysis of distributed learning approaches for AI-driven NIDS and offers practical guidelines for designing secure and efficient intrusion detection systems in decentralized environments.展开更多
Visual speech recognition(VSR)aims to infer spoken content from visual observations of articulatory movements.Despite significant progress,it remains a challenging task in computer vision and speech processing.Its dif...Visual speech recognition(VSR)aims to infer spoken content from visual observations of articulatory movements.Despite significant progress,it remains a challenging task in computer vision and speech processing.Its difficulty arises from pronounced speaker-to-speaker variability,the presence of homophenes(phonemes that are visually indistinguishable),changes in illumination,and the intrinsically high-dimensional nature of spatiotemporal lip dynamics.In this work,we propose NestLipGNN,a graph-based framework that integrates Graph Neural Networks(GNNs)with a nested multi-granularity learning strategy for visual speech recognition.We construct dynamic lip graphs from facial landmarks to model both spatial relationships between lip regions and their temporal motion during speech articulation.The proposed nested learning architecture supports hierarchical feature extraction across several levels of linguistic abstraction,spanning phoneme-level articulatory units,viseme-level visual speech categories,and word-level semantic representations.We further introduce a Temporal Graph Attention mechanism(T-GAT)that adaptively reweights the importance of distinct lip regions over time.We also introduce a graph-based contrastive learning objective to improve the discrimination of visually similar speech patterns,directly confronting the challenge of homophene resolution.Experiments on the LRW,LRS2,LRS3,and GRID datasets show that NestLipGNN improves recognition accuracy compared with existing methods,obtaining 92.3%word-level accuracy on LRW and delivering a 2.1%absolute performance gain over prior methods.Comprehensive ablation analyses confirm the contribution of each architectural component.展开更多
摘要Reconfigurable intelligent surface(RIS)have been cast as a promising alternative to alleviate blockage vulnerability and enhance coverage capability for terahertz(THz)communications.Owing to large-scale array elements at transceivers and RIS,the codebook based beamforming can be utilized in a computationally efficient manner.However,the codeword selection for analog beamforming is an intractable combinatorial optimization(CO)problem.To this end,by taking the CO problem as a classification problem,a multi-task learning based analog beam selection(MTL-ABS)framework is developed to implement cooperative beam selection concurrently at transceivers and RIS.In addition,residual network and self-attention mechanism are used to combat the network degradation and mine intrinsic THz channel features.Finally,the network convergence is analyzed from a blockwise perspective,and numerical results demonstrate that the MTL-ABS framework greatly decreases the beam selection overhead and achieves near optimal sum-rate compared with heuristic search based counterparts.
基金supported by the National Natural Science Foundation of China(Grant Nos.42130719 and 42177173)the Doctoral Direct Train Project of Chongqing Natural Science Foundation(Grant No.CSTB2023NSCQ-BSX0029).
摘要Underground engineering projects such as deep tunnel excavation often encounter rockburst disasters accompanied by numerous microseismic events.Rapid interpretation of microseismic signals is crucial for the timely identification of rockbursts.However,conventional processing encompasses multi-step workflows,including classification,denoising,picking,locating,and computational analysis,coupled with manual intervention,which collectively compromise the reliability of early warnings.To address these challenges,this study innovatively proposes the“microseismic stethoscope"-a multi-task machine learning and deep learning model designed for the automated processing of massive microseismic signals.This model efficiently extracts three key parameters that are necessary for recognizing rockburst disasters:rupture location,microseismic energy,and moment magnitude.Specifically,the model extracts raw waveform features from three dedicated sub-networks:a classifier for source zone classification,and two regressors for microseismic energy and moment magnitude estimation.This model demonstrates superior efficiency compared to traditional processing and semi-automated processing,reducing per-event processing time from 0.71 s to 0.49 s to merely 0.036 s.It concurrently achieves 98%accuracy in source zone classification,with microseismic energy and moment magnitude estimation errors of 0.13 and 0.05,respectively.This model has been well applied and validated in the Daxiagu Tunnel case in Sichuan,China.The application results indicate that the model is as accurate as traditional methods in determining source parameters,and thus can be used to identify potential geomechanical processes of rockburst disasters.By enhancing the signal processing reliability of microseismic events,the proposed model in this study presents a significant advancement in the identification of rockburst disasters.
基金supported by the Ministry of Education(MOE)Singapore,Academic Research Fund(AcRF)Tier 1(RG65/22)。
摘要Convolutional neural networks(CNNs)have shown remarkable success across numerous tasks such as image classification,yet the theoretical understanding of their convergence remains underdeveloped compared to their empirical achievements.In this paper,the first filter learning framework with convergence-guaranteed learning laws for end-to-end learning of deep CNNs is proposed.Novel update laws with convergence analysis are formulated based on the mathematical representation of each layer in convolutional neural networks.The proposed learning laws enable concurrent updates of weights across all layers of the deep convolutional neural network and the analysis shows that the training errors converge to certain bounds which are dependent on the approximation errors.Case studies are conducted on benchmark datasets and the results show that the proposed concurrent filter learning framework guarantees the convergence and offers more consistent and reliable results during training with a trade-off in performance compared to stochastic gradient descent methods.This framework represents a significant step towards enhancing the reliability and effectiveness of deep convolutional neural network by developing a theoretical analysis which allows practical implementation of the learning laws with automatic tuning of the learning rate to guarantee the convergence during training.
基金financially supported by the National Natural Science Foundation of China(Grant Nos.52202156 and 52303306)the support from Anhui Project(Grant No.Z010118169)+3 种基金The University Synergy Innovation Program of Anhui Province(Grant No.GXXT-2022-012)Key Natural Science Research Projects in Colleges and Universities in Anhui Province(Grant No.KJ2021A1088)Scientific Research Project of Colleges and Universities in Anhui Province(Grant No.2022AH050113)Postdoctoral Daily Public Start-Up Funds of Anhui University(Grant No.S202418001/069)。
摘要The edge deployment of artificial intelligence has driven the exploitation of compact,energy-efficient information processing systems that integrate sensing,memory,and multi-task processing functions.However,conventional vision systems suffer from significant energyime overhead,extra hardware costs,and an unaffordable algorithm.Herein,we demonstrate an in-sensor computing system employing reconfigurable optoelectronic transistors(ROETs)for multi-task learning.These transistors exhibit reconfigurable volatile and nonvolatile characteristics under both optical and electrical stimuli.Capitalizing on this reconfigurability,we establish an in-sensor reservoir computing(RC)system operating in multi-signal modes:volatile dynamics function as the reservoir,whereas nonvolatile properties configure the readout layer.The abundant optoelectronic reservoir states display exceptional feature separability and prolonged stability in the ambient atmosphere.Such a reliable RC system successfully achieves multi-task processing of images.Notably,under the optoelectronic coordination mode,it effectively alleviates feature degradation while sustaining consistently high recognition accuracy.Furthermore,the system exhibits remarkable dynamic information processing capabilities,achieving recognition accuracies of 89.02%for dynamic gestures and 96.04%for moving vehicles recognition,respectively.Supplemental functionalities,including light adaptation and image sharpening,are also implemented.This work presents a configurable multimodal platform featuring a flexible in-sensor reservoir computing architecture,providing a potential solution for efficient multi-task processing.
基金supported by the National Natural Science Foundation of China(No.62201226)the Guangdong Basic and Applied Basic Research Foundation(No.2022A1515240021),China.
摘要Real-time and accurate dynamic wake information is essential for wind resource assessment and the optimization of wind farm operations.To further understand the wake characteristics of wind turbines,we propose a hierarchical learning approach integrated with a deep neural network-based prediction method.The integrated framework combines physical and mathematical models,enabling 3D spatiotemporal wind field predictions with minimal measured data requirements.Evaluation and validation results demonstrate that the proposed method achieves accurate ultra-short-term wake predictions across the entire domain with minimal prediction errors.Compared with conventional methods,the proposed hierarchical learning framework markedly lowers the training-data requirements of physics-informed neural networks for large-scale flow-field prediction while maintaining high accuracy.In addition,it demonstrates superior performance in both local and global wake forecasts,offering practical insights for efficient turbine operation and wake analysis.
基金supported by the MSIT(Ministry of Science and ICT),Korea,under the ITRC(Information Technology Research Center)support program(IITP-2023-2018-0-01431)supervised by the IITP(Institute for Information&Communications Technology Planning&Evaluation)the Brain Korea 21(BK21)FOUR program of the National Research Foundation of Korea funded by the Ministry of Education(NRF5199991514504).
摘要Efficient and accurate anomaly detection in a network is of great significance for maintaining network and device security.Most anomaly detection methods assume that different anomalous network data distributions are the same or similar and ignore data privacy preservation.In this paper,a novel Federated Learning(FL)is proposed that it can quickly detect different types of anomalies in Non-Independent and Identically Distributed(Non-IID)data.First,we design a multi-domain machine learning model for multi-domain data,named Aegean,which consists of two modules:an ensemble AutoEncoder(AE)and a Generative Adversarial Network(GAN).Second,because data from different domains are non-IID,we model the anomaly detection problem as a dual problem,which can be recast as a robust optimization problem.The robust optimization problem is non-convex and therefore difficult to solve.As a remedy,we formulate and solve a dual problem by taking the Lagrangian dual function of the original problem.Experiments demonstrate that Aegean significantly outperforms the current state-of-the-art methods,with a 16%F1 score improvement over that of a One-Class Support Vector Machine(OCSVM).The designed FL significantly reduces the communication overhead of FedAvg without sacrificing anomaly detection performance.
基金supported by the National Natural Science Foundation of China[NSFC,Grant Nos.U22A20597,42507217]the"Unveiling and Commanding"Project of Science and Technology Program of Tibet[Grant No.XZ202303ZY0006G]the"Key Research and Development Program"Project of Science and Technology Program of Tibet[Grant Nos.XZ202501ZY0104,XZ202501ZY0132]。
摘要Landslides are one of the most significant types of geological disasters worldwide.However,current landslide disaster recognition models generally face issues such as low training efficiency,reliance on large amounts of samples,and vague feature extraction.To address these problems,this paper proposes an enhanced intelligent recognition model based on an improved residual network.Taking remote sensing images from landslide-prone areas in Bijie City,Guizhou Province,China as the research subject,we implemented data augmentation techniques to expand the datasets,subsequently partitioning it into training and validation subsets at a73:27 ratio.During model training,a transfer learning strategy was employed to improve training efficiency through the utilization of ImageNet pre-trained weights.The Res Net50 network architecture was enhanced through the integration of channel attention and spatial attention mechanisms,enabling more effective extraction of critical landslide features.Comparative analysis indicates that after introducing the attention mechanisms,the enhanced model achieved an increase of 1.49% in accuracy,0.45% in precision,1.52% in F1 score,and 0.0297 in Kappa coefficient on the validation set.Notably,recall improved by 2.59%,indicating that the improved model has enhanced landslide disaster recognition capability and overall performance.This study successfully coupled the transfer learning strategy with a dual-attention mechanism into the ResNet50 architecture,allowing the rapid construction of an efficient recognition model under limited sample conditions,significantly improving the comprehensive performance and generalisation ability of the landslide recognition model.
基金supported by the National Natural Science Foundation of China under No.62202247.
摘要This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to capture complex spatiotemporal relationships in healthcare demand and resource allocation patterns,combined with centralised training and a decentralised execution approach.The system is modelled as a Markov game and solved using a deep reinforcement learning algorithm.Extensive simulations demonstrate that HealthNet outperforms eight state-of-the-art baseline methods across multiple network configurations and evaluation metrics.In a 4×4 grid network,HealthNet reduces average waiting times by 47.6%compared to model predictive control and 22.1%compared to the best-performing baseline.Traffic congestion rates are reduced to 16.7%compared to 42.3%for the worst baseline and 23.1%for the best baseline.Under irregular network topologies with stochastic disruptions,including demand surges and vehicle unavailability,HealthNet maintains superior performance with 42.1%lower average waiting time and 51.1%improvement in peak response times compared to competing approaches.These findings indicate that HealthNet can enhance both efficiency and resilience in healthcare transportation systems,potentially improving patient outcomes in complex urban environments.
基金supported by the National Key Research and Development Program of China(No.2022YFC3802302)the Project of Institute of Advanced Machines,Zhejiang University(No.KY202404-058).
摘要Shield tunneling is the method most commonly used for underground projects.Segment typesetting,which involves determining the optimal assembly point for each segment ring and sequentially assembling them into a complete tunnel,is a critical step in shield tunneling.Currently,this typesetting process relies heavily on the operator’s experience at construction sites,which does not guarantee quality.Furthermore,research focuses mainly on the commonly used 16-point segment typesetting,largely ignoring other segment types.In addition,the reliability of these studies in site applications remains unsatisfactory.To address these issues,we propose an intelligent method for segment typesetting using an artificial neural network(ANN)and transfer learning.Due to insufficient historical data for ANN training,a dataset creation method was devised based on the Monte Carlo method and manual annotation.An ANN model was then developed to typeset 16-point segments,with its hyperparameters optimized through Bayesian optimization.Subsequently,the trained model was adapted to other segment types via transfer learning,using 10-point segments as a case study.Based on the test set established in this study,our proposed method showed superior performance compared with several commonly used machine learning methods and a representative and well-validated segment typesetting method.It was also validated using real data collected from construction sites,achieving an accuracy of 93.75%for 16-point segments and 91.43%for 10-point segments,both of which significantly surpass the results from manual typesetting on-site.The proposed method achieves accurate,rapid,and intelligent segment typesetting,which is adaptable to various segment types.
基金supported by the National Natural Science Foundation of China(No.61905178)。
摘要Recently the depth estimation methods based on deep learning(DL)retain challenging to estimate a high-precision depth map in fringe projection structured light three-dimensional(3D)measurement with limited information from a single-frame fringe pattern.In this letter,we proposed a FDSUNet++convolutional neural network(CNN),which consists of a UNet++base model,an improved squeeze-andexcitation(ISE)block,a Fourier transform(FT)data preprocessing block,and a discrete wavelet transform(DWT)block.The proposed ISE block can improve the ability of feature extraction and the designed FT data preprocessing block preserves the key features of the fringe pattern by FT.The introduced DWT block reduces the complexity and training cost of the model.By integrating these three blocks into the UNet++,it can better achieve depth estimation.Experimental results from two structured light datasets demonstrate that the proposed FDSUNet++outperforms the state-of-the-art networks,achieving the best performance in both qualitative and quantitative evaluation.
基金funded by the Deanship of Scientific Research(DSR)at King Abdulaziz University,Jeddah,Saudi Arabia,grant number RG-2-611-42(A.O.A.).
摘要Wi-Fi technology has evolved significantly since its introduction in 1997,advancing to Wi-Fi 6 as the latest standard,with Wi-Fi 7 currently under development.Despite these advancements,integrating machine learning into Wi-Fi networks remains challenging,especially in decentralized environments with multiple access points(mAPs).This paper is a short review that summarizes the potential applications of federated reinforcement learning(FRL)across eight key areas of Wi-Fi functionality,including channel access,link adaptation,beamforming,multi-user transmissions,channel bonding,multi-link operation,spatial reuse,and multi-basic servic set(multi-BSS)coordination.FRL is highlighted as a promising framework for enabling decentralized training and decision-making while preserving data privacy.To illustrate its role in practice,we present a case study on link activation in a multi-link operation(MLO)environment with multiple APs.Through theoretical discussion and simulation results,the study demonstrates how FRL can improve performance and reliability,paving the way for more adaptive and collaborative Wi-Fi networks in the era of Wi-Fi 7 and beyond.
基金supported by the Korea Institute of Energy Technology Evaluation and Planning(KETEP)grant funded by the Korea government(MOTIE)(A Study on Development of Cyber-Physical Attack Response System and Security Management System for Maximizing Availability of Real-Time Distributed Resources,RS-2023-00303559).
摘要The rapid proliferation of Internet of Things(IoT)devices has increased the importance of network intrusion detection systems(NIDS)for protecting modern networks.However,many machine learning and deep learning based NIDS rely on large volumes of labeled attack data,which is often impractical to obtain for newly emerging or rare attacks.This paper presents a benchmark-style systematic evaluation of meta-learning-based FewShot Learning(FSL)classifiers for detecting previously unseen intrusions with limited labeled data.We investigate three representative FSL models,namely Prototypical Networks,Relation Networks,and MetaOptNet,and further examine two decision-level ensemble strategies based on majority voting and probability averaging.Experiments are conducted on the UQ-IoT-IDS-2021 dataset under four attack distributions and five K-shot settings(K∈{3,5,7,10,15}),with performance reported separately for seen and unseen attacks using class-balanced macro-averaged F1 score.The experimental results show that MetaOptNet consistently provides the strongest and most stable performance on unseen attacks.Furthermore,investigating decision-level ensembling reveals that probability averaging frequently matches or marginally outperforms the best base model by successfully mitigating weaker predictions,whereas majority voting demonstrates slight performance degradation due to strict consensus requirements.We also highlight prominent error modes,such as benign-attack ambiguity and cross-device attack correlations,alongside an analysis of model complexity and inference latency.These findings highlight the potential of few-shot learning for data-scarce intrusion detection and provide insights into model selection and architectural trade-offs under varying support set sizes and attack compositions.
摘要Purpose-Rail freight is widely recognized for its economic and environmental advantages,yet it remains weakly integrated into firms’supply chains,particularly in emerging economies.This study aims to investigate the conditions under which rail freight can be effectively integrated into multi-actor supply chains,with specific attention to the role of organizational coordination,information quality and artificial intelligence(AI)in shaping logistics integration outcomes.Design/methodology/approach-The study draws on a quantitative survey of 3,185 stakeholders involved in rail-based and multimodal supply chains in Morocco.The data are analyzed using a combination of machine learning,deep learning and artificial neural network models.These methods are used not only to identify the main determinants of rail freight integration but also to capture non-linear relationships,interaction effects and potential integration trajectories that cannot be addressed through conventional linear models.Findings-The results show that rail freight integration depends primarily on organizational and informational mechanisms rather than on infrastructure alone.Inter-organizational coordination and logistics information quality emerge as the most influential factors.AI contributes positively to rail freight integration,but its effect is conditional:AI tools significantly enhance integration only when adequate levels of coordination and information sharing are already in place.Scenario simulations further reveal that the strongest integration gains arise from the combined improvement of organizational practices and AI adoption.Originality/value-This research contributes to the literature by shifting the focus from infrastructure-centered explanations toward a systemic understanding of rail freight integration.It is among the first studies to empirically combine machine learning,deep learning and artificial neural networks to analyze logistics integration in an emerging-economy context and to show that AI functions as a complementary and amplifying mechanism rather than a standalone solution.
基金supported by the National Key R&D Program(No.2021YFB3501002)supported by the Ministry of Science and Technology of China,National Natural Science Foundation of China(No.51825101,52127801).
摘要Surface properties of crystals are critical in many fields,including electrochemistry and photoelectronics,the efficient prediction of which can expedite the design and optimization of catalysts,batteries,alloys etc.However,we are still far from realizing this vision due to the rarity of surface property-related databases,especially for multicomponent compounds,due to the large sample spaces and limited computing resources.In this work,we present a surface emphasized multi-task crystal graph convolutional neural network(SEM-CGCNN)to predict multiple surface properties simultaneously from crystal structures.The model is evaluated on a dataset of 3526 surface energies and work functions of binary magnesium intermetallics obtained through first-principles calculations,and obvious improvements are observed both in efficiency and accuracy over the original CGCNN model.By transferring the pre-trained model to the datasets of pure metals and other intermetallics,the fine-tuned SEM-CGCNN outperforms learning from scratch and can be further applied to other surface properties and materials systems.This study could be a paradigm for the end-to-end mapping of atomic structures to anisotropic surface properties of crystals,which provides an efficient framework to understand and screen materials with desired surface characteristics.
摘要The increasing adoption of 5G cellular networks has introduced significant challenges for network operators.The main challenge lies in the management of seamless handoff(HO),which occurs owing to the rapid expansion of equipment,data,and network complexity.To address this challenge,a model named optimal HO management deep learning neural network(OHMDLNN)is proposed.The model is trained on network activity data,and it uses KPIs(key performance indicators)and system-level parameters to make HO decisions.As demonstrated in the article,OHMDLNN is successful in analyzing the effect and interdependence of KPIs from both the network and user equipment(UE)perspectives.Moreover,the model is evaluated for accuracy(the percentage of correct decisions made by a model on a dataset)in comparison with existing neural network-based HO decision models.These include temporal convolution networks(TCN),recurrent neural networks(RNN),long short-term memory(LSTM),gated recurrent units(GRU),and convolutional neural networks(CNN).The dataset used to evaluate the performance of the model consisted of 65,000 records.The model demonstrates superior performance,with an average improvement on accuracy of 8 percent over TCN,18 percent over RNN,6 percent over LSTM,14 percent over GRU and 4 percent over CNN.Along with accuracy,the model is also tested on important performance indicators,including the packet loss rate,the success rate,latency,and throughput at the time of handover.These results affirm its efficiency in the HO decision-making process.Future research will consider the use of advanced deep learning architectures and simplify the process of integrating system-level inputs to optimize system performance during HO events.
基金supported by the National Natural Science Foundation of China(Grant NO.62176046)Noncommunicable Chronic Diseases-National Science and Technology Major Project(2023ZD0501806)。
摘要Communication infrastructure is often among the first casualties in natural or human-induced disasters,severely impairing the coordination and efficiency of rescue operations.Rapid deployment of Unmanned Aerial Vehicles(UAVs)and satellite systems has thus become essential for establishing robust communication links to support rescue-critical tasks.However,existing emergency communication networks rely heavily on domain expertise for topology design,thereby suffering from issues such as inefficient resource allocation and network congestion,among others.To address these challenges,we present TopoLLM,a framework that leverages Large Language Models(LLMs)for tool-driven optimization of emergency network topologies.This framework effectively combines the reasoning capabilities of the LLM with TopoTool,a domain-specific optimization toolkit engineered for high-precision and load-balanced network planning in disaster scenarios.Guided by an adaptive toolselection mechanism,TopoLLM autonomously generates resilient topologies and allocates resources intelligently,reducing the need for extensive human interventions.Experimental evaluations on simulated disaster scenarios verify that TopoLLM can rapidly generate high-accuracy and robust topologies,achieving notable performance improvements compared with existing approaches.
摘要Dear Editor,This letter proposes a fully distributed multi-agent reinforcement learning(DMARL)algorithm for the coordinated optimization and scheduling of source-load-storage in networked microgrids.To accommodate the rapid development of networked microgrids,we have designed a DMARL algorithm with an event-triggered mechanism(ETM).Unlike centralized approaches,DMARL empowers individual agents to cooperatively learn and optimize based only on local observations,while reducing the communication burden.Simulation results validate the performance of the proposed algorithm,demonstrating its effectiveness in efficient microgrid resource management.
基金supported by the National Natural Science Foundation of China(Grant Nos.12574220 and 52031016)。
摘要Fe-based amorphous alloys are promising soft magnetic materials for developing next-generation devices with high frequency and efficiency.However,optimizing Fe-based alloys with ultra-high saturation magnetic flux density(Bs),ultralow coercivity(Hc),and good glass-forming ability remains a notorious challenge owing to the vast composition space and complex trade-offs among these properties.Thus,conventional design methods face great challenges.Here,we develop a generative multi-task deep learning(GMTDL)approach to achieve simultaneous optimization of compositions and tradeoff properties.The GMTDL can sufficiently exploit and share knowledge from datasets across different tasks,despite the limitations and imbalances of these datasets.Therefore,it exhibits superior performance in predicting alloys with multiple targeted properties,outperforming previous machine learning-based design strategies.Moreover,the GMTDL can also tailor compositions,providing an efficient way to regulate properties and generate desired candidates for further experimental processing.The validity and reliability of GMTDL are rigorously tested by benchmarking against Fe-based alloys reported very recently.Moreover,some new alloys with ultra-high Bs and ultra-low Hc are predicted.The optimal content windows of key elements and their synergistic effects are also unraveled,providing practical guidance.Thus,our study establishes an effective and reliable paradigm for simultaneous prediction and optimization of high-performance materials with multiple properties.
基金supported by the Research year project of the KongjuNational University in 2025 and the Institute of Information&Communications Technology Planning&Evaluation(IITP)grant funded by the Korea government(MSIT)(No.RS-2024-00444170,Research and International Collaboration on Trust Model-Based Intelligent Incident Response Technologies in 6G Open Network Environment).
摘要With the growing complexity and decentralization of network systems,the attack surface has expanded,which has led to greater concerns over network threats.In this context,artificial intelligence(AI)-based network intrusion detection systems(NIDS)have been extensively studied,and recent efforts have shifted toward integrating distributed learning to enable intelligent and scalable detection mechanisms.However,most existing works focus on individual distributed learning frameworks,and there is a lack of systematic evaluations that compare different algorithms under consistent conditions.In this paper,we present a comprehensive evaluation of representative distributed learning frameworks—Federated Learning(FL),Split Learning(SL),hybrid collaborative learning(SFL),and fully distributed learning—in the context of AI-driven NIDS.Using recent benchmark intrusion detection datasets,a unified model backbone,and controlled distributed scenarios,we assess these frameworks across multiple criteria,including detection performance,communication cost,computational efficiency,and convergence behavior.Our findings highlight distinct trade-offs among the distributed learning frameworks,demonstrating that the optimal choice depends strongly on systemconstraints such as bandwidth availability,node resources,and data distribution.This work provides the first holistic analysis of distributed learning approaches for AI-driven NIDS and offers practical guidelines for designing secure and efficient intrusion detection systems in decentralized environments.
基金funded by Ho Chi Minh City Open University(HCMCOU)the Ministry of Education and Training(Vietnam)under grant number B2025-MBS-01.
摘要Visual speech recognition(VSR)aims to infer spoken content from visual observations of articulatory movements.Despite significant progress,it remains a challenging task in computer vision and speech processing.Its difficulty arises from pronounced speaker-to-speaker variability,the presence of homophenes(phonemes that are visually indistinguishable),changes in illumination,and the intrinsically high-dimensional nature of spatiotemporal lip dynamics.In this work,we propose NestLipGNN,a graph-based framework that integrates Graph Neural Networks(GNNs)with a nested multi-granularity learning strategy for visual speech recognition.We construct dynamic lip graphs from facial landmarks to model both spatial relationships between lip regions and their temporal motion during speech articulation.The proposed nested learning architecture supports hierarchical feature extraction across several levels of linguistic abstraction,spanning phoneme-level articulatory units,viseme-level visual speech categories,and word-level semantic representations.We further introduce a Temporal Graph Attention mechanism(T-GAT)that adaptively reweights the importance of distinct lip regions over time.We also introduce a graph-based contrastive learning objective to improve the discrimination of visually similar speech patterns,directly confronting the challenge of homophene resolution.Experiments on the LRW,LRS2,LRS3,and GRID datasets show that NestLipGNN improves recognition accuracy compared with existing methods,obtaining 92.3%word-level accuracy on LRW and delivering a 2.1%absolute performance gain over prior methods.Comprehensive ablation analyses confirm the contribution of each architectural component.