Background:In recent years,deep convolutional neural networks(CNNs)have achieved great successes in medical imaging.However,it is difficult to obtain accurate pathological information for clinical diagnosis and treatm...Background:In recent years,deep convolutional neural networks(CNNs)have achieved great successes in medical imaging.However,it is difficult to obtain accurate pathological information for clinical diagnosis and treatment by leveraging single-modality medical images.This study aims to provide an efficient multimodality whole heart segmentation method for the diagnosis of coronary heart disease.Methods:We propose SFAM-TransUnet for multimodality whole heart segmentation,a novel deep learning framework combining CNNs and transformers.Primarily,the method integrates CNNs and visual transformers(Vits)into a unified fusion framework.Specifically,the shallow feature fusion module is designed to connect MRI and CT images,thereby providing a powerful and efficient multimodality fusion backbone for semantic segmentation.Furthermore,we propose a fusion ViT(FViT)module including self-attention(SA)and adaptive mutual boost attention(Ada-MBA)to enhance contextual information within and across modalities.The Ada-MBA module assigns attention to semantic perception regions by calculating SA and cross-attention,which improves the ability to understand context from the different modalities.Extensive experiments are con-ducted on the clinical Multi-Modality Whole Heart Segmentation datasets.Results:We successfully improved the whole heart segmentation DSCs to 0.902(AA),0.920(LV-blood),0.863(LA-blood),and 0.837(LV-myo),the HDs to 9.886(AA),9.947(LV-blood),11.911(LA-blood),and 13.599(LV-myo),the PSNR values to 33.577(AA),30.091(LV-blood),32.055(LA-blood),and 29.837(LV-myo),SSMI values to 0.901(AA),0.818(LV-blood),0.765(LA-blood),and 0.743(LV-myo).This demonstrate SFAM-TransUnet outperforms various alternative methods.Conclusions:We propose SFAM-TransUnet,an efficient framework tailored for whole heart segmentation that combines CNNs and transformers.It provides a powerful multimodality fusion network to improve the performance of whole heart semantic segmentation.These results demonstrate the efficacy of SFAM-TransUnet in integrating relevant information between different modalities in multimodal tasks.展开更多
This article systematically integrates the powerful generative features of Generative Artificial Intelligence(GAI)with the synergistic features of multi-modal technology and explores a new pedagogical approach to effe...This article systematically integrates the powerful generative features of Generative Artificial Intelligence(GAI)with the synergistic features of multi-modal technology and explores a new pedagogical approach to effective language teaching under the context of a lack of active engagement and motivation,and limited accessibility and dynamism of educational materials in traditional English courses in China’s advanced education.This study proposes a novel language teaching model(a three-tier structure)based on a critical review of the usage of GAI and multi-modality in educational environments.This new language teaching model combines GAI with multi-modal technology and centers around the G-M4 cycle(Generation-Input-Interaction-Output-Monitor/Feedback).This model means empowered generative capabilities,a more dynamic and interactive learning environment,multi-modal and creative output,and effective evaluation and prompt feedback.Furthermore,critical aspects that require attention,such as data privacy and ethical responsibilities,are also illustrated.展开更多
With the increasing of the elderly population and the growing hearth care cost, the role of service robots in aiding the disabled and the elderly is becoming important. Many researchers in the world have paid much att...With the increasing of the elderly population and the growing hearth care cost, the role of service robots in aiding the disabled and the elderly is becoming important. Many researchers in the world have paid much attention to heaRthcare robots and rehabilitation robots. To get natural and harmonious communication between the user and a service robot, the information perception/feedback ability, and interaction ability for service robots become more important in many key issues.展开更多
For the analysis of spinal and disc diseases,automated tissue segmentation of the lumbar spine is vital.Due to the continuous and concentrated location of the target,the abundance of edge features,and individual diffe...For the analysis of spinal and disc diseases,automated tissue segmentation of the lumbar spine is vital.Due to the continuous and concentrated location of the target,the abundance of edge features,and individual differences,conventional automatic segmentation methods perform poorly.Since the success of deep learning in the segmentation of medical images has been shown in the past few years,it has been applied to this task in a number of ways.The multi-scale and multi-modal features of lumbar tissues,however,are rarely explored by methodologies of deep learning.Because of the inadequacies in medical images availability,it is crucial to effectively fuse various modes of data collection for model training to alleviate the problem of insufficient samples.In this paper,we propose a novel multi-modality hierarchical fusion network(MHFN)for improving lumbar spine segmentation by learning robust feature representations from multi-modality magnetic resonance images.An adaptive group fusion module(AGFM)is introduced in this paper to fuse features from various modes to extract cross-modality features that could be valuable.Furthermore,to combine features from low to high levels of cross-modality,we design a hierarchical fusion structure based on AGFM.Compared to the other feature fusion methods,AGFM is more effective based on experimental results on multi-modality MR images of the lumbar spine.To further enhance segmentation accuracy,we compare our network with baseline fusion structures.Compared to the baseline fusion structures(input-level:76.27%,layer-level:78.10%,decision-level:79.14%),our network was able to segment fractured vertebrae more accurately(85.05%).展开更多
A new coarse-to-fine strategy was proposed for nonrigid registration of computed tomography(CT) and magnetic resonance(MR) images of a liver.This hierarchical framework consisted of an affine transformation and a B-sp...A new coarse-to-fine strategy was proposed for nonrigid registration of computed tomography(CT) and magnetic resonance(MR) images of a liver.This hierarchical framework consisted of an affine transformation and a B-splines free-form deformation(FFD).The affine transformation performed a rough registration targeting the mismatch between the CT and MR images.The B-splines FFD transformation performed a finer registration by correcting local motion deformation.In the registration algorithm,the normalized mutual information(NMI) was used as similarity measure,and the limited memory Broyden-Fletcher- Goldfarb-Shannon(L-BFGS) optimization method was applied for optimization process.The algorithm was applied to the fully automated registration of liver CT and MR images in three subjects.The results demonstrate that the proposed method not only significantly improves the registration accuracy but also reduces the running time,which is effective and efficient for nonrigid registration.展开更多
In recent years, the research on identity has brought about a great importance in sociolinguistic research. However, the research on the construction of peasant identity in comic sketches or in real life is rare. The ...In recent years, the research on identity has brought about a great importance in sociolinguistic research. However, the research on the construction of peasant identity in comic sketches or in real life is rare. The topic of this research is "The Peasant Identity in ZHAO's Comic Sketches in the Perspective of Multi-modality Theory", which is for the purpose to explore the construction of ZHAO's peasant identity in his performances. This paper takes both quantitative and qualitative research methods and the data for the research are collected from ZHAO's comic sketches in the Spring Festival Gala. Based on the multi-modality theory and for the purpose to find how peasant identity is constructed in his comic sketches, the research finds that ZHAO's peasant identity in his comic sketches could be constructed through the language, body language, costumes and stage design modalities, which are revealed in detail in section "Research Results". In addition, this research also achieves the fact that people could use different kinds of semiotics to communicate with each other in real life, and benefits us with a pretty new understanding of peasant emotion and peasant identity in China. What is more, it enables us to perceive that language cannot only reflect people's life, their social activities, but also reveal people's social identity展开更多
Listening is the breakthrough for conquering English castle, it is not only the requirement of English test, but also the practical use of English knowledge and the embodiment of English comprehensive ability. Listeni...Listening is the breakthrough for conquering English castle, it is not only the requirement of English test, but also the practical use of English knowledge and the embodiment of English comprehensive ability. Listening teaching plays a crucial role in foreign language teaching. However, the effect of listening teaching is undesirable. In recent years, multi-modality theory has been focused by many researchers. In view of particularity of the listening teaching, it is urgent to apply the multi-modality theory to English listening teaching which will produce very good teaching result.展开更多
In this work, we propose a new variational model for multi-modal image registration and present an efficient numerical implementation. The model minimizes a new functional based on using reformulated normalized gradie...In this work, we propose a new variational model for multi-modal image registration and present an efficient numerical implementation. The model minimizes a new functional based on using reformulated normalized gradients of the images as the fidelity term and higher-order derivatives as the regularizer. A key feature of the model is its ability of guaranteeing a diffeomorphic transformation which is achieved by a control term motivated by the quasi-conformal map and Beltrami coefficient. The existence of the solution of this model is established. To solve the model numerically, we design a Gauss-Newton method to solve the resulting discrete optimization problem and prove its convergence;a multilevel technique is employed to speed up the initialization and avoid likely local minima of the underlying functional. Finally, numerical experiments demonstrate that this new model can deliver good performances for multi-modal image registration and simultaneously generate an accurate diffeomorphic transformation.展开更多
Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combini...Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combining silhouette and skeleton data is a promising direction,effectively fusing these heterogeneous modalities and adaptively weighting their contributions in response to diverse conditions remains a central problem.This paper introduces GaitMAFF,a novelMulti-modal Adaptive Feature Fusion Network,to address this challenge.Our approach first transforms discrete skeleton joints into a dense SkeletonMap representation to align with silhouettes,then employs an attention-based module to dynamically learn the fusion weights between the two modalities.These fused features are processed by a powerful spatio-temporal backbone withWeighted Global-Local Feature FusionModules(WFFM)to learn a discriminative representation.Extensive experiments on the challenging CCPG and Gait3D datasets show that GaitMAFF achieves state-of-the-art performance,with an average Rank-1 accuracy of 84.6%on CCPG and 58.7%on Gait3D.These results demonstrate that our adaptive fusion strategy effectively integrates complementary multimodal information,significantly enhancing gait recognition robustness and accuracy in complex scenes and providing a practical solution for real-world applications.展开更多
The flow behavior of molten steel in the thin slab mold under high casting speed conditions was investigated,with a focus on the multi-mode continuous casting and rolling mold.A steel-slag two-phase flow model was est...The flow behavior of molten steel in the thin slab mold under high casting speed conditions was investigated,with a focus on the multi-mode continuous casting and rolling mold.A steel-slag two-phase flow model was established using large eddy simulation,the volume of fluid,and magnetohydrodynamics methods through numerical simulation.The maximum flow velocity and wave height at the steel-slag interface within the mold are critical evaluation criteria for analyzing asymmetric flow under varying casting speeds and electromagnetic braking.The results indicate that the asymmetric flows within the mold do not occur synchronously.The severity of the asymmetric flow correlates with the velocity difference across the steel-slag interface.A greater biased flow prolongs the time required to revert to a steady state.When the magnetic field intensity is set to 0.24 T and the magnetic pole position is at 390 mm from the steel-slag interface,this configuration can reduce the velocity of the steel-slag interface,thereby mitigating the asymmetric flow.Additionally,it can diminish the velocity,impact depth,and impact intensity on the narrow face of the jet,thus improving the distribution of velocity and turbulent kinetic energy within the mold.This configuration prolongs the time required for the steel-slag interface to transition from a stable state to its maximum velocity and shortens the time for the interface to return to stability from an unstable state.Moreover,it ensures the positional stability of the steel-slag interface,confining its position within−3 mm.展开更多
In multi-modal emotion recognition,excessive reliance on historical context often impedes the detection of emotional shifts,while modality heterogeneity and unimodal noise limit recognition performance.Existing method...In multi-modal emotion recognition,excessive reliance on historical context often impedes the detection of emotional shifts,while modality heterogeneity and unimodal noise limit recognition performance.Existing methods struggle to dynamically adjust cross-modal complementary strength to optimize fusion quality and lack effective mechanisms to model the dynamic evolution of emotions.To address these issues,we propose a multi-level dynamic gating and emotion transfer framework for multi-modal emotion recognition.A dynamic gating mechanism is applied across unimodal encoding,cross-modal alignment,and emotion transfer modeling,substantially improving noise robustness and feature alignment.First,we construct a unimodal encoder based on gated recurrent units and feature-selection gating to suppress intra-modal noise and enhance contextual representation.Second,we design a gated-attention crossmodal encoder that dynamically calibrates the complementary contributions of visual and audio modalities to the dominant textual features and eliminates redundant information.Finally,we introduce a gated enhanced emotion transfer module that explicitly models the temporal dependence of emotional evolution in dialogues via transfer gating and optimizes continuity modeling with a comparative learning loss.Experimental results demonstrate that the proposed method outperforms state-of-the-art models on the public MELD and IEMOCAP datasets.展开更多
The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defec...The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defects pose potential threats to high-speed trains,thus necessitating timely and accurate track inspection.The majority of extant automatic inspection methods are predicated on the utilization of single visible light data,and the efficacy of the algorithmic processes is influenced by complex environments.Furthermore,due to the single information dimension,the detection accuracy of defects in similar,occluded,and small object categories is low.To address the aforementioned issues,this paper proposes a track defect detectionmethod based on dynamicmulti-modal fusion and challenging object enhanced perception.First,in light of the variances in the representation dimensions ofmultimodal information,this paper proposes a dynamic weighted multi-modal feature fusion module.The fused multi-modal features are assigned weights,and thenmultiplied with the extracted single-modal features atmultiple levels,achieving adaptive adjustment of the response degree of fusion features.Second,a novel stepwise multi-scale convolution feature aggregation module is proposed for challenging objects.The proposed method employs depth separable convolution and cross-scale aggregation operations of different receptive fields to enhance feature extraction and reuse,thereby reducing the degree of progressive loss of effective information.The experimental results demonstrate the efficacy of the proposed method in comparison to eight established methods,encompassing both single-modal and multi-modal methods,as evidenced by the extensive findings within the constructed RGBD dataset.展开更多
To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,w...To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,we proposed a hybrid framework integrating adaptive reinforcement learning(RL),multi-modal perception fusion,and enhanced pigeon flock optimization(PFO)with curiosity-driven exploration to enable robust autonomous and formation control.The framework leverages meta-learning to optimize RL policies for real-time adaptation,fuses sensor data for precise state estimation,and enhances PFO with learned leader-follower dynamics and exploration rewards to maintain cohesive formations and explore uncertain areas.For swarms of 10–30 UAVs,it achieves 34%faster convergence,61%reduced stability root mean square error(RMSE),88%fewer collisions and 85.6%–92.3%success rates in target detection and encirclement,outperforming standard multi-agent RL,pure PFO,and single-modality RL.Three-dimensional trajectory visualizations confirm cohesive formations,collision-free maneuvers,and efficient exploration in urban search-and-rescue scenarios.Innovations include meta-RL for rapid adaptation,multi-modal fusion for robust perception,and curiosity-driven PFO for scalable,decentralized control,advancing real-world multi-UAV swarm autonomy and coordination.展开更多
Metal organic framework(MOF) assembled with coordination bonds has the disadvantage of poor stability that limits its application in the field of stationary phase,while covalent organic framework(COF)assembled through...Metal organic framework(MOF) assembled with coordination bonds has the disadvantage of poor stability that limits its application in the field of stationary phase,while covalent organic framework(COF)assembled through covalent bonds exhibits excellent structural stability.It has been shown that the stationary phases prepared by combining MOF and COF can make up for the poor stability of MOF@SiO2,and the MOF/COF composites have superior chromatographic separation performance.However,the traditional methods for preparing COF/MOF based stationary phases are generally solvent thermal synthesis.In this study,a green and low-cost synthesis method was proposed for the preparation of MOF/COF@SiO2 stationary phase.Firstly,COF@SiO2 was prepared in a choline chloride/ethylene glycol based deep eutectic solvent(DES).Secondly,another acid-base tunable DES prepared by mixing p-toluenesulfonic acid(PTSA)and 2-methylimidazole in different proportions was introduced as the reaction solvent and reactant for rapid synthesis of MOF/COF@SiO2.Compared with the toxic transition metal-based MOFs selected in most previous studies,a lightweight and non-toxic S-zone metal(calcium) based MOF was employed in this study.PTSA and calcium will form the calcium/oxygen-containing organic acid framework in acidic DES,which assembles with terephthalic acid dissolved in basic DES to form MOF.The strong hydrogen bonding effect of DES can facilitate rapid assembly of Ca-MOF.The obtained Ca-MOF/COF@SiO2 can be used for multi-mode chromatography to efficiently separate multiple isomeric/hydrophilic/hydrophobic analytes.The synthesis method of Ca-MOF/COF@SiO2 is green and mild,especially the use of acid-base tunable DES promotes the rapid synthesis of non-toxic Ca-MOF/COF@silica composites,which offers an innovative approach of greenly synthesizing novel MOF/COF stationary phases and extends their applications in the field of chromatography.展开更多
This paper exploits multi-modal Physical(PHY)-layer features in terms of artificial fingerprint,In-phase/Quadrature(IQ)imbalance and Angle of Arrival(AoA)to propose a novel PHY-layer authentication framework for a Mil...This paper exploits multi-modal Physical(PHY)-layer features in terms of artificial fingerprint,In-phase/Quadrature(IQ)imbalance and Angle of Arrival(AoA)to propose a novel PHY-layer authentication framework for a Millimeter Wave(mmWave)Multiple-Input Multiple-Output(MIMO)Unmanned Aerial Vehicle(UAV)-enabled communication system.First,we resort to the AoA-based spatial fingerprint to effectively address the challenge of channel fingerprint instability induced by high-speed UAV mobility.To further enhance the low discriminability of hardware fingerprints caused by refined manufacturing techniques,artificial Gaussian noise is injected into the transmission signals to assist the receiver in better distinguishing between legitimate and illegitimate UAVs.Then,we jointly combine with inherent IQ imbalance and AoA features to design a hybrid authentication scheme and thus construct a multi-dimensional fingerprint space for a comprehensive characterization of UAV identities.To theoretically evaluate the effectiveness of the proposed authentication framework,the analytical closed-form expressions of performance metrics like false alarm and detection probabilities are also exactly derived based on the statistical signal processing technology and composite hypothesis testing.Finally,we provide large simulation results to validate the correctness and feasibility of the proposed theoretical models,and also discuss the relation between system security and communication service quality under different artificial fingerprint level.展开更多
Modern malware is increasingly employing polymorphism,packing,and metamorphism to evade traditional signature-based detection.Because of this,there is an urgency to have more reliable classification systems.Visual mal...Modern malware is increasingly employing polymorphism,packing,and metamorphism to evade traditional signature-based detection.Because of this,there is an urgency to have more reliable classification systems.Visual malware analysis,where binaries are converted into grayscale images,has demonstrated potential in revealing structural patterns of malware family classification.However,recent methods mostly rely on single-stream,lightweight Convolutional Neural Networks(CNNs).These models have a major blind spot.The visual representation textures can be heavily obscured without changing the underlying malicious code,causing severe performance drops on newer or even rare malware classes.This paper presents a Hybrid Multi-Modal Deep Learning framework to fix this vulnerability.The proposed dual-stream architecture uses image recognition via EfficientNetB0 alongside metadata analysis using 1D-convolutional byte embeddings.This paper evaluated the framework on the modern MalwareVision-2025 dataset(approximately 125000 samples)and the legacy Malimg dataset(9339 samples).On MalwareVision-2025,the model reached a weighted accuracy of 88.44%and achieved 100%benign recall on the evaluated split.The testing across both datasets shows that combining visual and structural features reduces modality collapse.This creates a much stronger system compared to using either input type on its own.In particular,the hybrid approach improves detection performance on deeply hidden contemporary threats,including the AveMariaRAT and CobaltStrike families.展开更多
Aiming at the problems of data sparsity,uneven behavior weight allocation,and insufficient timeliness modeling existing in traditional recommendation systems in the scenario of personalized fashion recommendation,this...Aiming at the problems of data sparsity,uneven behavior weight allocation,and insufficient timeliness modeling existing in traditional recommendation systems in the scenario of personalized fashion recommendation,this paper proposes a personalized recommendation method that integrates multi-behavior weights and multi-modal features.A dynamic weighted collaborative filtering algorithm is designed,which comprehensively considers the multi-dimensional behaviors of users,and introduces a time attenuation factor to construct a time-sensitive user-item scoring matrix,so as to more accurately depict the dynamic changes of user interests.A multi-modal deep fusion framework is built:ResNet-50 is used to extract commodity image features,and the pre-trained BERT model is combined to extract text features;meanwhile,the multi-head self-attention mechanism is adopted to realize semantic-level interaction and adaptive fusion of cross-modal features,thereby enhancing the expressive ability of commodity representation.Then,user preference score prediction is carried out based on the deep predictive network to generate a personalized recommendation list.Experimental results on real e-commerce datasets show that the method in this paper achieves 0.703 and 0.491 on HR@5 and NDCG@5,respectively,which is significantly superior to other baseline models.Ablation experiments further verify the effectiveness of each module including time attenuation,multi-behavior weights and multi-modal features.This study provides a more accurate,dynamic and transparently interpretable personalized recommendation solution for e-commerce platforms,and has certain theoretical value and practical significance.展开更多
This study systematically evaluates the performance of 15 conventional global single-threshold segmentation algorithms and three representative convolutional neural network(CNN)models across multi-modal geotechnical m...This study systematically evaluates the performance of 15 conventional global single-threshold segmentation algorithms and three representative convolutional neural network(CNN)models across multi-modal geotechnical material images,including computed tomography(CT)and scanning electron microscopy(SEM)data.Based on their characteristics,thresholding methods are classified into three categories:histogram-based,entropy-based,and other approaches.Four types of geotechnical material CT images and two types of SEM images were selected as the evaluation datasets,and an objective assessment criterion that does not require manual annotation was proposed.The results indicate that the performance of different thresholding algorithms varies considerably across imaging modalities:the Otsu method performs best on coal and sandstone CT images,the Liu-S method(implemented in the JHNY-DPM software)excels on sandy soil CT images,and the Yen method demonstrates strong robustness on SEM images.However,all thresholding methods fail to effectively segment granite images with uneven grayscale distributions.In contrast,deep learning models exhibit superior performance across all modalities,with U-Net achieving the highest accuracy and stability during both training and validation without noticeable overfitting,significantly outperforming Fcn and Deeplabv3.Further experiments on a combined CT-SEM dataset reveal that despite domain adaptation challenges,U-Net can consistently segment complex geotechnical structures across different imaging modalities.Overall,the analysis demonstrates that deep learning models substantially enhance the accuracy and robustness of multi-modal geotechnical image segmentation,providing guidance for algorithm selection and supporting the unified processing of multi-source imaging data toward automation and intelligent analysis in digital geotechnical research.展开更多
This paper proposes an efficient algorithm for real-time multi-modal image matching based on a lightweight feature fusion network,targeting the challenges of multi-modal image matching in multi-source data analysis.Th...This paper proposes an efficient algorithm for real-time multi-modal image matching based on a lightweight feature fusion network,targeting the challenges of multi-modal image matching in multi-source data analysis.The algorithm addresses significant multi-modal feature differences and real-time processing limitations by incorporating key technologies including reparameterization in convolutional neural networks,multi-scale image pyramids,and feature fusion modules.The matching process employs a coarse-to-fine strategy,ensuring robust performance in complex environments.Experimental results using multi-modal datasets demonstrate that the proposed algorithm achieves superior accuracy and speed,with a success rate of 98.3%and an average matching time of 30.51 ms per 500×500 image pair.These results highlight the practical value and strong generalization capability of the algorithm in real-time applications.展开更多
To address the challenges of dusty,foggy and other complex construction site environments leading to the failure of visible light imaging and difficulties in small target detection,as well as the high resource consump...To address the challenges of dusty,foggy and other complex construction site environments leading to the failure of visible light imaging and difficulties in small target detection,as well as the high resource consumption hindering model deployment,an enhanced and lightweight algorithm is proposed.This algorithm employs a hybrid architecture,integrating red green blue(RGB)(visible light)and thermal infrared(RGBT)multi-modal images through a fusion framework based on you only look once(YOLO)version 8 and Mamba-Transformer(MT).We refer to this integrated model as YOLOv8-RGBT-MT.In terms of network improvements,a frequency enhancement module is first employed to enhance visible light and infrared images.And then,a module integrating Mamba and Transformer components is designed to replace base convolutional blocks in the backbone network,thereby expanding the receptive field of the model and improving feature extraction in complex backgrounds.Finally,a multi-modal feature fusion mechanism is introduced,through which complementary information from visible and infrared images is effectively integrated via an adaptive weighting strategy,so that both the detection accuracy and robustness for small targets are enhanced.Experimental results demonstrate that,compared to YOLOv8-RGBT,the enhanced algorithm achieves an improvement of 18.7%in mAP50,while reducing the number of inference time by 79.7%.展开更多
基金supported by the Henan Province Science and Technology Research Project(Grant 252102311276)Henan Province Key Scientific Research Projects of Universities(Grant 25B520002)+1 种基金the Fund of the Institute of Complexity Science from Henan University of Technology(Grant CSKFJJ-2025-13)the 2023 Research Nursery Engineering Project of Henan University of Chinese Medicine(Grant MP2023-10).
摘要Background:In recent years,deep convolutional neural networks(CNNs)have achieved great successes in medical imaging.However,it is difficult to obtain accurate pathological information for clinical diagnosis and treatment by leveraging single-modality medical images.This study aims to provide an efficient multimodality whole heart segmentation method for the diagnosis of coronary heart disease.Methods:We propose SFAM-TransUnet for multimodality whole heart segmentation,a novel deep learning framework combining CNNs and transformers.Primarily,the method integrates CNNs and visual transformers(Vits)into a unified fusion framework.Specifically,the shallow feature fusion module is designed to connect MRI and CT images,thereby providing a powerful and efficient multimodality fusion backbone for semantic segmentation.Furthermore,we propose a fusion ViT(FViT)module including self-attention(SA)and adaptive mutual boost attention(Ada-MBA)to enhance contextual information within and across modalities.The Ada-MBA module assigns attention to semantic perception regions by calculating SA and cross-attention,which improves the ability to understand context from the different modalities.Extensive experiments are con-ducted on the clinical Multi-Modality Whole Heart Segmentation datasets.Results:We successfully improved the whole heart segmentation DSCs to 0.902(AA),0.920(LV-blood),0.863(LA-blood),and 0.837(LV-myo),the HDs to 9.886(AA),9.947(LV-blood),11.911(LA-blood),and 13.599(LV-myo),the PSNR values to 33.577(AA),30.091(LV-blood),32.055(LA-blood),and 29.837(LV-myo),SSMI values to 0.901(AA),0.818(LV-blood),0.765(LA-blood),and 0.743(LV-myo).This demonstrate SFAM-TransUnet outperforms various alternative methods.Conclusions:We propose SFAM-TransUnet,an efficient framework tailored for whole heart segmentation that combines CNNs and transformers.It provides a powerful multimodality fusion network to improve the performance of whole heart semantic segmentation.These results demonstrate the efficacy of SFAM-TransUnet in integrating relevant information between different modalities in multimodal tasks.
基金supported by the 2025 Collaborative Education Project of Industry-Academia Cooperation(Grant Number:2506184205):“Exploration of Teaching Reform Path of Advanced English Course Under the Perspective of AI Empowerment”.
摘要This article systematically integrates the powerful generative features of Generative Artificial Intelligence(GAI)with the synergistic features of multi-modal technology and explores a new pedagogical approach to effective language teaching under the context of a lack of active engagement and motivation,and limited accessibility and dynamism of educational materials in traditional English courses in China’s advanced education.This study proposes a novel language teaching model(a three-tier structure)based on a critical review of the usage of GAI and multi-modality in educational environments.This new language teaching model combines GAI with multi-modal technology and centers around the G-M4 cycle(Generation-Input-Interaction-Output-Monitor/Feedback).This model means empowered generative capabilities,a more dynamic and interactive learning environment,multi-modal and creative output,and effective evaluation and prompt feedback.Furthermore,critical aspects that require attention,such as data privacy and ethical responsibilities,are also illustrated.
摘要With the increasing of the elderly population and the growing hearth care cost, the role of service robots in aiding the disabled and the elderly is becoming important. Many researchers in the world have paid much attention to heaRthcare robots and rehabilitation robots. To get natural and harmonious communication between the user and a service robot, the information perception/feedback ability, and interaction ability for service robots become more important in many key issues.
基金supported in part by the Technology Innovation 2030 under Grant 2022ZD0211700.
摘要For the analysis of spinal and disc diseases,automated tissue segmentation of the lumbar spine is vital.Due to the continuous and concentrated location of the target,the abundance of edge features,and individual differences,conventional automatic segmentation methods perform poorly.Since the success of deep learning in the segmentation of medical images has been shown in the past few years,it has been applied to this task in a number of ways.The multi-scale and multi-modal features of lumbar tissues,however,are rarely explored by methodologies of deep learning.Because of the inadequacies in medical images availability,it is crucial to effectively fuse various modes of data collection for model training to alleviate the problem of insufficient samples.In this paper,we propose a novel multi-modality hierarchical fusion network(MHFN)for improving lumbar spine segmentation by learning robust feature representations from multi-modality magnetic resonance images.An adaptive group fusion module(AGFM)is introduced in this paper to fuse features from various modes to extract cross-modality features that could be valuable.Furthermore,to combine features from low to high levels of cross-modality,we design a hierarchical fusion structure based on AGFM.Compared to the other feature fusion methods,AGFM is more effective based on experimental results on multi-modality MR images of the lumbar spine.To further enhance segmentation accuracy,we compare our network with baseline fusion structures.Compared to the baseline fusion structures(input-level:76.27%,layer-level:78.10%,decision-level:79.14%),our network was able to segment fractured vertebrae more accurately(85.05%).
基金Project(61240010)supported by the National Natural Science Foundation of ChinaProject(20070007070)supported by Specialized Research Fund for the Doctoral Program of Higher Education of China
摘要A new coarse-to-fine strategy was proposed for nonrigid registration of computed tomography(CT) and magnetic resonance(MR) images of a liver.This hierarchical framework consisted of an affine transformation and a B-splines free-form deformation(FFD).The affine transformation performed a rough registration targeting the mismatch between the CT and MR images.The B-splines FFD transformation performed a finer registration by correcting local motion deformation.In the registration algorithm,the normalized mutual information(NMI) was used as similarity measure,and the limited memory Broyden-Fletcher- Goldfarb-Shannon(L-BFGS) optimization method was applied for optimization process.The algorithm was applied to the fully automated registration of liver CT and MR images in three subjects.The results demonstrate that the proposed method not only significantly improves the registration accuracy but also reduces the running time,which is effective and efficient for nonrigid registration.
摘要In recent years, the research on identity has brought about a great importance in sociolinguistic research. However, the research on the construction of peasant identity in comic sketches or in real life is rare. The topic of this research is "The Peasant Identity in ZHAO's Comic Sketches in the Perspective of Multi-modality Theory", which is for the purpose to explore the construction of ZHAO's peasant identity in his performances. This paper takes both quantitative and qualitative research methods and the data for the research are collected from ZHAO's comic sketches in the Spring Festival Gala. Based on the multi-modality theory and for the purpose to find how peasant identity is constructed in his comic sketches, the research finds that ZHAO's peasant identity in his comic sketches could be constructed through the language, body language, costumes and stage design modalities, which are revealed in detail in section "Research Results". In addition, this research also achieves the fact that people could use different kinds of semiotics to communicate with each other in real life, and benefits us with a pretty new understanding of peasant emotion and peasant identity in China. What is more, it enables us to perceive that language cannot only reflect people's life, their social activities, but also reveal people's social identity
摘要Listening is the breakthrough for conquering English castle, it is not only the requirement of English test, but also the practical use of English knowledge and the embodiment of English comprehensive ability. Listening teaching plays a crucial role in foreign language teaching. However, the effect of listening teaching is undesirable. In recent years, multi-modality theory has been focused by many researchers. In view of particularity of the listening teaching, it is urgent to apply the multi-modality theory to English listening teaching which will produce very good teaching result.
摘要In this work, we propose a new variational model for multi-modal image registration and present an efficient numerical implementation. The model minimizes a new functional based on using reformulated normalized gradients of the images as the fidelity term and higher-order derivatives as the regularizer. A key feature of the model is its ability of guaranteeing a diffeomorphic transformation which is achieved by a control term motivated by the quasi-conformal map and Beltrami coefficient. The existence of the solution of this model is established. To solve the model numerically, we design a Gauss-Newton method to solve the resulting discrete optimization problem and prove its convergence;a multilevel technique is employed to speed up the initialization and avoid likely local minima of the underlying functional. Finally, numerical experiments demonstrate that this new model can deliver good performances for multi-modal image registration and simultaneously generate an accurate diffeomorphic transformation.
基金funded by the Natural Science Foundation of Chongqing Municipality,grant number CSTB2022NSCQ-MSX0503.
摘要Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combining silhouette and skeleton data is a promising direction,effectively fusing these heterogeneous modalities and adaptively weighting their contributions in response to diverse conditions remains a central problem.This paper introduces GaitMAFF,a novelMulti-modal Adaptive Feature Fusion Network,to address this challenge.Our approach first transforms discrete skeleton joints into a dense SkeletonMap representation to align with silhouettes,then employs an attention-based module to dynamically learn the fusion weights between the two modalities.These fused features are processed by a powerful spatio-temporal backbone withWeighted Global-Local Feature FusionModules(WFFM)to learn a discriminative representation.Extensive experiments on the challenging CCPG and Gait3D datasets show that GaitMAFF achieves state-of-the-art performance,with an average Rank-1 accuracy of 84.6%on CCPG and 58.7%on Gait3D.These results demonstrate that our adaptive fusion strategy effectively integrates complementary multimodal information,significantly enhancing gait recognition robustness and accuracy in complex scenes and providing a practical solution for real-world applications.
基金support from the National Natural Science Foundation of China(Grant Nos.52174313 and 52304350)thank all members of the Hebei High Quality Steel Continuous Casting Engineering Technology Research Center at North China University of Science and Technology,Tangshan,China.
摘要The flow behavior of molten steel in the thin slab mold under high casting speed conditions was investigated,with a focus on the multi-mode continuous casting and rolling mold.A steel-slag two-phase flow model was established using large eddy simulation,the volume of fluid,and magnetohydrodynamics methods through numerical simulation.The maximum flow velocity and wave height at the steel-slag interface within the mold are critical evaluation criteria for analyzing asymmetric flow under varying casting speeds and electromagnetic braking.The results indicate that the asymmetric flows within the mold do not occur synchronously.The severity of the asymmetric flow correlates with the velocity difference across the steel-slag interface.A greater biased flow prolongs the time required to revert to a steady state.When the magnetic field intensity is set to 0.24 T and the magnetic pole position is at 390 mm from the steel-slag interface,this configuration can reduce the velocity of the steel-slag interface,thereby mitigating the asymmetric flow.Additionally,it can diminish the velocity,impact depth,and impact intensity on the narrow face of the jet,thus improving the distribution of velocity and turbulent kinetic energy within the mold.This configuration prolongs the time required for the steel-slag interface to transition from a stable state to its maximum velocity and shortens the time for the interface to return to stability from an unstable state.Moreover,it ensures the positional stability of the steel-slag interface,confining its position within−3 mm.
基金funded by“the Fanying Special Program of the National Natural Science Foundation of China,grant number 62341307”“the Scientific research project of Jiangxi Provincial Department of Education,grant number GJJ200839”“the Doctoral startup fund of Jiangxi University of Technology,grant number 205200100402”.
摘要In multi-modal emotion recognition,excessive reliance on historical context often impedes the detection of emotional shifts,while modality heterogeneity and unimodal noise limit recognition performance.Existing methods struggle to dynamically adjust cross-modal complementary strength to optimize fusion quality and lack effective mechanisms to model the dynamic evolution of emotions.To address these issues,we propose a multi-level dynamic gating and emotion transfer framework for multi-modal emotion recognition.A dynamic gating mechanism is applied across unimodal encoding,cross-modal alignment,and emotion transfer modeling,substantially improving noise robustness and feature alignment.First,we construct a unimodal encoder based on gated recurrent units and feature-selection gating to suppress intra-modal noise and enhance contextual representation.Second,we design a gated-attention crossmodal encoder that dynamically calibrates the complementary contributions of visual and audio modalities to the dominant textual features and eliminates redundant information.Finally,we introduce a gated enhanced emotion transfer module that explicitly models the temporal dependence of emotional evolution in dialogues via transfer gating and optimizes continuity modeling with a comparative learning loss.Experimental results demonstrate that the proposed method outperforms state-of-the-art models on the public MELD and IEMOCAP datasets.
基金funded by Beijing Natural Science Foundation,grant number L241078.
摘要The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defects pose potential threats to high-speed trains,thus necessitating timely and accurate track inspection.The majority of extant automatic inspection methods are predicated on the utilization of single visible light data,and the efficacy of the algorithmic processes is influenced by complex environments.Furthermore,due to the single information dimension,the detection accuracy of defects in similar,occluded,and small object categories is low.To address the aforementioned issues,this paper proposes a track defect detectionmethod based on dynamicmulti-modal fusion and challenging object enhanced perception.First,in light of the variances in the representation dimensions ofmultimodal information,this paper proposes a dynamic weighted multi-modal feature fusion module.The fused multi-modal features are assigned weights,and thenmultiplied with the extracted single-modal features atmultiple levels,achieving adaptive adjustment of the response degree of fusion features.Second,a novel stepwise multi-scale convolution feature aggregation module is proposed for challenging objects.The proposed method employs depth separable convolution and cross-scale aggregation operations of different receptive fields to enhance feature extraction and reuse,thereby reducing the degree of progressive loss of effective information.The experimental results demonstrate the efficacy of the proposed method in comparison to eight established methods,encompassing both single-modal and multi-modal methods,as evidenced by the extensive findings within the constructed RGBD dataset.
基金supported by the National Natural Science Foundation of China(No.62350048)。
摘要To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,we proposed a hybrid framework integrating adaptive reinforcement learning(RL),multi-modal perception fusion,and enhanced pigeon flock optimization(PFO)with curiosity-driven exploration to enable robust autonomous and formation control.The framework leverages meta-learning to optimize RL policies for real-time adaptation,fuses sensor data for precise state estimation,and enhances PFO with learned leader-follower dynamics and exploration rewards to maintain cohesive formations and explore uncertain areas.For swarms of 10–30 UAVs,it achieves 34%faster convergence,61%reduced stability root mean square error(RMSE),88%fewer collisions and 85.6%–92.3%success rates in target detection and encirclement,outperforming standard multi-agent RL,pure PFO,and single-modality RL.Three-dimensional trajectory visualizations confirm cohesive formations,collision-free maneuvers,and efficient exploration in urban search-and-rescue scenarios.Innovations include meta-RL for rapid adaptation,multi-modal fusion for robust perception,and curiosity-driven PFO for scalable,decentralized control,advancing real-world multi-UAV swarm autonomy and coordination.
基金supported by National Natural Science Foundation of China (Nos.21906124,32302202)Natural Science Foundation of Hubei Province (No.2017CFB220)Natural Science Foundation of Shandong Province (No.ZR2023MH278)。
摘要Metal organic framework(MOF) assembled with coordination bonds has the disadvantage of poor stability that limits its application in the field of stationary phase,while covalent organic framework(COF)assembled through covalent bonds exhibits excellent structural stability.It has been shown that the stationary phases prepared by combining MOF and COF can make up for the poor stability of MOF@SiO2,and the MOF/COF composites have superior chromatographic separation performance.However,the traditional methods for preparing COF/MOF based stationary phases are generally solvent thermal synthesis.In this study,a green and low-cost synthesis method was proposed for the preparation of MOF/COF@SiO2 stationary phase.Firstly,COF@SiO2 was prepared in a choline chloride/ethylene glycol based deep eutectic solvent(DES).Secondly,another acid-base tunable DES prepared by mixing p-toluenesulfonic acid(PTSA)and 2-methylimidazole in different proportions was introduced as the reaction solvent and reactant for rapid synthesis of MOF/COF@SiO2.Compared with the toxic transition metal-based MOFs selected in most previous studies,a lightweight and non-toxic S-zone metal(calcium) based MOF was employed in this study.PTSA and calcium will form the calcium/oxygen-containing organic acid framework in acidic DES,which assembles with terephthalic acid dissolved in basic DES to form MOF.The strong hydrogen bonding effect of DES can facilitate rapid assembly of Ca-MOF.The obtained Ca-MOF/COF@SiO2 can be used for multi-mode chromatography to efficiently separate multiple isomeric/hydrophilic/hydrophobic analytes.The synthesis method of Ca-MOF/COF@SiO2 is green and mild,especially the use of acid-base tunable DES promotes the rapid synthesis of non-toxic Ca-MOF/COF@silica composites,which offers an innovative approach of greenly synthesizing novel MOF/COF stationary phases and extends their applications in the field of chromatography.
基金supported in part by the National Key R&D Program of China under Grant 2023YFB3107500in part by the National Natural Science Foundation of China under Grant 62272241+1 种基金in part by the State Key Laboratory of Integrated Services Networks(Xidian University),under Grant ISN24-18in part by the Nanjing University of Posts and Telecommunications Scientific Research Foundation under Grant NY221122。
摘要This paper exploits multi-modal Physical(PHY)-layer features in terms of artificial fingerprint,In-phase/Quadrature(IQ)imbalance and Angle of Arrival(AoA)to propose a novel PHY-layer authentication framework for a Millimeter Wave(mmWave)Multiple-Input Multiple-Output(MIMO)Unmanned Aerial Vehicle(UAV)-enabled communication system.First,we resort to the AoA-based spatial fingerprint to effectively address the challenge of channel fingerprint instability induced by high-speed UAV mobility.To further enhance the low discriminability of hardware fingerprints caused by refined manufacturing techniques,artificial Gaussian noise is injected into the transmission signals to assist the receiver in better distinguishing between legitimate and illegitimate UAVs.Then,we jointly combine with inherent IQ imbalance and AoA features to design a hybrid authentication scheme and thus construct a multi-dimensional fingerprint space for a comprehensive characterization of UAV identities.To theoretically evaluate the effectiveness of the proposed authentication framework,the analytical closed-form expressions of performance metrics like false alarm and detection probabilities are also exactly derived based on the statistical signal processing technology and composite hypothesis testing.Finally,we provide large simulation results to validate the correctness and feasibility of the proposed theoretical models,and also discuss the relation between system security and communication service quality under different artificial fingerprint level.
摘要Modern malware is increasingly employing polymorphism,packing,and metamorphism to evade traditional signature-based detection.Because of this,there is an urgency to have more reliable classification systems.Visual malware analysis,where binaries are converted into grayscale images,has demonstrated potential in revealing structural patterns of malware family classification.However,recent methods mostly rely on single-stream,lightweight Convolutional Neural Networks(CNNs).These models have a major blind spot.The visual representation textures can be heavily obscured without changing the underlying malicious code,causing severe performance drops on newer or even rare malware classes.This paper presents a Hybrid Multi-Modal Deep Learning framework to fix this vulnerability.The proposed dual-stream architecture uses image recognition via EfficientNetB0 alongside metadata analysis using 1D-convolutional byte embeddings.This paper evaluated the framework on the modern MalwareVision-2025 dataset(approximately 125000 samples)and the legacy Malimg dataset(9339 samples).On MalwareVision-2025,the model reached a weighted accuracy of 88.44%and achieved 100%benign recall on the evaluated split.The testing across both datasets shows that combining visual and structural features reduces modality collapse.This creates a much stronger system compared to using either input type on its own.In particular,the hybrid approach improves detection performance on deeply hidden contemporary threats,including the AveMariaRAT and CobaltStrike families.
基金funded by the Shandong University of Technology Science and Technology Doctor Startup Fund(Project:UsingThe Real-Time Services of Variable Rate and Pauseable Non-Real-Time ServicesGrant number:4041/422022)+2 种基金the Research Fund(Project:Patent Assignment for User Identification System of a Neural Network-Based VR DeviceGrant number:9101/22502551)the Research Fund(Project:Patent Rights Transfer for the Generating 3D Facial Images with a Deep Learning Method,Grant number:9101/22502561).
摘要Aiming at the problems of data sparsity,uneven behavior weight allocation,and insufficient timeliness modeling existing in traditional recommendation systems in the scenario of personalized fashion recommendation,this paper proposes a personalized recommendation method that integrates multi-behavior weights and multi-modal features.A dynamic weighted collaborative filtering algorithm is designed,which comprehensively considers the multi-dimensional behaviors of users,and introduces a time attenuation factor to construct a time-sensitive user-item scoring matrix,so as to more accurately depict the dynamic changes of user interests.A multi-modal deep fusion framework is built:ResNet-50 is used to extract commodity image features,and the pre-trained BERT model is combined to extract text features;meanwhile,the multi-head self-attention mechanism is adopted to realize semantic-level interaction and adaptive fusion of cross-modal features,thereby enhancing the expressive ability of commodity representation.Then,user preference score prediction is carried out based on the deep predictive network to generate a personalized recommendation list.Experimental results on real e-commerce datasets show that the method in this paper achieves 0.703 and 0.491 on HR@5 and NDCG@5,respectively,which is significantly superior to other baseline models.Ablation experiments further verify the effectiveness of each module including time attenuation,multi-behavior weights and multi-modal features.This study provides a more accurate,dynamic and transparently interpretable personalized recommendation solution for e-commerce platforms,and has certain theoretical value and practical significance.
基金supported by the National Key Research and Development Program of China(No.2025YFE0116200)the National Natural Science Foundation of China(Nos.52474155 and W2521168)+1 种基金the Natural Science Foundation of Jiangsu Province(Grant No.BK20240107)the Scientific Research Innovation Capability Support Project for Young Faculty(SRICSPYF-ZY2025043).
摘要This study systematically evaluates the performance of 15 conventional global single-threshold segmentation algorithms and three representative convolutional neural network(CNN)models across multi-modal geotechnical material images,including computed tomography(CT)and scanning electron microscopy(SEM)data.Based on their characteristics,thresholding methods are classified into three categories:histogram-based,entropy-based,and other approaches.Four types of geotechnical material CT images and two types of SEM images were selected as the evaluation datasets,and an objective assessment criterion that does not require manual annotation was proposed.The results indicate that the performance of different thresholding algorithms varies considerably across imaging modalities:the Otsu method performs best on coal and sandstone CT images,the Liu-S method(implemented in the JHNY-DPM software)excels on sandy soil CT images,and the Yen method demonstrates strong robustness on SEM images.However,all thresholding methods fail to effectively segment granite images with uneven grayscale distributions.In contrast,deep learning models exhibit superior performance across all modalities,with U-Net achieving the highest accuracy and stability during both training and validation without noticeable overfitting,significantly outperforming Fcn and Deeplabv3.Further experiments on a combined CT-SEM dataset reveal that despite domain adaptation challenges,U-Net can consistently segment complex geotechnical structures across different imaging modalities.Overall,the analysis demonstrates that deep learning models substantially enhance the accuracy and robustness of multi-modal geotechnical image segmentation,providing guidance for algorithm selection and supporting the unified processing of multi-source imaging data toward automation and intelligent analysis in digital geotechnical research.
基金supported by National Natural Science Foundation of China Projects of International Cooperation and Exchanges(No.W2411055)。
摘要This paper proposes an efficient algorithm for real-time multi-modal image matching based on a lightweight feature fusion network,targeting the challenges of multi-modal image matching in multi-source data analysis.The algorithm addresses significant multi-modal feature differences and real-time processing limitations by incorporating key technologies including reparameterization in convolutional neural networks,multi-scale image pyramids,and feature fusion modules.The matching process employs a coarse-to-fine strategy,ensuring robust performance in complex environments.Experimental results using multi-modal datasets demonstrate that the proposed algorithm achieves superior accuracy and speed,with a success rate of 98.3%and an average matching time of 30.51 ms per 500×500 image pair.These results highlight the practical value and strong generalization capability of the algorithm in real-time applications.
摘要To address the challenges of dusty,foggy and other complex construction site environments leading to the failure of visible light imaging and difficulties in small target detection,as well as the high resource consumption hindering model deployment,an enhanced and lightweight algorithm is proposed.This algorithm employs a hybrid architecture,integrating red green blue(RGB)(visible light)and thermal infrared(RGBT)multi-modal images through a fusion framework based on you only look once(YOLO)version 8 and Mamba-Transformer(MT).We refer to this integrated model as YOLOv8-RGBT-MT.In terms of network improvements,a frequency enhancement module is first employed to enhance visible light and infrared images.And then,a module integrating Mamba and Transformer components is designed to replace base convolutional blocks in the backbone network,thereby expanding the receptive field of the model and improving feature extraction in complex backgrounds.Finally,a multi-modal feature fusion mechanism is introduced,through which complementary information from visible and infrared images is effectively integrated via an adaptive weighting strategy,so that both the detection accuracy and robustness for small targets are enhanced.Experimental results demonstrate that,compared to YOLOv8-RGBT,the enhanced algorithm achieves an improvement of 18.7%in mAP50,while reducing the number of inference time by 79.7%.