期刊文献+
共找到75,478篇文章
< 1 2 250 >
每页显示 20 50 100
Anti-interference diffractive deep neural networks for multi-object recognition 认领 引用
1
作者 Zhiqi Huang Yufei Liu +11 位作者 Nan Zhang Zian Zhang Qiming Liao Cong He Shendong Liu Youhai Liu Hongtao Wang Xingdu Qiao Joel K.W.Yang Yan Zhang Lingling Huang Yongtian Wang 《Light: Science & Applications》 SCIE EI CAS CSCD 2026年第4期1172-1181,共10页
Optical neural networks(ONNs)are emerging as a promising neuromorphic computing paradigm for object recognition,offering unprecedented advantages in light-speed computation,ultra-low power consumption,and inherent par... Optical neural networks(ONNs)are emerging as a promising neuromorphic computing paradigm for object recognition,offering unprecedented advantages in light-speed computation,ultra-low power consumption,and inherent parallelism.However,most of ONNs are only capable of performing simple object classification tasks.These tasks are typically constrained to single-object scenarios,which limits their practical applications in multi-object recognition tasks.Here,we propose an anti-interference diffractive deep neural network(AI D2NN)that can accurately and robustly recognize targets in multi-object scenarios,including intra-class,inter-class,and dynamic interference.By employing different deep-learning-based training strategies for targets and interference,two transmissive diffractive layers form a physical network that maps the spatial information of targets all-optically into the power spectrum of the output light,while dispersing all interference as background noise.We demonstrate the effectiveness of this framework in classifying unknown handwritten digits under dynamic scenarios involving 40 categories of interference,achieving a simulated blind testing accuracy of 87.4%using terahertz waves.The presented framework can be physically scaled to operate at any electromagnetic wavelength by simply scaling the diffractive features in proportion to the wavelength range of interest.This work can greatly advance the practical application of ONNs in target recognition and pave the way for the development of real-time,high-throughput,low-power all-optical computing systems,which are expected to be applied to autonomous driving perception,precision medical diagnosis,and intelligent security monitoring. 展开更多
关键词 Diffractive Deep Neural Networks object recognitionoffering Multi object Recognition object classification optical neural networks onns neuromorphic computing paradigm Anti interference Light speed Computation
暂未订购 下载PDF
Self-powered Posture Recognition Based on Thermoelectric Hydrogel 认领 引用 被引量:1
2
作者 Yu-Hao Zhang Guo-Shun Cui +5 位作者 Saeed Ahmed Khan Lang-Hua Shi Hao-Zhe Zhang Kun Yang Xuan-Sen Zhao Hu-Lin Zhang 《Chinese Journal of Polymer Science》 SCIE EI CAS CSCD 2026年第6期1780-1789,I0014,共10页
Posture recognition technology plays a crucial role in health monitoring,motion analysis,and other related fields.However,traditional recognition devices are limited by a lack of comfort,stability,and environmental ad... Posture recognition technology plays a crucial role in health monitoring,motion analysis,and other related fields.However,traditional recognition devices are limited by a lack of comfort,stability,and environmental adaptability,which significantly restrict their performance and range of applications.In response,this study proposes a self-powered,strain-driven,intelligent posture recognition method based on a thermoelectric hydrogel.This approach enables high-precision posture recognition through simultaneous acquisition of dynamic strain signals from multiple joints.The proposed poly(vinyl alcohol)(PVA)/starch bi-network hydrogel,synthesized in a binary H2O/glycerol solvent system,exhibited outstanding flexibility and tensile strength,allowing it to conform closely to joint movements and ensure wearing comfort.Simultaneously,the hydrogel generated stable electrical signals via the redox reaction of[Fe(CN)6]3-/4-,supporting continuous monitoring and data collection.Moreover,its self-powered nature allows effective joint strain signal acquisition,even in open environments.Using a machine learning algorithm that incorporates signal processing and analysis,the proposed method achieved a posture recognition accuracy of 96.82%.This study presents a novel technological route for posture recognition in wearable devices,demonstrating its strong application potential in real-time feedback for sports rehabilitation and performance analysis of athletes. 展开更多
关键词 Self-powered Thermoelectric hydrogel Machine learning Posture recognition
暂未订购 下载PDF
Research on the visualization method of lithology intelligent recognition based on deep learning using mine tunnel images 认领 引用 被引量:2
3
作者 Aiai Wang Shuai Cao +1 位作者 Erol Yilmaz Hui Cao 《International Journal of Minerals,Metallurgy and Materials》 SCIE EI CAS CSCD 2026年第1期141-152,共12页
An image processing and deep learning method for identifying different types of rock images was proposed.Preprocessing,such as rock image acquisition,gray scaling,Gaussian blurring,and feature dimensionality reduction... An image processing and deep learning method for identifying different types of rock images was proposed.Preprocessing,such as rock image acquisition,gray scaling,Gaussian blurring,and feature dimensionality reduction,was conducted to extract useful feature information and recognize and classify rock images using Tensor Flow-based convolutional neural network(CNN)and Py Qt5.A rock image dataset was established and separated into workouts,confirmation sets,and test sets.The framework was subsequently compiled and trained.The categorization approach was evaluated using image data from the validation and test datasets,and key metrics,such as accuracy,precision,and recall,were analyzed.Finally,the classification model conducted a probabilistic analysis of the measured data to determine the equivalent lithological type for each image.The experimental results indicated that the method combining deep learning,Tensor Flow-based CNN,and Py Qt5 to recognize and classify rock images has an accuracy rate of up to 98.8%,and can be successfully utilized for rock image recognition.The system can be extended to geological exploration,mine engineering,and other rock and mineral resource development to more efficiently and accurately recognize rock samples.Moreover,it can match them with the intelligent support design system to effectively improve the reliability and economy of the support scheme.The system can serve as a reference for supporting the design of other mining and underground space projects. 展开更多
关键词 rock picture recognition convolutional neural network intelligent support for roadways deep learning lithology determination
暂未订购 下载PDF
Hybridtube with endo-functionalized cavity for highly selective recognition:Subtle structure changes in guest lead to significant binding affinity variations 认领 引用
4
作者 Yan-Fang Wang Jia Liu +5 位作者 Wei Li Song-Meng Wang Jian Qin Cheng-Da Zhao Liu-Pan Yang Li-Li Wang 《Chinese Chemical Letters》 SCIE CAS CSCD 2026年第6期573-579,共7页
Achieving highly selective recognition of structurally similar substrates in water has been,and still remains,challenging.Herein,we report highly selective recognition of adenosine(A)and its analogs by using a hybridt... Achieving highly selective recognition of structurally similar substrates in water has been,and still remains,challenging.Herein,we report highly selective recognition of adenosine(A)and its analogs by using a hybridtube(HT)with endo-functionalized cavity.Fluorescence titration data and density-functional theory calculations reveal that the macrocycle's hydrophobic cavity and its internal hydrogen-bonding sites are crucial for attaining this high binding selectivity to A and its analogs.Decreasing the number of hydrophilic hydroxyl groups on the ribose ring while increasing the number of hydrophobic methyl groups on the purine ring can significantly enhance the hydrophobic effect between host and guest,thereby strengthening the binding affinity.Furthermore,different hydrogen bond acceptors on the guest can greatly affect host-guest binding,leading to a substantial enhancement in binding selectivity(A/dA up to 61.7-fold).Based on the high binding selectivity of HT,a substrate-selective fluorescent supramolecular tandem assay was developed for real-time and continuous monitoring of the enzyme activity of adenosine deaminase(ADA).Finally,we demonstrated the potential of this tandem assay for inhibitor screening,which holds significant implications for drug design and medical diagnostics. 展开更多
关键词 Biomimetic macrocycle Molecular recognition Selectivity Enzyme assay
暂未订购 下载PDF
Dynamic Facial Expression Recognition of Learners via Adaptive Global Attention and Differential Temporal Transformer 认领 引用
5
作者 Wei Liu Lujia Li +4 位作者 Chun Yan Yulin Zhang Xiaochun Cheng Xinyan Zhao Mingshi Liu 《CAAI Transactions on Intelligence Technology》 SCIE EI CSCD 2026年第2期514-528,共15页
Analysing learners'facial expressions during learning and exploring their learning processes and emotional changes are of great significance for assisting teachers'teaching and promoting smart education.In com... Analysing learners'facial expressions during learning and exploring their learning processes and emotional changes are of great significance for assisting teachers'teaching and promoting smart education.In complex learning environments,static facial expression recognition fails to capture the dynamic changes of learners'expressions losing the continuous features in the learning process,and its recognition effect is easily interfered with by factors such as occlusion and lighting variations during learning.To address the above issues,a network model based on adaptive global attention and temporal difference is proposed to recognise learners'dynamic expression sequences.Firstly,we have designed an Adaptive Global Attention(AGA)block,which adaptively models inter-channel relationships to dynamically enhance key channels that are highly correlated with learners'states while suppressing redundant information,thereby improving the model's feature representation capability under noisy environments.Secondly,we have designed a Differential Temporal Transformer(DTFormer)to extract differential information between consecutive frames,increasing the model's sensitivity to learners'facial expression dynamics and improving recognition performance.The two components complement each other in terms of spatial feature enhancement and temporal dynamic modelling effectively improving the model's overall capability for representing learners'dynamic facial expressions.Experiments were conducted on public datasets DFEW,FERV39k and the learner E-learning emotional state data set DAiSEE,and comparisons were made with classical methods using objective indicators.The results demonstrate that the proposed method outperforms the comparison methods in multiple performance indicators,thereby verifying its effectiveness. 展开更多
关键词 face analysis facial expression recognition spatial-temporal feature transformer
暂未订购 下载PDF
Real-Time Emotion Recognition System Using Adaptive Distillation Technique 认领 引用
6
作者 Mustaqeem Khan Ufaq Khan +3 位作者 Mamoun Awad Nazar Zaki Guiyoung Son Soonil Kwon 《Computer Modeling in Engineering & Sciences》 SCIE EI 2026年第4期1025-1042,共18页
Knowledge distillation has shown impressive results in different fields,including detection,recognition,and generation.These models are excellent at tasks such as speech recognition,but they need to be shrunk down usi... Knowledge distillation has shown impressive results in different fields,including detection,recognition,and generation.These models are excellent at tasks such as speech recognition,but they need to be shrunk down using adaptive knowledge distillation(AKD).The use of AKD can improve human-computer interactions and streamline data collection in the field of Speech Emotion Recognition(SER).This study presents a high-level approach that employs a novel adaptive knowledge distillation(AKD)with spatio-temporal transformers to acquire advanced semantic features from the input signal.This method uses an instance-by-instance correlation between the teacher and a student to determine the teacher’s importance.Additionally,this work proposes a knowledge-transfer strategy to integrate soft targets between teachers and students,aiming to provide deeper insight for the final prediction.Our light-weight model AKD is an efficient solution for edge devices and learns the synergistic information for respective tasks,as discussed in the results and analysis section.Our proposed model AKD outperforms the SOTA models of SER systems on the benchmark datasets,IEMOCAP,EmoDB,and RAVDESS,with an absolute gain of 4%-6%in overall recognition rate. 展开更多
关键词 Affective computing edge electronics emotion recognition knowledge distillation speech signal
暂未订购 下载PDF
Exploring Generation of Pronunciation Lexicon for Low-Resource Language Automatic Speech Recognition Based on Generic Phone Recognizer 认领 引用
7
作者 LI Jinpeng CHEN Xie ZHANG Weiqiang 《Journal of Shanghai Jiaotong university(Science)》 EI 2026年第2期265-272,共8页
The lexicon is an essential component in the hybrid automatic speech recognition(ASR)system.However,a high-quality lexicon requires significant efforts from the linguistic experts and is difficult to obtain,especially... The lexicon is an essential component in the hybrid automatic speech recognition(ASR)system.However,a high-quality lexicon requires significant efforts from the linguistic experts and is difficult to obtain,especially for low-resource languages.This paper addresses the problem of using a well-trained universal phone recognizer,obtained through the training of multilingual speech data and pronunciation lexicons,to generate pronunciation lexicons for low-resource languages driven by speech data.We propose a simple pipeline that utilizes this approach to generate pronunciation lexicons and apply them into ASR systems.The steps to generate the lexicon are simple and generic:applying the International Phonetic Alphabet(IPA)phone recognizer on the speech,then aligning it with the reference word sequence,followed by filtering to obtain a series of AUTO-subwords,using them to generate the AUTO-subword lexicon and the AUTO-IPA lexicon.We used the pronunciation lexicon generated for the hybrid system and for fine-tuning the pre-trained model.According to the experiment results,we are able to construct the lexicon without resourcing to linguistic experts.Furthermore,the generated lexicon is able to outperform grapheme-based lexicon and is comparable to expert lexicon. 展开更多
关键词 International Phonetic Alphabet(IPA) lexicon learning phone recognition low-resource speech recognition
暂未订购 下载PDF
MmPiFNN:A multi-mode physics-informed fuzzy neural network for passive recognition of surface ships by underwater equipment using ship radiated noise signals 认领 引用
8
作者 Feng Liu Zipeng Li +2 位作者 Kunde Yang Fuhu Chen Junru Yu 《Defence Technology(防务技术)》 SCIE EI CAS CSCD 2026年第5期243-266,共24页
Ship radiated noise(SRN)is a key acoustic cue for underwater platforms such as submarines to detect,identify,and track surface vessels in long-range sonar confrontation scenarios.Accurate classification of SRN signals... Ship radiated noise(SRN)is a key acoustic cue for underwater platforms such as submarines to detect,identify,and track surface vessels in long-range sonar confrontation scenarios.Accurate classification of SRN signals is thus critical for underwater target recognition and maritime situational awareness.However,under complex and dynamic marine environments,SRN recognition remains highly challenging due to strong background noise,sample imbalance,and limited availability of labeled data.To enhance recognition performance under these constraints,this paper proposes a novel multi-mode physics-informed fuzzy neural network(MmPiFNN)that integrates multi-mode features,fuzzy inference,and physics-based constraints.The model applies Wasserstein generative adversarial networkbased data augmentation to address class imbalance and data scarcity.It then extracts time domain,time-frequency domain,and spatial domain features in parallel,followed by a fuzzy inference mechanism that adaptively fuses multi-mode information,improving interpretability.The fused features are input into a physics-informed neural network enhanced with three physics-based constraints:classification loss,multi-mode consistency loss,and physics-informed residual loss,enabling end-to-end physically consistent learning.The experimental results demonstrate that the proposed MmPiFNN achieves a classification precision of 91.22%on the DeepShip Dataset,outperforming existing models.Moreover,it maintains stable and high recognition performance even under small sample conditions,indicating strong practical value and promising application potential. 展开更多
关键词 Ship radiated noise Physical constraint Mode fusion Fuzzy system Passive recognition
暂未订购 下载PDF
A Fine-Grained RecognitionModel based on Discriminative Region Localization and Efficient Second-Order Feature Encoding 认领 引用
9
作者 Xiaorui Zhang Yingying Wang +3 位作者 Wei Sun Shiyu Zhou Haoming Zhang Pengpai Wang 《Computers, Materials & Continua》 SCIE EI 2026年第4期946-965,共20页
Discriminative region localization and efficient feature encoding are crucial for fine-grained object recognition.However,existing data augmentation methods struggle to accurately locate discriminative regions in comp... Discriminative region localization and efficient feature encoding are crucial for fine-grained object recognition.However,existing data augmentation methods struggle to accurately locate discriminative regions in complex backgrounds,small target objects,and limited training data,leading to poor recognition.Fine-grained images exhibit“small inter-class differences,”and while second-order feature encoding enhances discrimination,it often requires dual Convolutional Neural Networks(CNN),increasing training time and complexity.This study proposes a model integrating discriminative region localization and efficient second-order feature encoding.By ranking feature map channels via a fully connected layer,it selects high-importance channels to generate an enhanced map,accurately locating discriminative regions.Cropping and erasing augmentations further refine recognition.To improve efficiency,a novel second-order feature encoding module generates an attention map from the fourth convolutional group of Residual Network 50 layers(ResNet-50)and multiplies it with features from the fifth group,producing second-order features while reducing dimensionality and training time.Experiments on Caltech-University of California,San Diego Birds-200-2011(CUB-200-2011),Stanford Car,and Fine-Grained Visual Classification of Aircraft(FGVC Aircraft)datasets show state-of-the-art accuracy of 88.9%,94.7%,and 93.3%,respectively. 展开更多
关键词 Fine-grained recognition feature encoding data augmentation second-order feature discriminative regions
暂未订购 下载PDF
Hybrid Quantum Gate Enabled CNN Framework with Optimized Features for Human-Object Detection and Recognition 认领 引用
10
作者 Nouf Abdullah Almujally Tanvir Fatima Naik Bukht +3 位作者 Shuaa S.Alharbi Asaad Algarni Ahmad Jalal Jeongmin Park 《Computers, Materials & Continua》 SCIE EI 2026年第4期2254-2271,共18页
Recognising human-object interactions(HOI)is a challenging task for traditional machine learning models,including convolutional neural networks(CNNs).Existing models show limited transferability across complex dataset... Recognising human-object interactions(HOI)is a challenging task for traditional machine learning models,including convolutional neural networks(CNNs).Existing models show limited transferability across complex datasets such as D3D-HOI and SYSU 3D HOI.The conventional architecture of CNNs restricts their ability to handle HOI scenarios with high complexity.HOI recognition requires improved feature extraction methods to overcome the current limitations in accuracy and scalability.This work proposes a Novel quantum gate-enabled hybrid CNN(QEH-CNN)for effectiveHOI recognition.Themodel enhancesCNNperformance by integrating quantumcomputing components.The framework begins with bilateral image filtering,followed bymulti-object tracking(MOT)and Felzenszwalb superpixel segmentation.A watershed algorithm refines object boundaries by cleaning merged superpixels.Feature extraction combines a histogram of oriented gradients(HOG),Global Image Statistics for Texture(GIST)descriptors,and a novel 23-joint keypoint extractionmethod using relative joint angles and joint proximitymeasures.A fuzzy optimization process refines the extracted features before feeding them into the QEH-CNNmodel.The proposed model achieves 95.06%accuracy on the 3D-D3D-HOI dataset and 97.29%on the SYSU3DHOI dataset.Theintegration of quantum computing enhances feature optimization,leading to improved accuracy and overall model efficiency. 展开更多
关键词 Pattern recognition image segmentation computer vision object detection
暂未订购 下载PDF
Memristive neural network circuit with fault tolerance for character recognition 认领 引用
11
作者 Mei Guo Jikang Liu Jingzhi Xu 《Chinese Physics B》 SCIE EI CAS CSCD 2026年第6期325-342,共18页
Memristor-based neural networks are one of the most promising approaches for the hardware implementation of artificial neural networks.In this paper,a memristor-based neural network circuit based on a one-memristor–o... Memristor-based neural networks are one of the most promising approaches for the hardware implementation of artificial neural networks.In this paper,a memristor-based neural network circuit based on a one-memristor–one-resistor(1M1R)synaptic array structure is designed for character recognition.Compared with other memristive synaptic arrays,the 1M1R structure can reduce the number of memristors used.However,memristors may malfunction due to fabrication defects and the influence of external factors,resulting in a decrease in the accuracy of the circuit's character recognition,and a suitable solution needs to be found to improve the stability and durability of the circuit.Therefore,in this paper,a fault-tolerant module with feedback adjustment capability is designed in the memristive neural network circuit that can readjust the weights of the memristors through in-situ training to solve multiple faults in the memristive neural network.The effect of fault tolerance is verified by character recognition.The experimental results show that the designed memristive neural network circuit can accurately realize character recognition,and the designed fault-tolerant circuit can well tolerate multiple faults,ensuring stable operation of the circuit under fault conditions. 展开更多
关键词 memristor neural network circuit character recognition fault-tolerant feedback
暂未订购 下载PDF
Advances in intelligent defect recognition method for oil and gas pipeline weld X-ray image 认领 引用
12
作者 Wei-Chao Qian Shao-Hua Dong +2 位作者 Meng Sun Zi-Cong Han Lin Chen 《Petroleum Science》 SCIE EI CAS CSCD 2026年第4期2136-2174,共39页
Long-distance pipelines are essential for transporting oil and gas,with the quality of welds directly affecting their safety and reliability.Weld defects can emerge during the welding process due to improper technique... Long-distance pipelines are essential for transporting oil and gas,with the quality of welds directly affecting their safety and reliability.Weld defects can emerge during the welding process due to improper techniques or environmental factors,which can result in pipeline leakage or rupture that pose a public safety and environmental risk and can lead to significant economic losses.Therefore,effective weld defect detection is crucial to ensure the safe operation of long-distance pipelines.Although traditional X-ray inspection is commonly used to detect weld defects,it is inefficient and subjective due to its dependence on manual analysis.New developments in computer vision have significantly improved the efficiency and accuracy of automated technology for defect recognition in pipeline weld Xray images.As a result,there is an urgent need for a comprehensive review to support the development of this field and provide valuable insights.This paper comprehensively evaluates the progress in the technology for the intelligent recognition of defects in pipeline weld X-ray images,focusing on preprocessing and defect detection techniques.This review explores three key directions for intelligent weld defect recognition:signal processing,feature design,and deep learning-based methods.Deep learning-based defect recognition techniques were examined in detail from five primary perspectives:dataset creation,image classification,semantic segmentation,object detection,and performance evaluation.Finally,the challenges and future development trends in the intelligent recognition of defects in pipeline weld X-ray images are discussed,emphasizing areas that require further research and innovative advancements. 展开更多
关键词 Weld defect recognition Image preprocessing Signal processing Feature design Deep learning
暂未订购 下载PDF
A machine learning-based depression recognition model integrating spiritexpression features from traditional Chinese medicine 认领 引用
13
作者 Minghui Yao Rongrong Zhu +4 位作者 Peng Qian Huilin Liu Xirong Sun Limin Gao Fufeng Li 《Digital Chinese Medicine》 CAS CSCD 2026年第1期68-79,共12页
Objective To develop a depression recognition model by integrating the spirit-expression diagnostic framework of traditional Chinese medicine(TCM)with machine learning algorithms.The proposed model seeks to establish ... Objective To develop a depression recognition model by integrating the spirit-expression diagnostic framework of traditional Chinese medicine(TCM)with machine learning algorithms.The proposed model seeks to establish a TCM-informed tool for early depression screening,thereby bridging traditional diagnostic principles with modern computational approaches.Methods The study included patients with depression who visited the Shanghai Pudong New Area Mental Health Center from October 1,2022 to October 1,2023,as well as students and teachers from Shanghai University of Traditional Chinese Medicine during the same period as the healthy control group.Videos of 3–10 s were captured using a Xiaomi Pad 5,and the TCM spirit and expressions were determined by TCM experts(at least 3 out of 5 experts agreed to determine the category of TCM spirit and expressions).Basic information,facial images,and interview information were collected through a portable TCM intelligent analysis and diagnosis device,and facial diagnosis features were extracted using the Open CV computer vision library technology.Statistical analysis methods such as parametric and non-parametric tests were used to analyze the baseline data,TCM spirit and expression features,and facial diagnosis feature parameters of the two groups,to compare the differences in TCM spirit and expression and facial features.Five machine learning algorithms,including extreme gradient boosting(XGBoost),decision tree(DT),Bernoulli naive Bayes(BernoulliNB),support vector machine(SVM),and k-nearest neighbor(KNN)classification,were used to construct a depression recognition model based on the fusion of TCM spirit and expression features.The performance of the model was evaluated using metrics such as accuracy,precision,and the area under the receiver operating characteristic(ROC)curve(AUC).The model results were explained using the Shapley Additive exPlanations(SHAP).Results A total of 93 depression patients and 87 healthy individuals were ultimately included in this study.There was no statistically significant difference in the baseline characteristics between the two groups(P>0.05).The differences in the characteristics of the spirit and expressions in TCM and facial features between the two groups were shown as follows.(i)Quantispirit facial analysis revealed that depression patients exhibited significantly reduced facial spirit and luminance compared with healthy controls(P<0.05),with characteristic features such as sad expressions,facial erythema,and changes in the lip color ranging from erythematous to cyanotic.(ii)Depressed patients exhibited significantly lower values in facial complexion L,lip L,and a values,and gloss index,but higher values in facial complexion a and b,lip b,low gloss index,and matte index(all P<0.05).(iii)The results of multiple models show that the XGBoost-based depression recognition model,integrating the TCM“spirit-expression”diagnostic framework,achieved an accuracy of 98.61%and significantly outperformed four benchmark algorithms—DT,BernoulliNB,SVM,and KNN(P<0.01).(iv)The SHAP visualization results show that in the recognition model constructed by the XGBoost algorithm,the complexion b value,categories of facial spirit,high gloss index,low gloss index,categories of facial expression and texture features have significant contribution to the model.Conclusion This study demonstrates that integrating TCM spirit-expression diagnostic features with machine learning enables the construction of a high-precision depression detection model,offering a novel paradigm for objective depression diagnosis. 展开更多
关键词 Traditional Chinese medicine Spirit Expression Feature fusion Depression Recognition model
暂未订购 下载PDF
Improving Person Recognition for Single-Person-in-Photos:Intimacy in Photo Collections 认领 引用
14
作者 Xiaoyi Duan Tianqi Zou +2 位作者 Chenyang Wang Yu Gu Xiuying Li 《Computers, Materials & Continua》 SCIE EI 2026年第2期2089-2112,共24页
Person recognition in photo collections is a critical yet challenging task in computer vision.Previous studies have used social relationships within photo collections to address this issue.However,these methods often ... Person recognition in photo collections is a critical yet challenging task in computer vision.Previous studies have used social relationships within photo collections to address this issue.However,these methods often fail when performing single-person-in-photos recognition in photo collections,as they cannot rely on social connections for recognition.In this work,we discard social relationships and instead measure the relationships between photos to solve this problem.We designed a new model that includes a multi-parameter attention network for adaptively fusing visual features and a unified formula for measuring photo intimacy.This model effectively recognizes individuals in single photo within the collection.Due to outdated annotations and missing photos in the existing PIPA(Person in Photo Album)dataset,wemanually re-annotated it and added approximately ten thousand photos of Asian individuals to address the underrepresentation issue.Our results on the re-annotated PIPA dataset are superior to previous studies in most cases,and experiments on the supplemented dataset further demonstrate the effectiveness of our method.We have made the PIPA dataset publicly available on Zenodo,with the DOI:10.5281/zenodo.12508096(accessed on 15 October 2025). 展开更多
关键词 Deep learning computer vision person recognition photo intimacy PIPA dataset
暂未订购 下载PDF
GaitMAFF:Adaptive Multi-Modal Fusion of Skeleton Maps and Silhouettes for Robust Gait Recognition in Complex Scenarios 认领 引用
15
作者 Zhongbin Luo Zhaoyang Guan +2 位作者 Wenxing You Yunteng Wang Yanqiu Bi 《Computers, Materials & Continua》 SCIE EI 2026年第5期540-558,共19页
Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combini... Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combining silhouette and skeleton data is a promising direction,effectively fusing these heterogeneous modalities and adaptively weighting their contributions in response to diverse conditions remains a central problem.This paper introduces GaitMAFF,a novelMulti-modal Adaptive Feature Fusion Network,to address this challenge.Our approach first transforms discrete skeleton joints into a dense SkeletonMap representation to align with silhouettes,then employs an attention-based module to dynamically learn the fusion weights between the two modalities.These fused features are processed by a powerful spatio-temporal backbone withWeighted Global-Local Feature FusionModules(WFFM)to learn a discriminative representation.Extensive experiments on the challenging CCPG and Gait3D datasets show that GaitMAFF achieves state-of-the-art performance,with an average Rank-1 accuracy of 84.6%on CCPG and 58.7%on Gait3D.These results demonstrate that our adaptive fusion strategy effectively integrates complementary multimodal information,significantly enhancing gait recognition robustness and accuracy in complex scenes and providing a practical solution for real-world applications. 展开更多
关键词 Gait recognition multi-modal fusion adaptive feature fusion skeleton map silhouette
暂未订购 下载PDF
RSG-Conformer:ReLU-Based Sparse and Grouped Conformer for Audio-Visual Speech Recognition 认领 引用
16
作者 Yewei Xiao Xin Du Wei Zeng 《Computers, Materials & Continua》 SCIE EI 2026年第3期1325-1348,共24页
Audio-visual speech recognition(AVSR),which integrates audio and visual modalities to improve recognition performance and robustness in noisy or adverse acoustic conditions,has attracted significant research interest.... Audio-visual speech recognition(AVSR),which integrates audio and visual modalities to improve recognition performance and robustness in noisy or adverse acoustic conditions,has attracted significant research interest.However,Conformer-based architectures remain computational expensive due to the quadratic increase in the spatial and temporal complexity of their softmax-based attention mechanisms with sequence length.In addition,Conformerbased architectures may not provide sufficient flexibility for modeling local dependencies at different granularities.To mitigate these limitations,this study introduces a novel AVSR framework based on a ReLU-based Sparse and Grouped Conformer(RSG-Conformer)architecture.Specifically,we propose a Global-enhanced Sparse Attention(GSA)module incorporating an efficient context restoration block to recover lost contextual cues.Concurrently,a Grouped-scale Convolution(GSC)module replaces the standard Conformer convolution module,providing adaptive local modeling across varying temporal resolutions.Furthermore,we integrate a Refined Intermediate Contextual CTC(RIC-CTC)supervision strategy.This approach applies progressively increasing loss weights combined with convolution-based context aggregation,thereby further relaxing the constraint of conditional independence inherent in standard CTC frameworks.Evaluations on the LRS2 and LRS3 benchmark validate the efficacy of our approach,with word error rates(WERs)reduced to 1.8%and 1.5%,respectively.These results further demonstrate and validate its state-of-the-art performance in AVSR tasks. 展开更多
关键词 Audio-visual speech recognition conformer CTC sparse attention
暂未订购 下载PDF
Efficient Video Emotion Recognition via Multi-Scale Region-Aware Convolution and Temporal Interaction Sampling 认领 引用
17
作者 Xiaorui Zhang Chunlin Yuan +1 位作者 Wei Sun Ting Wang 《Computers, Materials & Continua》 SCIE EI 2026年第2期2036-2054,共19页
Video emotion recognition is widely used due to its alignment with the temporal characteristics of human emotional expression,but existingmodels have significant shortcomings.On the one hand,Transformermultihead self-... Video emotion recognition is widely used due to its alignment with the temporal characteristics of human emotional expression,but existingmodels have significant shortcomings.On the one hand,Transformermultihead self-attention modeling of global temporal dependency has problems of high computational overhead and feature similarity.On the other hand,fixed-size convolution kernels are often used,which have weak perception ability for emotional regions of different scales.Therefore,this paper proposes a video emotion recognition model that combines multi-scale region-aware convolution with temporal interactive sampling.In terms of space,multi-branch large-kernel stripe convolution is used to perceive emotional region features at different scales,and attention weights are generated for each scale feature.In terms of time,multi-layer odd-even down-sampling is performed on the time series,and oddeven sub-sequence interaction is performed to solve the problem of feature similarity,while reducing computational costs due to the linear relationship between sampling and convolution overhead.This paper was tested on CMU-MOSI,CMU-MOSEI,and Hume Reaction.The Acc-2 reached 83.4%,85.2%,and 81.2%,respectively.The experimental results show that the model can significantly improve the accuracy of emotion recognition. 展开更多
关键词 Multi-scale region-aware convolution temporal interaction sampling video emotion recognition
暂未订购 下载PDF
A Hybrid CNN-BiLSTM Framework for Speech Emotion Recognition with TimeGAN-Augmented Data and Contrastive Learning 认领 引用
18
作者 Rashid Jahangir Muhammad Asif Nauman +1 位作者 Oumaima Saidani Faisal Ramzan 《Computers, Materials & Continua》 SCIE EI 2026年第9期775-794,共20页
Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains chal... Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains challenging due to the emotional ambiguity,lack of labeled data,class imbalance,and speaker variability.This study presents an effective SER framework that integrates contrastive representation learning,optimized spectrogram-based data augmentation,and selective synthetic data generation by using TimeGAN to enhance emotion classification performance.Contrastive learning enables the model to better discriminate acoustically similar emotions while Optuna automatically tunes augmentation strategies such as noise injection,time shifting,and time-frequency masking.Unlike existing approaches that apply synthetic generation uniformly across all classes,the proposed method targets only confusing or under-represented emotion classes to preserve the inter-class separability.A CNN-BiLSTM architecture is used to extract spectral and temporal information of the speech.The framework is evaluated with benchmark SER datasets—EMO-DB and RAVDESS—under speaker independent protocols.Experimental results demonstrate improved accuracy,robustness,and generalization under limited and imbalanced data conditions,supported by confusion matrices,UMAP,and t-SNE visualizations. 展开更多
关键词 Speech emotion recognition data augmentation optuna TimeGAN synthetic data contrastive learning
暂未订购 下载PDF
Recognition of tool wear in milling using Kolmogorov-Arnold networks and multi-sensor feature fusion 认领 引用
19
作者 Min Wan Mingwei Cao +1 位作者 Shuai Wang Xiaowei Zheng 《Chinese Journal of Mechanical Engineering》 SCIE EI CAS CSCD 2026年第3期90-106,共17页
Existing methods for tool wear recognition using online signal processing and deep learning techniques typically utilize support vector machines(SVM)and fully connected layers(FCL).These methods inherently struggle wi... Existing methods for tool wear recognition using online signal processing and deep learning techniques typically utilize support vector machines(SVM)and fully connected layers(FCL).These methods inherently struggle with capturing complex nonlinear wear patterns due to their dependence on linear transformations and fixed-weight architectures.To overcome these limitations,this study introduces a novel tool wear recognition method based on empirical wavelet transform(EWT)and Kolmogorov-Arnold Network(KAN).Through EWT,tool wear-re-lated component signals are adaptively extracted from multi-sensor data such as cutting forces,accelerations and acoustic emission signals.To further enhance feature extraction,this study designs a multi-scale convolu-tional network(MSCN)and an efficient channel attention(ECA)mechanism.The MSCN is aimed at isolating wear-sensitive features,while the ECA mechanism highlights critical wear indicators,thereby avoiding issues such as modal mixing and energy leakage.Based on these advancements,KAN has finally been adopted to construct the tool wear recognition model.Experimental results prove the feasibility of extracting tool wear component signals through EWT.Leveraging this foundation,the proposed method achieves superior recogni-tion accuracy and feasibility compared to existing approaches.The average recognition accuracy of the proposed model exceeds 99.9%,confirming its effectiveness in tool wear recognition. 展开更多
关键词 Tool wear recognition Empirical wavelet transform Deep learning Attention mechanism Kolmogorov-Arnold network
暂未订购 下载PDF
Enhanced Scene Recognition via Multi-Model Transfer Learning with Limited Labeled Data 认领 引用
20
作者 Samia Allaoua Chelloug Ahmed A.Abd El-Latif +1 位作者 Samah Al Shathri Mohamed Hammad 《Computers, Materials & Continua》 SCIE EI 2026年第5期1191-1211,共21页
Scene recognition is a critical component of computer vision,powering applications from autonomous vehicles to surveillance systems.However,its development is often constrained by a heavy reliance on large,expensively... Scene recognition is a critical component of computer vision,powering applications from autonomous vehicles to surveillance systems.However,its development is often constrained by a heavy reliance on large,expensively annotated datasets.This research presents a novel,efficient approach that leveragesmulti-model transfer learning from pre-trained deep neural networks—specifically DenseNet201 and Visual Geometry Group(VGG)—to overcome this limitation.Ourmethod significantly reduces dependency on vast labeled data while achieving high accuracy.Evaluated on the Aerial Image Dataset(AID)dataset,the model attained a validation accuracy of 93.6%with a loss of 0.35,demonstrating robust performance with minimal training data.These results underscore the viability of our approach for real-time,data-efficient scene recognition,offering a practical and cost-effective advancement for the field. 展开更多
关键词 Scene recognition transfer learning pre-trained deep models DenseNet201 VGG
暂未订购 下载PDF
上一页 1 2 250 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈