期刊文献+
共找到706篇文章
< 1 2 36 >
每页显示 20 50 100
Efficient Video Emotion Recognition via Multi-Scale Region-Aware Convolution and Temporal Interaction Sampling 认领 引用
1
作者 Xiaorui Zhang Chunlin Yuan +1 位作者 Wei Sun Ting Wang 《Computers, Materials & Continua》 SCIE EI 2026年第2期2036-2054,共19页
Video emotion recognition is widely used due to its alignment with the temporal characteristics of human emotional expression,but existingmodels have significant shortcomings.On the one hand,Transformermultihead self-... Video emotion recognition is widely used due to its alignment with the temporal characteristics of human emotional expression,but existingmodels have significant shortcomings.On the one hand,Transformermultihead self-attention modeling of global temporal dependency has problems of high computational overhead and feature similarity.On the other hand,fixed-size convolution kernels are often used,which have weak perception ability for emotional regions of different scales.Therefore,this paper proposes a video emotion recognition model that combines multi-scale region-aware convolution with temporal interactive sampling.In terms of space,multi-branch large-kernel stripe convolution is used to perceive emotional region features at different scales,and attention weights are generated for each scale feature.In terms of time,multi-layer odd-even down-sampling is performed on the time series,and oddeven sub-sequence interaction is performed to solve the problem of feature similarity,while reducing computational costs due to the linear relationship between sampling and convolution overhead.This paper was tested on CMU-MOSI,CMU-MOSEI,and Hume Reaction.The Acc-2 reached 83.4%,85.2%,and 81.2%,respectively.The experimental results show that the model can significantly improve the accuracy of emotion recognition. 展开更多
关键词 Multi-scale region-aware convolution temporal interaction sampling video emotion recognition
暂未订购 下载PDF
Multi-scale simplified residual convolutional neural network model for predicting compositions of binary magnesium alloys 认领 引用
2
作者 Xu Qin Qinghang Wang +6 位作者 Xinqian Zhao Shouxin Xia Li Wang Jiabao Long Yuhui Zhang Yanfu Chai Daolun Chen 《Journal of Magnesium and Alloys》 SCIE EI CAS CSCD 2026年第1期117-123,共7页
This study proposes a multi-scale simplified residual convolutional neural network(MS-SRCNN)for the precise prediction of Mg-Nd binary alloy compositions from scanning electron microscope(SEM)images.A multi-scale data... This study proposes a multi-scale simplified residual convolutional neural network(MS-SRCNN)for the precise prediction of Mg-Nd binary alloy compositions from scanning electron microscope(SEM)images.A multi-scale data structure is established by spatially aligning and stacking SEM images at different magnifications.The MS-SRCNN significantly reduces computational runtime by over 90%compared to traditional architectures like ResNet50,VGG16,and VGG19,without compromising prediction accuracy.The model demonstrates more excellent predictive performance,achieving a>5%increase in R2 compared to single-scale models.Furthermore,the MS-SRCNN exhibits robust composition prediction capability across other Mg-based binary alloys,including Mg-La,Mg-Sn,Mg-Ce,Mg-Sm,Mg-Ag,and Mg-Y,thereby emphasizing its generalization and extrapolation potential.This research establishes a non-destructive,microstructure-informed composition analysis framework,reduces characterization time compared to traditional experiment methods and provides insights into the composition-microstructure relationship in diverse material systems. 展开更多
关键词 Magnesium alloys Composition prediction Scanning electron microscope images Multi-scale simplified residual convolutional neural network
暂未订购 下载PDF
MSA-ConvNeXt:Predicting Magnetism of Doped Two-Dimensional Nanomaterials via Multi-Scale Convolution and Attention Mechanisms 认领 引用
3
作者 Yuxuan Feng Lili Liang +1 位作者 Guanglu Sun Yanrui Wei 《Computers, Materials & Continua》 SCIE EI 2026年第9期354-368,共15页
In doped two-dimensional nanomaterials,magnetism is one of the important physical properties.By introducing foreign doping atoms or molecules,the electronic structure of the material can be effectively regulated,leadi... In doped two-dimensional nanomaterials,magnetism is one of the important physical properties.By introducing foreign doping atoms or molecules,the electronic structure of the material can be effectively regulated,leading to changes in magnetic behavior.Currently,magnetic property prediction has achieved considerable results with the help of traditional CNNs,but there are still obvious limitations:(1)The feature extraction of dopant sites is constrained by fixed receptive fields,making it difficult to characterize local structural perturbations in the vicinity of dopant atoms and their spatial influence propagating to surrounding regions;(2)CNNs lack the capability to model long-range dependencies between non-neighboring atoms and their chemical bonds,thereby weakening the representation of long-range interactions within the material.In this study,we propose Multi-Scale and Attention ConvNeXt(MSA-ConvNeXt)based on multi-scale convolution and attention mechanisms,which consists of the following two core modules:(1)The Multi-scale Convolution Attention Block(MCAB),which models local structural perturbations around dopant atoms and their spatial effects via parallelmulti-scale convolutions.It uses a serial channel and spatial attention mechanism to adaptively recalibrate multi-scale features,highlighting the response of doping related regions and enhancing the ability to express dopant-site information;(2)The Visual Geometry Group–Swin Transformer(VGG-Swin)architecture extracts structural features of dopant sites using VGG convolutions to prevent the attenuation of structural information during global relationship modeling.Subsequently,the Swin Transformer is introduced,which uses the self-attention mechanism to dynamically weight and globally associate features at different spatial locations,in order to depict the long-range correlations between non-neighboring atoms and their chemical bonds with the dopant-site.Experiments conducted on a doped two-dimensional nanomaterial dataset constructed from the CMR database demonstrate that the proposed model outperforms existing methods in terms of accuracy and F1-score.Specifically,MSA-ConvNeXt achieves an accuracy of 91.66%,representing an improvement of 1.65%over the next best model.In addition,all experimental results are averaged over multiple independent runs(with five different random seeds),demonstrating the stability and reliability of themodel’s performance.Ablation studies further validate the effectiveness of each module design. 展开更多
关键词 Doped two-dimensional nanomaterials magnetic property prediction multi-scale convolution attention mechanism ConvNeXt
暂未订购 下载PDF
An Adaptive Multi-Scale Dilated Convolution Network for Real-Time Road Black Ice Detection 认领 引用
4
作者 Sun-Kyoung Kang Yeonwoo Lee 《Computers, Materials & Continua》 SCIE EI 2026年第8期1742-1768,共27页
Black ice formation on road surfaces presents a serious hazard due to its low visibility and high slipperiness,underscoring the critical need for timely and accurate detection in intelligent transportation systems.In ... Black ice formation on road surfaces presents a serious hazard due to its low visibility and high slipperiness,underscoring the critical need for timely and accurate detection in intelligent transportation systems.In this paper,we propose AdaMsDCNet,an adaptive multi-scale dilated convolution network designed for real-time black-ice semantic segmentation on resource-constrained edge platforms,applying a Convolutional Neural Network(CNN)with an adaptive Multi-Scale Dilated Convolution(MsDC)feature fusion encoder-decoder architecture.The key concept of AdaMsDCNet is to employ an encoder-decoder architecture with parallel multi-scale dilated convolutional paths that adjust dilation rates at different encoder depths using a systematic 4→2→1 progression,optimally capturing a wide range of receptive fields while mitigating checkerboard artifacts.The encoder dynamically fuses features from multiple dilation rates at each stage,enhancing segmentation accuracy.Simultaneously,the decoder uses transposed convolutions and skip connections to preserve fine spatial details.Experimental validation on a proprietary thermal infrared dataset of 1156 annotated images show that AdaMsDCNet_9 achieves 96.47%mIoU,95.48%Black-Ice IoU,97.55%Precision,97.82%Recall,and 97.69%F1-Score,outperforming U-Net(+26.78 pp mIoU,+29.88 pp Recall),DeepLabv3+(+2.82 pp mIoU),and LinkNet(+1.08 pp mIoU)while requiring only 1.86M parameters and maintaining real-time inference speeds of 3.94~5.63 FPS on the NVIDIA Jetson Nano embedded GPU.Ablation studies confirm the benefits of adaptive dilation,parallel feature fusion,and controlled channel growth for the accuracy–efficiency trade-off.Limitations including dataset generalization to uncontrolled outdoor conditions and the evaluation of imbalance-aware loss functions are identified as directions for future work. 展开更多
关键词 CNN multi-scale dilation convolution feature fusion black ice detection
暂未订购 下载PDF
MSSTGCN: Multi-Head Self-Attention and Spatial-Temporal Graph Convolutional Network for Multi-Scale Traffic Flow Prediction 认领 引用 被引量:2
5
作者 Xinlu Zong Fan Yu +1 位作者 Zhen Chen Xue Xia 《Computers, Materials & Continua》 SCIE EI 2025年第2期3517-3537,共21页
Accurate traffic flow prediction has a profound impact on modern traffic management. Traffic flow has complex spatial-temporal correlations and periodicity, which poses difficulties for precise prediction. To address ... Accurate traffic flow prediction has a profound impact on modern traffic management. Traffic flow has complex spatial-temporal correlations and periodicity, which poses difficulties for precise prediction. To address this problem, a Multi-head Self-attention and Spatial-Temporal Graph Convolutional Network (MSSTGCN) for multiscale traffic flow prediction is proposed. Firstly, to capture the hidden traffic periodicity of traffic flow, traffic flow is divided into three kinds of periods, including hourly, daily, and weekly data. Secondly, a graph attention residual layer is constructed to learn the global spatial features across regions. Local spatial-temporal dependence is captured by using a T-GCN module. Thirdly, a transformer layer is introduced to learn the long-term dependence in time. A position embedding mechanism is introduced to label position information for all traffic sequences. Thus, this multi-head self-attention mechanism can recognize the sequence order and allocate weights for different time nodes. Experimental results on four real-world datasets show that the MSSTGCN performs better than the baseline methods and can be successfully adapted to traffic prediction tasks. 展开更多
关键词 Graph convolutional network traffic flow prediction multi-scale traffic flow spatial-temporal model
暂未订购 下载PDF
Occluded Gait Emotion Recognition Based on Multi-Scale Suppression Graph Convolutional Network 认领 引用 被引量:1
6
作者 Yuxiang Zou Ning He +2 位作者 Jiwu Sun Xunrui Huang Wenhua Wang 《Computers, Materials & Continua》 SCIE EI 2025年第1期1255-1276,共22页
In recent years,gait-based emotion recognition has been widely applied in the field of computer vision.However,existing gait emotion recognition methods typically rely on complete human skeleton data,and their accurac... In recent years,gait-based emotion recognition has been widely applied in the field of computer vision.However,existing gait emotion recognition methods typically rely on complete human skeleton data,and their accuracy significantly declines when the data is occluded.To enhance the accuracy of gait emotion recognition under occlusion,this paper proposes a Multi-scale Suppression Graph ConvolutionalNetwork(MS-GCN).TheMS-GCN consists of three main components:Joint Interpolation Module(JI Moudle),Multi-scale Temporal Convolution Network(MS-TCN),and Suppression Graph Convolutional Network(SGCN).The JI Module completes the spatially occluded skeletal joints using the(K-Nearest Neighbors)KNN interpolation method.The MS-TCN employs convolutional kernels of various sizes to comprehensively capture the emotional information embedded in the gait,compensating for the temporal occlusion of gait information.The SGCN extracts more non-prominent human gait features by suppressing the extraction of key body part features,thereby reducing the negative impact of occlusion on emotion recognition results.The proposed method is evaluated on two comprehensive datasets:Emotion-Gait,containing 4227 real gaits from sources like BML,ICT-Pollick,and ELMD,and 1000 synthetic gaits generated using STEP-Gen technology,and ELMB,consisting of 3924 gaits,with 1835 labeled with emotions such as“Happy,”“Sad,”“Angry,”and“Neutral.”On the standard datasets Emotion-Gait and ELMB,the proposed method achieved accuracies of 0.900 and 0.896,respectively,attaining performance comparable to other state-ofthe-artmethods.Furthermore,on occlusion datasets,the proposedmethod significantly mitigates the performance degradation caused by occlusion compared to other methods,the accuracy is significantly higher than that of other methods. 展开更多
关键词 KNN interpolation multi-scale temporal convolution suppression graph convolutional network gait emotion recognition human skeleton
暂未订购 下载PDF
An efficient projection defocus algorithm based on multi-scale convolution kernel templates 认领 引用 被引量:1
7
作者 Bo ZHU Li-jun XIE +1 位作者 Guang-hua SONG Yao ZHENG 《Journal of Zhejiang University-Science C(Computers and Electronics)》 2013年第12期930-940,共11页
The focal problems of projection include out-of-focus projection images from the projector caused by incomplete mechanical focus and screen-door effects produced by projection pixilation. To eliminate these defects an... The focal problems of projection include out-of-focus projection images from the projector caused by incomplete mechanical focus and screen-door effects produced by projection pixilation. To eliminate these defects and enhance the imaging quality and clarity of projectors, a novel adaptive projection defocus algorithm is proposed based on multi-scale convolution kernel templates. This algorithm applies the improved Sobel-Tenengrad focus evaluation function to calculate the sharpness degree of intensity equalization and then constructs multi-scale defocus convolution kernels to remap and render the defocus projection image. The resulting projection defocus corrected images can eliminate out-of-focus effects and improve the sharpness of uncorrected images. Experiments show that the algorithm works quickly and robustly and that it not only effectively eliminates visual artifacts and can run on a self-designed smart projection system in real time but also significantly improves the resolution and clarity of the observer's visual perception. 展开更多
关键词 Projection focal Sobel-Tenengrad evaluation function Projector defocus Multi-scale convolution kernels
AMVT-NMN:Adaptive Multi-Scale Vision Transformer with Neuromorphic Memory Networks for Enhanced Lung Cancer Detection 认领 引用
8
作者 Wariyo Godana Arero Yaqin Zhao +6 位作者 Mudasir Ahmad Wani Pir Noman Ahmad Kashish Ara Shakil Sadique Ahmad Sidrak Habtemariam Teredda Merhawit Berhane Teklu Longwen Wu 《Computer Modeling in Engineering & Sciences》 SCIE EI 2026年第4期1236-1262,共27页
Lung cancer accounts for the highest number of cancer deaths globally,underscoring the urgent need for early and precise detection to enhance patient outcomes.While deep learning has made remarkable strides in analyzi... Lung cancer accounts for the highest number of cancer deaths globally,underscoring the urgent need for early and precise detection to enhance patient outcomes.While deep learning has made remarkable strides in analyzing medical images,current approaches face a fundamental challenge.They cannot adequately capture detailed local patterns and broader contextual relationships within lung Computed tomography(CT)scans.To address this limitation,we introduce AMVT-NMN(adaptive multi-scale vision transformer with neuromorphic memory networks),which combines three complementary mechanisms.The dynamic adaptive kernel networks component intelligently adjusts receptive field sizes based on input characteristics,enabling flexible feature capture across multiple scales.The neuromorphic contextual memory attention module draws inspiration from how human memory systems process information,maintaining a dynamic record of diagnostically relevant patterns to inform current predictions.The hierarchical cross-scale fusion mechanism with learnable weights synthesizes information from different resolution levels through adaptive weighting.Testing on the Iraq-Oncology Teaching Hospital/National Center for Cancer Diseases(IQOTHNCCD)dataset demonstrates strong performance:97.9%accuracy,96.5%sensitivity,98.7%specificity,and 99.2%Area under the Curve-Receiver Operating Characteristic(AUC-ROC).These results surpass existing methods such as CNN-GD,which achieved 97.2%accuracy.Notably,the high specificity translates to fewer false alarms,potentially reducing unnecessary biopsies and follow-up imaging outcomes that matter considerably in clinical practice.Result of AMVT-NMN generalization to the Lung Image Database Consortium and Image Database Resource Initiative(LIDC-IDRI),Lung nodule analysis(LUNA16),and Non-Small Cell Lung Cancer(NSCLC)-Radiomics datasets showed AUCs of 96.5%,92.8%,and 97.2%,respectively.Ablation experiments confirm that each architectural element of AMVT-NMN contributes meaningfully to overall performance.Five-fold cross-validation yielded consistent results(97.71±0.57%),indicating reliable performance across different patient subsets.The memory-augmented design shows particular promise for handling diagnostically ambiguous cases.It is focused on pattern recognition and computational intelligence,which is useful for coping with uncertain information in intelligent diagnosis systems,meeting the growing trend for trusted artificial intelligence(AI)in decision-making. 展开更多
关键词 Adaptive multi-scale vision transformer dynamic adaptive kernel networks hierarchical cross-scale fusion lung cancer detection neuromorphic memory networks
暂未订购 下载PDF
A multi-scale convolutional auto-encoder and its application in fault diagnosis of rolling bearings 认领 引用 被引量:12
9
作者 Ding Yunhao Jia Minping 《Journal of Southeast University(English Edition)》 EI CAS 2019年第4期417-423,共7页
Aiming at the difficulty of fault identification caused by manual extraction of fault features of rotating machinery,a one-dimensional multi-scale convolutional auto-encoder fault diagnosis model is proposed,based on ... Aiming at the difficulty of fault identification caused by manual extraction of fault features of rotating machinery,a one-dimensional multi-scale convolutional auto-encoder fault diagnosis model is proposed,based on the standard convolutional auto-encoder.In this model,the parallel convolutional and deconvolutional kernels of different scales are used to extract the features from the input signal and reconstruct the input signal;then the feature map extracted by multi-scale convolutional kernels is used as the input of the classifier;and finally the parameters of the whole model are fine-tuned using labeled data.Experiments on one set of simulation fault data and two sets of rolling bearing fault data are conducted to validate the proposed method.The results show that the model can achieve 99.75%,99.3%and 100%diagnostic accuracy,respectively.In addition,the diagnostic accuracy and reconstruction error of the one-dimensional multi-scale convolutional auto-encoder are compared with traditional machine learning,convolutional neural networks and a traditional convolutional auto-encoder.The final results show that the proposed model has a better recognition effect for rolling bearing fault data. 展开更多
关键词 fault diagnosis deep learning convolutional auto-encoder multi-scale convolutional kernel feature extraction
暂未订购 下载PDF
Land cover classification from remote sensing images based on multi-scale fully convolutional network 认领 引用 被引量:26
10
作者 Rui Li Shunyi Zheng +2 位作者 Chenxi Duan Libo Wang Ce Zhang 《Geo-Spatial Information Science》 SCIE EI CSCD 2022年第2期278-294,共17页
Although the Convolutional Neural Network(CNN)has shown great potential for land cover classification,the frequently used single-scale convolution kernel limits the scope of informa-tion extraction.Therefore,we propos... Although the Convolutional Neural Network(CNN)has shown great potential for land cover classification,the frequently used single-scale convolution kernel limits the scope of informa-tion extraction.Therefore,we propose a Multi-Scale Fully Convolutional Network(MSFCN)with a multi-scale convolutional kernel as well as a Channel Attention Block(CAB)and a Global Pooling Module(GPM)in this paper to exploit discriminative representations from two-dimensional(2D)satellite images.Meanwhile,to explore the ability of the proposed MSFCN for spatio-temporal images,we expand our MSFCN to three-dimension using three-dimensional(3D)CNN,capable of harnessing each land cover category’s time series interac-tion from the reshaped spatio-temporal remote sensing images.To verify the effectiveness of the proposed MSFCN,we conduct experiments on two spatial datasets and two spatio-temporal datasets.The proposed MSFCN achieves 60.366%on the WHDLD dataset and 75.127%on the GID dataset in terms of mIoU index while the figures for two spatio-temporal datasets are 87.753%and 77.156%.Extensive comparative experiments and abla-tion studies demonstrate the effectiveness of the proposed MSFCN. 展开更多
关键词 Spatio-temporal remote sensing images Multi-Scale Fully Convolutional Network land cover classification
暂未订购 下载PDF
High-Quality Single-Pixel Imaging Based on Large-Kernel Convolution under Low-Sampling Conditions 认领 引用
11
作者 Chenyu Yuan Yuanhao Su Chunfang Wang 《Chinese Physics Letters》 SCIE EI CAS CSCD 2025年第4期55-61,共7页
In recent years,deep learning has been introduced into the field of Single-pixel imaging(SPI),garnering significant attention.However,conventional networks still exhibit limitations in preserving image details.To addr... In recent years,deep learning has been introduced into the field of Single-pixel imaging(SPI),garnering significant attention.However,conventional networks still exhibit limitations in preserving image details.To address this issue,we integrate Large Kernel Convolution(LKconv)into the U-Net framework,proposing an enhanced network structure named U-LKconv network,which significantly enhances the capability to recover image details even under low sampling conditions. 展开更多
关键词 large kernel convolution lkconv recover image details U lkconv network high quality single pixel imaging U Net low sampling conditions enhanced network structure large kernel convolution
暂未订购 下载PDF
Pedestrian attribute classification with multi-scale and multi-label convolutional neural networks 认领 引用
12
作者 朱建清 Zeng Huanqiang +2 位作者 Zhang Yuzhao Zheng Lixin Cai Canhui 《High Technology Letters》 EI CAS 2018年第1期53-61,共9页
Pedestrian attribute classification from a pedestrian image captured in surveillance scenarios is challenging due to diverse clothing appearances,varied poses and different camera views. A multiscale and multi-label c... Pedestrian attribute classification from a pedestrian image captured in surveillance scenarios is challenging due to diverse clothing appearances,varied poses and different camera views. A multiscale and multi-label convolutional neural network( MSMLCNN) is proposed to predict multiple pedestrian attributes simultaneously. The pedestrian attribute classification problem is firstly transformed into a multi-label problem including multiple binary attributes needed to be classified. Then,the multi-label problem is solved by fully connecting all binary attributes to multi-scale features with logistic regression functions. Moreover,the multi-scale features are obtained by concatenating those featured maps produced from multiple pooling layers of the MSMLCNN at different scales. Extensive experiment results show that the proposed MSMLCNN outperforms state-of-the-art pedestrian attribute classification methods with a large margin. 展开更多
关键词 pedestrian attribute classification multi-scale features multi-label classification convolutional neural network (CNN)
暂未订购 下载PDF
Improved multi-scale feature fusion for infrared small target detection based on YOLOv8 认领 引用
13
作者 DING Shangsi YANG Guiqin GAN Bingkun 《Journal of Measurement Science and Instrumentation》 CAS CSCD 2026年第2期208-218,共11页
Aiming at the problems of low target pixels and intricate background in small target detection in infrared scenes,a target detection model based on multi-scale feature extraction with YOLOv8 was proposed.Firstly,all d... Aiming at the problems of low target pixels and intricate background in small target detection in infrared scenes,a target detection model based on multi-scale feature extraction with YOLOv8 was proposed.Firstly,all downsampling convolutions in the network were replaced with the Haar wavelet downsampling(HWD)module to better preserve fine-grained details in infrared imagery during downsampling.Secondly,the spatial pyramid pooling-fast(SPPF)module was improved by introducing separable convolutions,which expanded the receptive field in both horizontal and vertical directions,enabling more comprehensive spatial information capture.Furthermore,a novel C2f_CDWR module was designed using dilated convolutions with varying dilation rates to achieve adaptive feature extraction across multiple receptive fields,thus enhancing detection performance for objects of different sizes.Finally,to improve localization accuracy,the original CIoU loss in YOLOv8 was replaced with Inner-SIoU,which effectively improved bounding box regression accuracy and significantly boosted the model’s capability in detecting small infrared targets.The experimental evaluation on the HIT-UAV dataset shows that the precision of the enhanced YOLOv8 model is 90.5%,the recall rate is 75.9%,and the mean average precision is 85.7%.In terms of infrared target detection,its performance was significantly better than that of the baseline YOLOv8 model and other benchmark models. 展开更多
关键词 infrared image small object detection multi-scale feature extraction dilation convolution YOLOv8v8
暂未订购 下载PDF
Multi-Scale Convolutional Gated Recurrent Unit Networks for Tool Wear Prediction in Smart Manufacturing 认领 引用 被引量:7
14
作者 Weixin Xu Huihui Miao +3 位作者 Zhibin Zhao Jinxin Liu Chuang Sun Ruqiang Yan 《Chinese Journal of Mechanical Engineering》 SCIE EI CAS CSCD 2021年第3期130-145,共16页
As an integrated application of modern information technologies and artificial intelligence,Prognostic and Health Management(PHM)is important for machine health monitoring.Prediction of tool wear is one of the symboli... As an integrated application of modern information technologies and artificial intelligence,Prognostic and Health Management(PHM)is important for machine health monitoring.Prediction of tool wear is one of the symbolic applications of PHM technology in modern manufacturing systems and industry.In this paper,a multi-scale Convolutional Gated Recurrent Unit network(MCGRU)is proposed to address raw sensory data for tool wear prediction.At the bottom of MCGRU,six parallel and independent branches with different kernel sizes are designed to form a multi-scale convolutional neural network,which augments the adaptability to features of different time scales.These features of different scales extracted from raw data are then fed into a Deep Gated Recurrent Unit network to capture long-term dependencies and learn significant representations.At the top of the MCGRU,a fully connected layer and a regression layer are built for cutting tool wear prediction.Two case studies are performed to verify the capability and effectiveness of the proposed MCGRU network and results show that MCGRU outperforms several state-of-the-art baseline models. 展开更多
关键词 Tool wear prediction Multi-scale Convolutional neural networks Gated recurrent unit
暂未订购 下载PDF
MSSTNet:Multi-scale facial videos pulse extraction network based on separable spatiotemporal convolution and dimension separable attention 认领 引用
15
作者 Changchen ZHAO Hongsheng WANG Yuanjing FENG 《虚拟现实与智能硬件(中英文)》 EI 2023年第2期124-141,共18页
Background The use of remote photoplethysmography(rPPG)to estimate blood volume pulse in a noncontact manner has been an active research topic in recent years.Existing methods are primarily based on a singlescale regi... Background The use of remote photoplethysmography(rPPG)to estimate blood volume pulse in a noncontact manner has been an active research topic in recent years.Existing methods are primarily based on a singlescale region of interest(ROI).However,some noise signals that are not easily separated in a single-scale space can be easily separated in a multi-scale space.Also,existing spatiotemporal networks mainly focus on local spatiotemporal information and do not emphasize temporal information,which is crucial in pulse extraction problems,resulting in insufficient spatiotemporal feature modelling.Methods Here,we propose a multi-scale facial video pulse extraction network based on separable spatiotemporal convolution(SSTC)and dimension separable attention(DSAT).First,to solve the problem of a single-scale ROI,we constructed a multi-scale feature space for initial signal separation.Second,SSTC and DSAT were designed for efficient spatiotemporal correlation modeling,which increased the information interaction between the long-span time and space dimensions;this placed more emphasis on temporal features.Results The signal-to-noise ratio(SNR)of the proposed network reached 9.58dB on the PURE dataset and 6.77dB on the UBFC-rPPG dataset,outperforming state-of-the-art algorithms.Conclusions The results showed that fusing multi-scale signals yielded better results than methods based on only single-scale signals.The proposed SSTC and dimension-separable attention mechanism will contribute to more accurate pulse signal extraction. 展开更多
关键词 Remote photoplethysmography Heart rate Separable spatiotemporal convolution Dimension separable attention Multi-scale Neural network
暂未订购 下载PDF
A Lightweight Convolutional Neural Network with Hierarchical Multi-Scale Feature Fusion for Image Classification 认领 引用 被引量:2
16
作者 Adama Dembele Ronald Waweru Mwangi Ananda Omutokoh Kube 《Journal of Computer and Communications》 2024年第2期173-200,共28页
Convolutional neural networks (CNNs) are widely used in image classification tasks, but their increasing model size and computation make them challenging to implement on embedded systems with constrained hardware reso... Convolutional neural networks (CNNs) are widely used in image classification tasks, but their increasing model size and computation make them challenging to implement on embedded systems with constrained hardware resources. To address this issue, the MobileNetV1 network was developed, which employs depthwise convolution to reduce network complexity. MobileNetV1 employs a stride of 2 in several convolutional layers to decrease the spatial resolution of feature maps, thereby lowering computational costs. However, this stride setting can lead to a loss of spatial information, particularly affecting the detection and representation of smaller objects or finer details in images. To maintain the trade-off between complexity and model performance, a lightweight convolutional neural network with hierarchical multi-scale feature fusion based on the MobileNetV1 network is proposed. The network consists of two main subnetworks. The first subnetwork uses a depthwise dilated separable convolution (DDSC) layer to learn imaging features with fewer parameters, which results in a lightweight and computationally inexpensive network. Furthermore, depthwise dilated convolution in DDSC layer effectively expands the field of view of filters, allowing them to incorporate a larger context. The second subnetwork is a hierarchical multi-scale feature fusion (HMFF) module that uses parallel multi-resolution branches architecture to process the input feature map in order to extract the multi-scale feature information of the input image. Experimental results on the CIFAR-10, Malaria, and KvasirV1 datasets demonstrate that the proposed method is efficient, reducing the network parameters and computational cost by 65.02% and 39.78%, respectively, while maintaining the network performance compared to the MobileNetV1 baseline. 展开更多
关键词 MobileNet Image Classification Lightweight Convolutional Neural Network Depthwise Dilated Separable Convolution Hierarchical Multi-Scale Feature Fusion
暂未订购 下载PDF
M2ANet:Multi-branch and multi-scale attention network for medical image segmentation 认领 引用 被引量:1
17
作者 Wei Xue Chuanghui Chen +3 位作者 Xuan Qi Jian Qin Zhen Tang Yongsheng He 《Chinese Physics B》 SCIE EI CAS CSCD 2025年第8期547-559,共13页
Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to ... Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to the inability to effectively capture global information from images,CNNs can easily lead to loss of contours and textures in segmentation results.Notice that the transformer model can effectively capture the properties of long-range dependencies in the image,and furthermore,combining the CNN and the transformer can effectively extract local details and global contextual features of the image.Motivated by this,we propose a multi-branch and multi-scale attention network(M2ANet)for medical image segmentation,whose architecture consists of three components.Specifically,in the first component,we construct an adaptive multi-branch patch module for parallel extraction of image features to reduce information loss caused by downsampling.In the second component,we apply residual block to the well-known convolutional block attention module to enhance the network’s ability to recognize important features of images and alleviate the phenomenon of gradient vanishing.In the third component,we design a multi-scale feature fusion module,in which we adopt adaptive average pooling and position encoding to enhance contextual features,and then multi-head attention is introduced to further enrich feature representation.Finally,we validate the effectiveness and feasibility of the proposed M2ANet method through comparative experiments on four benchmark medical image segmentation datasets,particularly in the context of preserving contours and textures. 展开更多
关键词 medical image segmentation convolutional neural network multi-branch attention multi-scale feature fusion
暂未订购 下载PDF
Deep Multi-Scale and Attention-Based Architectures for Semantic Segmentation in Biomedical Imaging 认领 引用
18
作者 Majid Harouni Vishakha Goyal +2 位作者 Gabrielle Feldman Sam Michael Ty C.Voss 《Computers, Materials & Continua》 SCIE EI 2025年第10期331-366,共36页
Semantic segmentation plays a foundational role in biomedical image analysis, providing precise information about cellular, tissue, and organ structures in both biological and medical imaging modalities. Traditional a... Semantic segmentation plays a foundational role in biomedical image analysis, providing precise information about cellular, tissue, and organ structures in both biological and medical imaging modalities. Traditional approaches often fail in the face of challenges such as low contrast, morphological variability, and densely packed structures. Recent advancements in deep learning have transformed segmentation capabilities through the integration of fine-scale detail preservation, coarse-scale contextual modeling, and multi-scale feature fusion. This work provides a comprehensive analysis of state-of-the-art deep learning models, including U-Net variants, attention-based frameworks, and Transformer-integrated networks, highlighting innovations that improve accuracy, generalizability, and computational efficiency. Key architectural components such as convolution operations, shallow and deep blocks, skip connections, and hybrid encoders are examined for their roles in enhancing spatial representation and semantic consistency. We further discuss the importance of hierarchical and instance-aware segmentation and annotation in interpreting complex biological scenes and multiplexed medical images. By bridging methodological developments with diverse application domains, this paper outlines current trends and future directions for semantic segmentation, emphasizing its critical role in facilitating annotation, diagnosis, and discovery in biomedical research. 展开更多
关键词 Biomedical semantic segmentation multi-scale feature fusion fine-and coarse-scale features convolution operations shallow and deep blocks skip connections
暂未订购 下载PDF
A Kernel-Based Convolution Method to Calculate Sparse Aerial Image Intensity for Lithography Simulation 认领 引用 被引量:3
19
作者 史峥 王国雄 +2 位作者 严晓浪 陈志锦 高根生 《Journal of Semiconductors》 CAS 北大核心 2003年第4期357-361,共5页
Optical proximity correction (OPC) systems require an accurate and fast way to predict how patterns will be transferred to the wafer.Based on Gabor's 'reduction to principal waves',a partially coherent ima... Optical proximity correction (OPC) systems require an accurate and fast way to predict how patterns will be transferred to the wafer.Based on Gabor's 'reduction to principal waves',a partially coherent imaging system can be represented as a superposition of coherent imaging systems,so an accurate and fast sparse aerial image intensity calculation algorithm for lithography simulation is presented based on convolution kernels,which also include simulating the lateral diffusion and some mask processing effects via Gaussian filter.The simplicity of this model leads to substantial computational and analytical benefits.Efficiency of this method is also shown through simulation results. 展开更多
关键词 lithography simulation optical proximity correction convolution kernels
暂未订购 下载PDF
Optimization of optical convolution kernel of optoelectronic hybrid convolution neural network 认领 引用 被引量:1
20
作者 XU Xiaofeng ZHU Lianqing +3 位作者 ZHUANG Wei ZHANG Dongliang LU Lidan YUAN Pei 《Optoelectronics Letters》 EI 2022年第3期181-186,共6页
To enhance the optical computation’s utilization efficiency, we develop an optimization method for optical convolution kernel in the optoelectronic hybrid convolution neural network(OHCNN). To comply with the actual ... To enhance the optical computation’s utilization efficiency, we develop an optimization method for optical convolution kernel in the optoelectronic hybrid convolution neural network(OHCNN). To comply with the actual calculation process, the convolution kernel is expanded from single-channel to two-channel, containing positive and negative weights. The Fashion-MNIST dataset is used to test the network architecture’s accuracy, and the accuracy is improved by 7.5% with the optimized optical convolution kernel. The energy efficiency ratio(EER) of two-channel network is 46.7% higher than that of the single-channel network, and it is 2.53 times of that of traditional electronic products. 展开更多
关键词 convolution kernel weights
暂未订购 下载PDF
上一页 1 2 36 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈