期刊文献+
共找到4,134篇文章
< 1 2 207 >
每页显示 20 50 100
GLMCNet: A Global-Local Multiscale Context Network for High-Resolution Remote Sensing Image Semantic Segmentation 认领 引用
1
作者 Yanting Zhang Qiyue Liu +4 位作者 Chuanzhao Tian Xuewen Li Na Yang Feng Zhang Hongyue Zhang 《Computers, Materials & Continua》 SCIE EI 2026年第1期2086-2110,共25页
High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes an... High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes and wealth of spatial details pose challenges for semantic segmentation.While convolutional neural networks(CNNs)excel at capturing local features,they are limited in modeling long-range dependencies.Conversely,transformers utilize multihead self-attention to integrate global context effectively,but this approach often incurs a high computational cost.This paper proposes a global-local multiscale context network(GLMCNet)to extract both global and local multiscale contextual information from HRSIs.A detail-enhanced filtering module(DEFM)is proposed at the end of the encoder to refine the encoder outputs further,thereby enhancing the key details extracted by the encoder and effectively suppressing redundant information.In addition,a global-local multiscale transformer block(GLMTB)is proposed in the decoding stage to enable the modeling of rich multiscale global and local information.We also design a stair fusion mechanism to transmit deep semantic information from deep to shallow layers progressively.Finally,we propose the semantic awareness enhancement module(SAEM),which further enhances the representation of multiscale semantic features through spatial attention and covariance channel attention.Extensive ablation analyses and comparative experiments were conducted to evaluate the performance of the proposed method.Specifically,our method achieved a mean Intersection over Union(mIoU)of 86.89%on the ISPRS Potsdam dataset and 84.34%on the ISPRS Vaihingen dataset,outperforming existing models such as ABCNet and BANet. 展开更多
关键词 Multiscale context attention mechanism remote sensing images semantic segmentation
暂未订购 下载PDF
A Study on Improving the Accuracy of Semantic Segmentation for Autonomous Driving 认领 引用
2
作者 Bin Zhang Zhancheng Xu 《Computers, Materials & Continua》 SCIE EI 2026年第2期321-332,共12页
This study aimed to enhance the performance of semantic segmentation for autonomous driving by improving the 2DPASS model.Two novel improvements were proposed and implemented in this paper:dynamically adjusting the lo... This study aimed to enhance the performance of semantic segmentation for autonomous driving by improving the 2DPASS model.Two novel improvements were proposed and implemented in this paper:dynamically adjusting the loss function ratio and integrating an attention mechanism(CBAM).First,the loss function weights were adjusted dynamically.The grid search method is used for deciding the best ratio of 7:3.It gives greater emphasis to the cross-entropy loss,which resulted in better segmentation performance.Second,CBAM was applied at different layers of the 2Dencoder.Heatmap analysis revealed that introducing it after the second block of 2D image encoding produced the most effective enhancement of important feature representation.The training epoch was chosen for optimizing the best value by experiments,which improved model convergence and overall accuracy.To evaluate the proposed approach,experiments were conducted based on the SemanticKITTI database.The results showed that the improved model achieved higher segmentation accuracy by 64.31%,improved 11.47% in mIoU compared with the conventional 2DPASS model(baseline:52.84%).It was more effective at detecting small and distant objects and clearly identifying boundaries between different classes.Issues such as noise and variations in data distribution affected its accuracy,indicating the need for further refinement.Overall,the proposed improvements to the 2DPASS model demonstrated the potential to advance semantic segmentation technology and contributed to a more reliable perception of complex,dynamic environments in autonomous vehicles.Accurate segmentation enhances the vehicle’s ability to distinguish different objects,and this improvement directly supports safer navigation,robust decision-making,and efficient path planning,making it highly applicable to real-world deployment of autonomous systems in urban and highway settings. 展开更多
关键词 Autonomous driving system semantic segmentation 2DPASS deep learning model
暂未订购 下载PDF
PointNMSA: An Improved PointNeXt Network with Non-Local Multi-Scale Aggregation for 3D Point Cloud Semantic Segmentation 认领 引用
3
作者 Aihua Wu Chenlu Huang 《Computers, Materials & Continua》 SCIE EI 2026年第8期1632-1649,共18页
Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Althoug... Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts. 展开更多
关键词 3D point cloud semantic segmentation indoor scene understanding multi-scale feature aggregation non-local context integration PointNeXt
暂未订购 下载PDF
A Semantic Segmentation Network for Colorectal Polyp Images With Progressive Fusion of Dual-Branch Features 认领 引用
4
作者 Tianxu Yan Jiabin Yu +6 位作者 Zheng Li Liangyu Chen Hongmei Mi Luyang Chen Wei Si Dongping Zhang Hui Lin 《CAAI Transactions on Intelligence Technology》 SCIE EI CSCD 2026年第3期900-919,共20页
Accurate segmentation of colorectal polyps is essential for early colorectal cancer screening,yet remains challenging due to weak foreground–background contrast,disrupted boundaries caused by specular reflections and... Accurate segmentation of colorectal polyps is essential for early colorectal cancer screening,yet remains challenging due to weak foreground–background contrast,disrupted boundaries caused by specular reflections and intestinal folds,and pronounced scale variation among polyps.These factors make it difficult for existing methods to jointly preserve fine boundary details and robust global semantic context.To address these task‐specific challenges,we propose a Dual‐branch Feature Progressive Fusion Network(DFPF‐Net)for colorectal polyp segmentation.DFPF‐Net adopts a dual‐encoder architecture that integrates a CNN‐based encoder for local and boundary‐sensitive representation for global semantic modelling.A boundaryaware branch equipped with stacked Inversely Perceive Information Layers(IPILs)enhances ambiguous and fragmented contours,while the semantic branch incorporates Misalignment Fusion Modules(MFMs)and a Misaligned Single‐layer Reinforcement Module(MSRM)to alleviate semantic misalignment and insufficient cross‐scale interaction.Furthermore,a Perceptual Information Fusion Module(PIFM)enables effective semantic–boundary collaboration,and a Multi‐level Residual Decoding Module(MRDM)progressively reconstructs structurally consistent segmentation outputs.Extensive experiments on multiple public colonoscopy datasets demonstrate that DFPF‐Net achieves competitive and robust segmentation performance.In particular,on the challenging ETIS dataset,DFPF‐Net attains 0.785 mDice and 0.704 mIoU,indicating its capability in handling complex structures and ambiguous boundaries in colorectal polyp segmentation. 展开更多
关键词 boundary‐aware learning colorectal polyp segmentation multi‐scale fusion semantic segmentation vision transformer
暂未订购 下载PDF
Importance-Aware Image Segmentation-Based Semantic Communication for Autonomous Driving 认领 引用
5
作者 Lyu Jie Tong Haonan +4 位作者 Pan Qiang Zhang Zhilong He Xinxin Luo Tao Yin Changchuan 《China Communications》 SCIE EI CSCD 2026年第2期228-243,共16页
This article studies the problem of image segmentation-based semantic communication in autonomous driving.In real traffic scenes,the detecting of objects(e.g.,vehicles and pedestrians)is more important to guarantee dr... This article studies the problem of image segmentation-based semantic communication in autonomous driving.In real traffic scenes,the detecting of objects(e.g.,vehicles and pedestrians)is more important to guarantee driving safety,which is always ignored in existing works.Therefore,we propose a vehicular image segmentation-oriented semantic communication system,termed VIS-SemCom,focusing on transmitting and recovering image semantic features of high-important objects to reduce transmission redundancy.First,we develop a semantic codec based on Swin Transformer architecture,which expands the perceptual field thus improving the segmentation accuracy.To highlight the important objects'accuracy,we propose a multi-scale semantic extraction method by assigning the number of Swin Transformer blocks for diverse resolution semantic features.Also,an importance-aware loss incorporating important levels is devised,and an online hard example mining(OHEM)strategy is proposed to handle small sample issues in the dataset.Finally,experimental results demonstrate that the proposed VIS-SemCom can achieve a significant mean intersection over union(mIoU)performance in the SNR regions,a reduction of transmitted data volume by about 60%at 60%mIoU,and improve the segmentation accuracy of important objects,compared to baseline image communication. 展开更多
关键词 autonomous driving image segmentation semantic communication Swin Transformer
暂未订购 下载PDF
Subnetwork-based federated few-shot semantic segmentation of organ images 认领 引用
6
作者 Junpeng WU Meng ZHAO Huanping ZHANG 《Optoelectronics Letters》 EI 2026年第5期275-281,共7页
Federated learning(FL),as a distributed learning paradigm,allows multiple medical institutions to collaborate on learning without the need to centralize all client data.However,existing methods pay little attention to... Federated learning(FL),as a distributed learning paradigm,allows multiple medical institutions to collaborate on learning without the need to centralize all client data.However,existing methods pay little attention to more challenging medical image semantic segmentation tasks,especially in the scenario of the imbalanced dataset in federated few-shot learning(FSL).In this paper,we propose a subnetwork-based federated few-shot organ image segmentation method.Firstly,individual clients train using local training samples and then upload local model gradients to the server.The server utilizes their respective local model gradients to update the subnetwork maintained on the server and generate aggregation weights for forming personalized model parameters.Through this method,we can learn the similarities between different clients to address data heterogeneity issues.In addition,to enhance the communication efficiency between clients and the server,we have also designed a personalized layer aggregation strategy,which only transmits partial layer model parameters during the communication process to improve communication efficiency.Finally,we conducted experiments on abdomen magnetic resonance imaging(ABD-MRI)and abdomen computed tomography(ABD-CT)datasets to demonstrate the effectiveness of our method. 展开更多
关键词 subnetwork based federated learning semantic segmentation few shot learning imbalanced dataset medical image semantic segmentation local training federated learning fl collaborate learning
暂未订购 下载PDF
Context Patch Fusion with Class Token Enhancement for Weakly Supervised Semantic Segmentation 认领 引用
7
作者 Yiyang Fu Hui Li Wangyu Wu 《Computer Modeling in Engineering & Sciences》 SCIE EI 2026年第1期1130-1150,共21页
Weakly Supervised Semantic Segmentation(WSSS),which relies only on image-level labels,has attracted significant attention for its cost-effectiveness and scalability.Existing methods mainly enhance inter-class distinct... Weakly Supervised Semantic Segmentation(WSSS),which relies only on image-level labels,has attracted significant attention for its cost-effectiveness and scalability.Existing methods mainly enhance inter-class distinctions and employ data augmentation to mitigate semantic ambiguity and reduce spurious activations.However,they often neglect the complex contextual dependencies among image patches,resulting in incomplete local representations and limited segmentation accuracy.To address these issues,we propose the Context Patch Fusion with Class Token Enhancement(CPF-CTE)framework,which exploits contextual relations among patches to enrich feature repre-sentations and improve segmentation.At its core,the Contextual-Fusion Bidirectional Long Short-Term Memory(CF-BiLSTM)module captures spatial dependencies between patches and enables bidirectional information flow,yield-ing a more comprehensive understanding of spatial correlations.This strengthens feature learning and segmentation robustness.Moreover,we introduce learnable class tokens that dynamically encode and refine class-specific semantics,enhancing discriminative capability.By effectively integrating spatial and semantic cues,CPF-CTE produces richer and more accurate representations of image content.Extensive experiments on PASCAL VOC 2012 and MS COCO 2014 validate that CPF-CTE consistently surpasses prior WSSS methods. 展开更多
关键词 Weakly supervised semantic segmentation context-fusion class enhancement
暂未订购 下载PDF
CAWASeg:Class Activation Graph Driven Adaptive Weight Adjustment for Semantic Segmentation 认领 引用
8
作者 Hailong Wang Minglei Duan +1 位作者 Lu Yao Hao Li 《Computers, Materials & Continua》 SCIE EI 2026年第3期1071-1091,共21页
In image analysis,high-precision semantic segmentation predominantly relies on supervised learning.Despite significant advancements driven by deep learning techniques,challenges such as class imbalance and dynamic per... In image analysis,high-precision semantic segmentation predominantly relies on supervised learning.Despite significant advancements driven by deep learning techniques,challenges such as class imbalance and dynamic performance evaluation persist.Traditional weighting methods,often based on pre-statistical class counting,tend to overemphasize certain classes while neglecting others,particularly rare sample categories.Approaches like focal loss and other rare-sample segmentation techniques introduce multiple hyperparameters that require manual tuning,leading to increased experimental costs due to their instability.This paper proposes a novel CAWASeg framework to address these limitations.Our approach leverages Grad-CAM technology to generate class activation maps,identifying key feature regions that the model focuses on during decision-making.We introduce a Comprehensive Segmentation Performance Score(CSPS)to dynamically evaluate model performance by converting these activation maps into pseudo mask and comparing them with Ground Truth.Additionally,we design two adaptive weights for each class:a Basic Weight(BW)and a Ratio Weight(RW),which the model adjusts during training based on real-time feedback.Extensive experiments on the COCO-Stuff,CityScapes,and ADE20k datasets demonstrate that our CAWASeg framework significantly improves segmentation performance for rare sample categories while enhancing overall segmentation accuracy.The proposed method offers a robust and efficient solution for addressing class imbalance in semantic segmentation tasks. 展开更多
关键词 Semantic segmentation class activation graph adaptive weight adjustment pseudo mask
暂未订购 下载PDF
Global context-aware multi-scale feature iterative refinement for aviation-road traffic semantic segmentation 认领 引用
9
作者 Mengyue ZHANG Shichun YANG +1 位作者 Xinjie FENG Yaoguang CAO 《Chinese Journal of Aeronautics》 SCIE EI CAS CSCD 2026年第2期429-441,共13页
Semantic segmentation for mixed scenes of aerial remote sensing and road traffic is one of the key technologies for visual perception of flying cars.The State-of-the-Art(SOTA)semantic segmentation methods have made re... Semantic segmentation for mixed scenes of aerial remote sensing and road traffic is one of the key technologies for visual perception of flying cars.The State-of-the-Art(SOTA)semantic segmentation methods have made remarkable achievements in both fine-grained segmentation and real-time performance.However,when faced with the huge differences in scale and semantic categories brought about by the mixed scenes of aerial remote sensing and road traffic,they still face great challenges and there is little related research.Addressing the above issue,this paper proposes a semantic segmentation model specifically for mixed datasets of aerial remote sensing and road traffic scenes.First,a novel decoding-recoding multi-scale feature iterative refinement structure is proposed,which utilizes the re-integration and continuous enhancement of multi-scale information to effectively deal with the huge scale differences between cross-domain scenes,while using a fully convolutional structure to ensure the lightweight and real-time requirements.Second,a welldesigned cross-window attention mechanism combined with a global information integration decoding block forms an enhanced global context perception,which can effectively capture the long-range dependencies and multi-scale global context information of different scenes,thereby achieving fine-grained semantic segmentation.The proposed method is tested on a large-scale mixed dataset of aerial remote sensing and road traffic scenes.The results confirm that it can effectively deal with the problem of large-scale differences in cross-domain scenes.Its segmentation accuracy surpasses that of the SOTA methods,which meets the real-time requirements. 展开更多
关键词 Aviation-road traffic Flying cars Global context-aware Multi-scale feature iterative refinement Semantic segmentation
暂未订购 下载PDF
Multiscale Long-Distance Feature Aggregation Network for Geospatial Semantic Segmentation in High-Resolution Remote Sensing Imagery 认领 引用
10
作者 Guangyu Xu Yuxi Ban +3 位作者 Legend Zhang Junmin Lyu Feng Bao Wenfeng Zheng 《Computer Modeling in Engineering & Sciences》 SCIE EI 2026年第7期935-957,共23页
High-resolution remote sensing semantic segmentation is a fundamental task in Geospatial Artificial Intelligence(GeoAI).Existing CNN-based methods are effective for local and multiscale feature extraction but often la... High-resolution remote sensing semantic segmentation is a fundamental task in Geospatial Artificial Intelligence(GeoAI).Existing CNN-based methods are effective for local and multiscale feature extraction but often lack progressive cross-scale semantic propagation,while attention-and Transformer-based methods improve global spatial modeling but generally ignore frequency-domain regularities.To address these limitations,this study proposes a Multiscale Long-Distance Feature Aggregation Network(MLFANet),a unified spatial-frequency segmentation framework for high-resolution remote sensing imagery.MLFANet introduces three key components:a Multiscale Global Dependency Extraction module for cascaded cross-scale contextual refinement,an FFT-based frequency-domain branch with learnable global filtering for capturing structural and texture regularities,and a bidirectional Spatial-Frequency Fusion module for adaptively aligning spatial details with frequency responses.Experiments on the ISPRS Potsdam and Vaihingen datasets demonstrate the effectiveness and feasibility of the proposed model.MLFANet achieves AF,MIoU,and OA values of 86.03%,76.21%,and 88.70%on Potsdam,and 83.17%,71.90%,and 86.33%on Vaihingen,respectively,outperforming representative CNN-based,attention-based,and hybrid models in overall metrics.In terms of computational complexity,MLFANet requires 17.49 GFLOPs under an input size of 256×256 pixels,indicating its practical feasibility for patch-based high-resolution remote sensing segmentation.Ablation studies further verify that multiscale dependency extraction,frequency-domain modeling,and adaptive spatial-frequency fusion each contribute to the final performance. 展开更多
关键词 Geospatial artificial intelligence(GeoAI) high-resolution remote sensing images semantic segmentation spatial-frequency fusion multiscale feature aggregation attention mechanism
暂未订购 下载PDF
Urban Point Cloud Semantic Segmentation Incorporating Hybrid and External Attention Modules 认领 引用
11
作者 MAO Jingyi WANG Jingxue BU Lijing 《Journal of Geodesy and Geoinformation Science》 CSCD 2026年第1期80-99,共20页
Imbalanced category distributions in the training data can significantly affect the performance of point clouds segmentation models based on deep learning.However,point clouds often exhibit significant class imbalance... Imbalanced category distributions in the training data can significantly affect the performance of point clouds segmentation models based on deep learning.However,point clouds often exhibit significant class imbalance in urban environments.This imbalance causes the network to under-learn minority categories during training,making it difficult to identify these classes during prediction accurately and thereby limiting classification accuracy.To address this issue,we propose a Multi-scale Hybrid Attention network(MHAnet),which integrates the hybrid and the external attention module.The hybrid attention module captures multi-scale features and highlights key regions,improving the network’s capability to differentiate minority classes.The external attention module introduces global context and dynamically adjusts the distribution of feature weights to reduce the imbalance caused by category imbalance.Additionally,to further extract common features among similar categories,a hybrid loss function is introduced to balance the contribution of different categories during training.Experimental results on the Semantic3D showed that MHAnet achieved excellent performance in urban point cloud semantic segmentation,with an overall accuracy(OA)of 93.9%and a mean intersection over union(mIoU)of 71.13%,outperforming mainstream methods. 展开更多
关键词 semantic segmentation static laser scanning hybrid attention module external attention module multi-scale features
暂未订购 下载PDF
Intelligent Semantic Segmentation with Vision Transformers for Aerial Vehicle Monitoring 认领 引用
12
作者 Moneerah Alotaibi 《Computers, Materials & Continua》 SCIE EI 2026年第1期1629-1648,共20页
Advanced traffic monitoring systems encounter substantial challenges in vehicle detection and classification due to the limitations of conventional methods,which often demand extensive computational resources and stru... Advanced traffic monitoring systems encounter substantial challenges in vehicle detection and classification due to the limitations of conventional methods,which often demand extensive computational resources and struggle with diverse data acquisition techniques.This research presents a novel approach for vehicle classification and recognition in aerial image sequences,integrating multiple advanced techniques to enhance detection accuracy.The proposed model begins with preprocessing using Multiscale Retinex(MSR)to enhance image quality,followed by Expectation-Maximization(EM)Segmentation for precise foreground object identification.Vehicle detection is performed using the state-of-the-art YOLOv10 framework,while feature extraction incorporates Maximally Stable Extremal Regions(MSER),Dense Scale-Invariant Feature Transform(Dense SIFT),and Zernike Moments Features to capture distinct object characteristics.Feature optimization is further refined through a Hybrid Swarm-based Optimization algorithm,ensuring optimal feature selection for improved classification performance.The final classification is conducted using a Vision Transformer,leveraging its robust learning capabilities for enhanced accuracy.Experimental evaluations on benchmark datasets,including UAVDT and the Unmanned Aerial Vehicle Intruder Dataset(UAVID),demonstrate the superiority of the proposed approach,achieving an accuracy of 94.40%on UAVDT and 93.57%on UAVID.The results highlight the efficacy of the model in significantly enhancing vehicle detection and classification in aerial imagery,outperforming existing methodologies and offering a statistically validated improvement for intelligent traffic monitoring systems compared to existing approaches. 展开更多
关键词 Machine learning semantic segmentation remote sensors deep learning object monitoring system
暂未订购 下载PDF
Enhancing convolution for Transformer-based weakly supervised semantic segmentation 认领 引用
13
作者 LIU Yu TAN Diaoyin +1 位作者 ZHOU Wen XIAO Huaxin 《Journal of Systems Engineering and Electronics》 SCIE CSCD 2026年第1期84-93,共10页
Weakly supervised semantic segmentation(WSSS)is a tricky task,which only provides category information for segmentation prediction.Thus,the key stage of WSSS is to generate the pseudo labels.For convolutional neural n... Weakly supervised semantic segmentation(WSSS)is a tricky task,which only provides category information for segmentation prediction.Thus,the key stage of WSSS is to generate the pseudo labels.For convolutional neural network(CNN)based methods,in which class activation mapping(CAM)is proposed to obtain the pseudo labels,and only concentrates on the most discriminative parts.Recently,transformer-based methods utilize attention map from the multi-headed self-attention(MHSA)module to predict pseudo labels,which usually contain obvious background noise and incoherent object area.To solve the above problems,we use the Conformer as our backbone,which is a parallel network based on convolutional neural network(CNN)and Transformer.The two branches generate pseudo labels and refine them independently,and can effectively combine the advantages of CNN and Transformer.However,the parallel structure is not close enough in the information communication.Thus,parallel structure can result in poor details about pseudo labels,and the background noise still exists.To alleviate this problem,we propose enhancing convolution CAM(ECCAM)model,which have three improved modules based on enhancing convolution,including deeper stem(DStem),convolutional feed-forward network(CFFN)and feature coupling unit with convolution(FCUConv).The ECCAM could make Conformer have tighter interaction between CNN and Transformer branches.After experimental verification,the improved modules we propose can help the network perceive more local information from images,making the final segmentation results more refined.Compared with similar architecture,our modules greatly improve the semantic segmentation performance and achieve70.2%mean intersection over union(mIoU)on the PASCAL VOC 2012 dataset. 展开更多
关键词 weakly supervised semantic segmentation transformer convolutional neural network
暂未订购 下载PDF
Improved SE-UNet network-based semantic segmentation and extraction of hidden geological significance in geological maps 认领 引用 被引量:1
14
作者 Kai Ma Jun-jie Liu +5 位作者 Si-qi Lu Ze-hua Huang Miao Tian Jun-yuan Deng Zhong Xie Qin-jun Qiu 《China Geology》 CAS CSCD 2025年第4期643-660,共18页
Automatic segmentation and recognition of content and element information in color geological map are of great significance for researchers to analyze the distribution of mineral resources and predict disaster informa... Automatic segmentation and recognition of content and element information in color geological map are of great significance for researchers to analyze the distribution of mineral resources and predict disaster information.This article focuses on color planar raster geological map(geological maps include planar geological maps,columnar maps,and profiles).While existing deep learning approaches are often used to segment general images,their performance is limited due to complex elements,diverse regional features,and complicated backgrounds for color geological map in the domain of geoscience.To address the issue,a color geological map segmentation model is proposed that combines the Felz clustering algorithm and an improved SE-UNet deep learning network(named GeoMSeg).Firstly,a symmetrical encoder-decoder structure backbone network based on UNet is constructed,and the channel attention mechanism SENet has been incorporated to augment the network’s capacity for feature representation,enabling the model to purposefully extract map information.The SE-UNet network is employed for feature extraction from the geological map and obtain coarse segmentation results.Secondly,the Felz clustering algorithm is used for super pixel pre-segmentation of geological maps.The coarse segmentation results are refined and modified based on the super pixel pre-segmentation results to obtain the final segmentation results.This study applies GeoMSeg to the constructed dataset,and the experimental results show that the algorithm proposed in this paper has superior performance compared to other mainstream map segmentation models,with an accuracy of 91.89%and a MIoU of 71.91%. 展开更多
关键词 Geological map UNet model Image segmentation Semantic segmentation Pixel pre-segmentation Clustering algorithm Attention mechanism Deep learning Artificial intelligence Geological survey engineering
暂未订购 下载PDF
CPEWS:Contextual Prototype-Based End-to-End Weakly Supervised Semantic Segmentation 认领 引用 被引量:1
15
作者 Xiaoyan Shao Jiaqi Han +2 位作者 Lingling Li Xuezhuan Zhao Jingjing Yan 《Computers, Materials & Continua》 SCIE EI 2025年第4期595-617,共23页
The primary challenge in weakly supervised semantic segmentation is effectively leveraging weak annotations while minimizing the performance gap compared to fully supervised methods.End-to-end model designs have gaine... The primary challenge in weakly supervised semantic segmentation is effectively leveraging weak annotations while minimizing the performance gap compared to fully supervised methods.End-to-end model designs have gained significant attention for improving training efficiency.Most current algorithms rely on Convolutional Neural Networks(CNNs)for feature extraction.Although CNNs are proficient at capturing local features,they often struggle with global context,leading to incomplete and false Class Activation Mapping(CAM).To address these limitations,this work proposes a Contextual Prototype-Based End-to-End Weakly Supervised Semantic Segmentation(CPEWS)model,which improves feature extraction by utilizing the Vision Transformer(ViT).By incorporating its intermediate feature layers to preserve semantic information,this work introduces the Intermediate Supervised Module(ISM)to supervise the final layer’s output,reducing boundary ambiguity and mitigating issues related to incomplete activation.Additionally,the Contextual Prototype Module(CPM)generates class-specific prototypes,while the proposed Prototype Discrimination Loss and Superclass Suppression Loss guide the network’s training,(LPDL)(LSSL)effectively addressing false activation without the need for extra supervision.The CPEWS model proposed in this paper achieves state-of-the-art performance in end-to-end weakly supervised semantic segmentation without additional supervision.The validation set and test set Mean Intersection over Union(MIoU)of PASCAL VOC 2012 dataset achieved 69.8%and 72.6%,respectively.Compared with ToCo(pre trained weight ImageNet-1k),MIoU on the test set is 2.1%higher.In addition,MIoU reached 41.4%on the validation set of the MS COCO 2014 dataset. 展开更多
关键词 End-to-end weakly supervised semantic segmentation vision transformer contextual prototype class activation map
暂未订购 下载PDF
Semantic segmentation of camouflage objects via fusing reconstructed multispectral and RGB images 认领 引用 被引量:1
16
作者 Feng Huang Gonghan Yang +5 位作者 Jing Chen Yixuan Xu Jingze Su Guimin Huang Shu Wang Wenxi Liu 《Defence Technology(防务技术)》 SCIE EI CAS CSCD 2025年第8期324-337,共14页
Accurate segmentation of camouflage objects in aerial imagery is vital for improving the efficiency of UAV-based reconnaissance and rescue missions.However,camouflage object segmentation is increasingly challenging du... Accurate segmentation of camouflage objects in aerial imagery is vital for improving the efficiency of UAV-based reconnaissance and rescue missions.However,camouflage object segmentation is increasingly challenging due to advances in both camouflage materials and biological mimicry.Although multispectral-RGB based technology shows promise,conventional dual-aperture multispectral-RGB imaging systems are constrained by imprecise and time-consuming registration and fusion across different modalities,limiting their performance.Here,we propose the Reconstructed Multispectral-RGB Fusion Network(RMRF-Net),which reconstructs RGB images into multispectral ones,enabling efficient multimodal segmentation using only an RGB camera.Specifically,RMRF-Net employs a divergentsimilarity feature correction strategy to minimize reconstruction errors and includes an efficient boundary-aware decoder to enhance object contours.Notably,we establish the first real-world aerial multispectral-RGB semantic segmentation of camouflage objects dataset,including 11 object categories.Experimental results demonstrate that RMRF-Net outperforms existing methods,achieving 17.38 FPS on the NVIDIA Jetson AGX Orin,with only a 0.96%drop in mIoU compared to the RTX 3090,showing its practical applicability in multimodal remote sensing. 展开更多
关键词 Camouflage object detection Reconstructed multispectral image(MSI) Unmanned aerial vehicle(UAV) Semantic segmentation Remote sensing
暂未订购 下载PDF
A 3D semantic segmentation network for accurate neuronal soma segmentation 认领 引用
17
作者 Li Ma Qi Zhong +2 位作者 Yezi Wang Xiaoquan Yang Qian Du 《Journal of Innovative Optical Health Sciences》 SCIE EI CSCD 2025年第1期67-83,共17页
Neuronal soma segmentation plays a crucial role in neuroscience applications.However,the fine structure,such as boundaries,small-volume neuronal somata and fibers,are commonly present in cell images,which pose a chall... Neuronal soma segmentation plays a crucial role in neuroscience applications.However,the fine structure,such as boundaries,small-volume neuronal somata and fibers,are commonly present in cell images,which pose a challenge for accurate segmentation.In this paper,we propose a 3D semantic segmentation network for neuronal soma segmentation to address this issue.Using an encoding-decoding structure,we introduce a Multi-Scale feature extraction and Adaptive Weighting fusion module(MSAW)after each encoding block.The MSAW module can not only emphasize the fine structures via an upsampling strategy,but also provide pixel-wise weights to measure the importance of the multi-scale features.Additionally,a dynamic convolution instead of normal convolution is employed to better adapt the network to input data with different distributions.The proposed MSAW-based semantic segmentation network(MSAW-Net)was evaluated on three neuronal soma images from mouse brain and one neuronal soma image from macaque brain,demonstrating the efficiency of the proposed method.It achieved an F1 score of 91.8%on Fezf2-2A-CreER dataset,97.1%on LSL-H2B-GFP dataset,82.8%on Thy1-EGFP-Mline dataset,and 86.9%on macaque dataset,achieving improvements over the 3D U-Net model by 3.1%,3.3%,3.9%,and 2.3%,respectively. 展开更多
关键词 Neuronal soma segmentation semantic segmentation network multi-scale feature extraction adaptive weighting fusion
暂未订购 下载PDF
Lightweight deep network and projection loss for eye semantic segmentation 认领 引用
18
作者 Qinjie Wang Tengfei Wang +1 位作者 Lizhuang Yang Hai Li 《中国科学技术大学学报》 CAS CSCD 北大核心 2025年第7期59-68,58,I0002,共10页
Semantic segmentation of eye images is a complex task with important applications in human–computer interaction,cognitive science,and neuroscience.Achieving real-time,accurate,and robust segmentation algorithms is cr... Semantic segmentation of eye images is a complex task with important applications in human–computer interaction,cognitive science,and neuroscience.Achieving real-time,accurate,and robust segmentation algorithms is crucial for computationally limited portable devices such as augmented reality and virtual reality.With the rapid advancements in deep learning,many network models have been developed specifically for eye image segmentation.Some methods divide the segmentation process into multiple stages to achieve model parameter miniaturization while enhancing output through post processing techniques to improve segmentation accuracy.These approaches significantly increase the inference time.Other networks adopt more complex encoding and decoding modules to achieve end-to-end output,which requires substantial computation.Therefore,balancing the model’s size,accuracy,and computational complexity is essential.To address these challenges,we propose a lightweight asymmetric UNet architecture and a projection loss function.We utilize ResNet-3 layer blocks to enhance feature extraction efficiency in the encoding stage.In the decoding stage,we employ regular convolutions and skip connections to upscale the feature maps from the latent space to the original image size,balancing the model size and segmentation accuracy.In addition,we leverage the geometric features of the eye region and design a projection loss function to further improve the segmentation accuracy without adding any additional inference computational cost.We validate our approach on the OpenEDS2019 dataset for virtual reality and achieve state-of-the-art performance with 95.33%mean intersection over union(mIoU).Our model has only 0.63M parameters and 350 FPS,which are 68%and 200%of the state-of-the-art model RITNet,respectively. 展开更多
关键词 lightweight deep network projection loss real-time semantic segmentation convolutional neural networks end-to-end
暂未订购 下载PDF
Deep Multi-Scale and Attention-Based Architectures for Semantic Segmentation in Biomedical Imaging 认领 引用
19
作者 Majid Harouni Vishakha Goyal +2 位作者 Gabrielle Feldman Sam Michael Ty C.Voss 《Computers, Materials & Continua》 SCIE EI 2025年第10期331-366,共36页
Semantic segmentation plays a foundational role in biomedical image analysis, providing precise information about cellular, tissue, and organ structures in both biological and medical imaging modalities. Traditional a... Semantic segmentation plays a foundational role in biomedical image analysis, providing precise information about cellular, tissue, and organ structures in both biological and medical imaging modalities. Traditional approaches often fail in the face of challenges such as low contrast, morphological variability, and densely packed structures. Recent advancements in deep learning have transformed segmentation capabilities through the integration of fine-scale detail preservation, coarse-scale contextual modeling, and multi-scale feature fusion. This work provides a comprehensive analysis of state-of-the-art deep learning models, including U-Net variants, attention-based frameworks, and Transformer-integrated networks, highlighting innovations that improve accuracy, generalizability, and computational efficiency. Key architectural components such as convolution operations, shallow and deep blocks, skip connections, and hybrid encoders are examined for their roles in enhancing spatial representation and semantic consistency. We further discuss the importance of hierarchical and instance-aware segmentation and annotation in interpreting complex biological scenes and multiplexed medical images. By bridging methodological developments with diverse application domains, this paper outlines current trends and future directions for semantic segmentation, emphasizing its critical role in facilitating annotation, diagnosis, and discovery in biomedical research. 展开更多
关键词 Biomedical semantic segmentation multi-scale feature fusion fine-and coarse-scale features convolution operations shallow and deep blocks skip connections
暂未订购 下载PDF
Bilateral Dual-Residual Real-Time Semantic Segmentation Network 认领 引用 被引量:1
20
作者 Shijie Xiang Dong Zhou +1 位作者 Dan Tian Zihao Wang 《Computers, Materials & Continua》 SCIE EI 2025年第4期497-515,共19页
Real-time semantic segmentation tasks place stringent demands on network inference speed,often requiring a reduction in network depth to decrease computational load.However,shallow networks tend to exhibit degradation... Real-time semantic segmentation tasks place stringent demands on network inference speed,often requiring a reduction in network depth to decrease computational load.However,shallow networks tend to exhibit degradation in feature extraction completeness and inference accuracy.Therefore,balancing high performance with real-time requirements has become a critical issue in the study of real-time semantic segmentation.To address these challenges,this paper proposes a lightweight bilateral dual-residual network.By introducing a novel residual structure combined with feature extraction and fusion modules,the proposed network significantly enhances representational capacity while reducing computational costs.Specifically,an improved compound residual structure is designed to optimize the efficiency of information propagation and feature extraction.Furthermore,the proposed feature extraction and fusion module enables the network to better capture multi-scale information in images,improving the ability to detect both detailed and global semantic features.Experimental results on the publicly available Cityscapes dataset demonstrate that the proposed lightweight dual-branch network achieves outstanding performance while maintaining low computational complexity.In particular,the network achieved a mean Intersection over Union(mIoU)of 78.4%on the Cityscapes validation set,surpassing many existing semantic segmentation models.Additionally,in terms of inference speed,the network reached 74.5 frames per second when tested on an NVIDIA GeForce RTX 3090 GPU,significantly improving real-time performance. 展开更多
关键词 Real-time residual structure semantic segmentation feature fusion
暂未订购 下载PDF
上一页 1 2 207 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈