Zanthoxylum bungeanum Maxim,generally called prickly ash,is widely grown in China.Zanthoxylum rust is the main disease affecting the growth and quality of Zanthoxylum.Traditional method for recognizing the degree of i...Zanthoxylum bungeanum Maxim,generally called prickly ash,is widely grown in China.Zanthoxylum rust is the main disease affecting the growth and quality of Zanthoxylum.Traditional method for recognizing the degree of infection of Zanthoxylum rust mainly rely on manual experience.Due to the complex colors and shapes of rust areas,the accuracy of manual recognition is low and difficult to be quantified.In recent years,the application of artificial intelligence technology in the agricultural field has gradually increased.In this paper,based on the DeepLabV2 model,we proposed a Zanthoxylum rust image segmentation model based on the FASPP module and enhanced features of rust areas.This paper constructed a fine-grained Zanthoxylum rust image dataset.In this dataset,the Zanthoxylum rust image was segmented and labeled according to leaves,spore piles,and brown lesions.The experimental results showed that the Zanthoxylum rust image segmentation method proposed in this paper was effective.The segmentation accuracy rates of leaves,spore piles and brown lesions reached 99.66%,85.16%and 82.47%respectively.MPA reached 91.80%,and MIoU reached 84.99%.At the same time,the proposed image segmentation model also had good efficiency,which can process 22 images per minute.This article provides an intelligent method for efficiently and accurately recognizing the degree of infection of Zanthoxylum rust.展开更多
Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity ...Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity and high memory consumption in self-attention,Mamba-UNet achieves efficient global modeling through linear-complexity state space modeling.This paper proposes TriLVM-UNet,a lightweight three-stage architecture that integrates parameter-efficient VMamba blocks and enhances cross-stage feature interaction via an improved skip-attention bridge(SAB)module inspired by UltraLight VM-UNet.The model incorporates a Lightweight Vision Mamba(LVM)layer for high-resolution feature extraction,alongside multi-scale dilated convolution(MSDC)and convolutional block attention module(CBAM)for enhanced feature fusion.Evaluated on the 3D ACDC dataset against six baseline models,TriLVM-UNet achieves 98.57%accuracy.The GitHub repository is available at:http://gffzz188fe103f8f1460asbc55wqwfwqno6ovq.ffgz.tsg.suse.edu.cn/730432ch/TriLVM-UNet.展开更多
Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to ...Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to the inability to effectively capture global information from images,CNNs can easily lead to loss of contours and textures in segmentation results.Notice that the transformer model can effectively capture the properties of long-range dependencies in the image,and furthermore,combining the CNN and the transformer can effectively extract local details and global contextual features of the image.Motivated by this,we propose a multi-branch and multi-scale attention network(M2ANet)for medical image segmentation,whose architecture consists of three components.Specifically,in the first component,we construct an adaptive multi-branch patch module for parallel extraction of image features to reduce information loss caused by downsampling.In the second component,we apply residual block to the well-known convolutional block attention module to enhance the network’s ability to recognize important features of images and alleviate the phenomenon of gradient vanishing.In the third component,we design a multi-scale feature fusion module,in which we adopt adaptive average pooling and position encoding to enhance contextual features,and then multi-head attention is introduced to further enrich feature representation.Finally,we validate the effectiveness and feasibility of the proposed M2ANet method through comparative experiments on four benchmark medical image segmentation datasets,particularly in the context of preserving contours and textures.展开更多
Organoids possess immense potential for unraveling the intricate functions of human tissues and facilitating preclinical disease treatment.Their applications span from high-throughput drug screening to the modeling of...Organoids possess immense potential for unraveling the intricate functions of human tissues and facilitating preclinical disease treatment.Their applications span from high-throughput drug screening to the modeling of complex diseases,with some even achieving clinical translation.Changes in the overall size,shape,boundary,and other morphological features of organoids provide a noninvasive method for assessing organoid drug sensitivity.However,the precise segmentation of organoids in bright-field microscopy images is made difficult by the complexity of the organoid morphology and interference,including overlapping organoids,bubbles,dust particles,and cell fragments.This paper introduces the precision organoid segmentation technique(POST),which is a deep-learning algorithm for segmenting challenging organoids under simple bright-field imaging conditions.Unlike existing methods,POST accurately segments each organoid and eliminates various artifacts encountered during organoid culturing and imaging.Furthermore,it is sensitive to and aligns with measurements of organoid activity in drug sensitivity experiments.POST is expected to be a valuable tool for drug screening using organoids owing to its capability of automatically and rapidly eliminating interfering substances and thereby streamlining the organoid analysis and drug screening process.展开更多
Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between ...Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between fragmented residue and soil,compounded by variable illumination and shadows in field imagery,often lead to poor segmentation performance.To overcome these limitations,we introduce RCTUnet,a novel deep learning architecture designed for robust crop-residue-soil segmentation and precise CRC estimation.RCTUnet’s architecture synergistically integrates three key components:(1)a ResNet50 backbone for deep,multi-scale feature extraction;(2)a convolutional block attention module(CBAM)to adaptively focus on salient residue features across both channel and spatial dimensions;and(3)a transformer-based global context fusion module(GCFM)to model long-range spatial dependencies,which is critical for interpreting heterogeneous residue patterns.We evaluated RCTUnet on a dataset of 1220 field-acquired images spanning four typical crop rotations.Experimental results show that,compared to traditional models:(1)RCTUnet achieves significantly higher crop-residue-soil segmentation accuracy than classic models including Unet,Unet++,DeepLabV3,segmentation network(SegNet),and fully convolutional network(FCN),with improvements of 3.24%,3.42%,4.88%,8.28%,and 6.05%in overall accuracy,respectively;(2)RCTUnet yields superior residue-soil segmentation performance,with increases in residue recall of 7.67%,7.37%,14.09%,27.05%,and 16.91%,respectively;(3)RCTUnet shows enhanced CRC estimation accuracy,achieving a root mean square error(RMSE)of 4.875,representing a 45.5%improvement over Unet(RMSE=8.941).These results demonstrate the efficacy of our hybrid approach,which combines deep hierarchical features,dual-domain attention,and global context modeling.RCTUnet provides a robust and reliable tool for automated CRC assessment,advancing the capabilities of in-field agricultural monitoring.展开更多
The stability of concrete-rock interfaces is a critical issue in underground engineering.This study investigated the strain localization mechanism and energy evolution of concrete-sandstone specimens containing single...The stability of concrete-rock interfaces is a critical issue in underground engineering.This study investigated the strain localization mechanism and energy evolution of concrete-sandstone specimens containing single and double interfacial cracks at various inclination angles.Acoustic emission(AE)technology and energy theory were used to analyze energy evolution,whereas digital image correlation was employed to examine strain development and fracture mechanisms.A new approach combining digital image processing and custom binarization was introduced to characterize the fractal properties of crack patterns using the box-counting method.Experimental results showed that the AE cumulative energy and stress-strain curves divided the loading process into three stages:crack closure(Ⅰ),stable crack growth(Ⅱ),and rapid crack propagation(Ⅲ).Fractal dimensions were computed for both singleand double-crack specimens using the Otsu method and the proposed binarization technique.The Otsu method yielded values of 1.457,1.482,1.131,1.512,1.489,1.536,1.171,and 1.491,whereas the new method produced higher values—1.6038,1.6643,1.2713,1.6806,1.5594,1.6282,1.2239,and 1.6565—indicating enhanced fractal characteristics.Furthermore,the proposed method detected a three-phase evolution in fractal dimension before failure,which follows an initial increase,a stable period,and a finalrapid rise.These findingsprovide theoretical support for the application of the proposed method in underground engineering.展开更多
The development of oil and gas is constrained by difficulties in dynamically characterizing pore structures.Traditional methods inadequately represent the complex interactions between mineral dissolution,precipitation...The development of oil and gas is constrained by difficulties in dynamically characterizing pore structures.Traditional methods inadequately represent the complex interactions between mineral dissolution,precipitation,and fluid flow.This study addresses these gaps by introducing a Transformer U-Neural Network(TransUNet)for computed tomography(CT)image segmentation.The integrated workflow combines conventional CT(Resolution of 5.4μm)and synchrotron radiation CT(Resolution of0.8μm)for dynamic flooding,imaging,segmentation,and precise 3D pore network extraction,overcoming resolution limits.TransUNet's strong global attention and feature extraction reduce overfitting and deliver high-accuracy segmentation of minerals,pores,and argillaceous microporous networks(AMN),achieving 74.92%intersection over union(IoU)for AMN.A porosity correction method improves conventional CT porosity accuracy to 94%of gas-measured values.Alkaline flooding experiments reveal:(1)initial clay swelling reduces small pore size by~50%as alkaline ions destabilize clay;(2)mineral dissolution,such as dolomite,creates secondary pores,increasing 80μm pores by 1.8 times;(3)silicate dissolution increases porosity and leads to a 93.7%rise in permeability.Clay reorganization enhances the AMN by 46.1%.The pore size distribution shifts to log-no rmal at steady state,and throat connectivity improves flow capacity.This work pioneers Transformer-based CT image segmentation,introduces cross-resolution prediction,and clarifies pore regulation by mineral phase changes,establishing a new paradigm for chemical flooding in sandstone reservoirs.展开更多
High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes an...High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes and wealth of spatial details pose challenges for semantic segmentation.While convolutional neural networks(CNNs)excel at capturing local features,they are limited in modeling long-range dependencies.Conversely,transformers utilize multihead self-attention to integrate global context effectively,but this approach often incurs a high computational cost.This paper proposes a global-local multiscale context network(GLMCNet)to extract both global and local multiscale contextual information from HRSIs.A detail-enhanced filtering module(DEFM)is proposed at the end of the encoder to refine the encoder outputs further,thereby enhancing the key details extracted by the encoder and effectively suppressing redundant information.In addition,a global-local multiscale transformer block(GLMTB)is proposed in the decoding stage to enable the modeling of rich multiscale global and local information.We also design a stair fusion mechanism to transmit deep semantic information from deep to shallow layers progressively.Finally,we propose the semantic awareness enhancement module(SAEM),which further enhances the representation of multiscale semantic features through spatial attention and covariance channel attention.Extensive ablation analyses and comparative experiments were conducted to evaluate the performance of the proposed method.Specifically,our method achieved a mean Intersection over Union(mIoU)of 86.89%on the ISPRS Potsdam dataset and 84.34%on the ISPRS Vaihingen dataset,outperforming existing models such as ABCNet and BANet.展开更多
Models based on U-shaped networks have achieved widespread success in the field of medical image segmentation,but their performance is generally limited by structural bottlenecks in the network.At this stage,feature m...Models based on U-shaped networks have achieved widespread success in the field of medical image segmentation,but their performance is generally limited by structural bottlenecks in the network.At this stage,feature maps experience a sharp decline in spatial resolution due to continuous downsampling,resulting in significant loss of critical boundaries and structural details.Additionally,the local receptive fields of convolutions limit the effective modelling of global context.To address this core issue,we propose a novel enhanced segmentation network called FMTNet.FMTNet fundamentally enhances the expressive power of deep features by integrating an innovative composite enhancement module at the bottleneck of the U-Net.This module consists of three synergistically working submodules:the Fourier spatial fusion module,which introduces a frequency-domain perspective to compensate for and reconstruct high-frequency structural information lost in the spatial domain;the hybrid mamba–transformer module,which efficiently captures cross-regional long-range dependencies to establish global context and the multi-scale context Aggregation module,which fuses features of different scales to adapt to objects of varying sizes.We conducted extensive experiments on multiple public multi-modal datasets,including colonoscopy polyps,dermatoscopy lesions,breast ultrasound and dental X-rays.The results demonstrate that FMTNet comprehensively outperforms SOTA methods across all key metrics,showcasing exceptional segmentation accuracy and generalisation capabilities.Our research study demonstrates that by synergistically enhancing deep features across three dimensions—frequency,global,and multi-scale—FMTNet provides a general and efficient solution to address the bottleneck issues of U-Net,significantly enhancing the accuracy and robustness of medical image segmentation.The source code and pre-trained weights are available at http://gffzz188fe103f8f1460asbc55wqwfwqno6ovq.ffgz.tsg.suse.edu.cn/shiguiling0-has/FMTNet.展开更多
Images taken in dim environments frequently exhibit issues like insufficient brightness,noise,color shifts,and loss of detail.These problems pose significant challenges to dark image enhancement tasks.Current approach...Images taken in dim environments frequently exhibit issues like insufficient brightness,noise,color shifts,and loss of detail.These problems pose significant challenges to dark image enhancement tasks.Current approaches,while effective in global illumination modeling,often struggle to simultaneously suppress noise and preserve structural details,especially under heterogeneous lighting.Furthermore,misalignment between luminance and color channels introduces additional challenges to accurate enhancement.In response to the aforementioned difficulties,we introduce a single-stage framework,M2ATNet,using the multi-scale multi-attention and Transformer architecture.First,to address the problems of texture blurring and residual noise,we design a multi-scale multi-attention denoising module(MMAD),which is applied separately to the luminance and color channels to enhance the structural and texture modeling capabilities.Secondly,to solve the non-alignment problem of the luminance and color channels,we introduce the multi-channel feature fusion Transformer(CFFT)module,which effectively recovers the dark details and corrects the color shifts through cross-channel alignment and deep feature interaction.To guide the model to learn more stably and efficiently,we also fuse multiple types of loss functions to form a hybrid loss term.We extensively evaluate the proposed method on various standard datasets,including LOL-v1,LOL-v2,DICM,LIME,and NPE.Evaluation in terms of numerical metrics and visual quality demonstrate that M2ATNet consistently outperforms existing advanced approaches.Ablation studies further confirm the critical roles played by the MMAD and CFFT modules to detail preservation and visual fidelity under challenging illumination-deficient environments.展开更多
Multilevel image segmentation is a critical task in image analysis,which imposes high requirements on the global search capability and convergence efficiency of segmentation algorithms.In this paper,an improved Artifi...Multilevel image segmentation is a critical task in image analysis,which imposes high requirements on the global search capability and convergence efficiency of segmentation algorithms.In this paper,an improved Artificial Protozoa Optimization algorithm,termed the two-stage Taguchi-assisted Gaussian–Levy Artificial Protozoa Optimization(TGAPO)algorithm,is proposed and applied tomultilevel image segmentation.The proposed algorithm adopts a two-stage evolutionary mechanism.In the first stage,Gaussian perturbation is introduced to enhance local search capability;in the second stage,Levy flight is incorporated to expand the global search range;and finally,the Taguchi strategy is employed to further refine the optimal solution.Consequently,the global optimization performance and robustness of the algorithm are significantly improved.To evaluate the effectiveness of the proposed TGAPO algorithm,comparative experiments are conducted with representative optimization algorithms,including the Grey Wolf Optimizer(GWO)and Particle Swarm Optimization(PSO),in the context ofmultilevel image segmentation.The segmentation quality is assessed using the minimum cross-entropy function as the performance metric.Experimental results demonstrate that the TGAPO algorithm outperforms the comparison algorithms in terms of segmentation accuracy and convergence speed,and exhibits superior stability in high-threshold segmentation tasks.Furthermore,the proposedmethod achieves excellentmulti-threshold segmentation performance for color images and shows strong potential for practical applications.展开更多
More accurate segmentation of skin cancers in dermoscopy images is crucial for clinical treatment.However,the prevalence of interfering noise in dermoscopy images poses a challenge to its accurate segmentation.For thi...More accurate segmentation of skin cancers in dermoscopy images is crucial for clinical treatment.However,the prevalence of interfering noise in dermoscopy images poses a challenge to its accurate segmentation.For this reason,this paper proposes an improved GLF-Segformer to improve segmentation.The model adds polarized self-attention(PSA)module and R-convolution and attention fusion module(R-CAFM)to the Segformer’s encoder to enhance the ability to capture local information and facilitate the effective fusion of local and global information.The decoder employs an innovative two-stage hybrid up-sampling to effectively reduce information loss.In addition,a new hybrid loss function is designed to further improve the segmentation accuracy of the model at complex boundaries.The experimental results show that GLF-Segformer achieves 90.73%and 89.85%mean intersection over union(mIoU)on two standard datasets,ISIC2017 and ISIC2018,respectively,and exhibits better segmentation performance compared to other comparison algorithms.展开更多
Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Althoug...Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts.展开更多
In order to address the challenges associated with poor semantic segmentation results of classical semantic segmentation networks in high-resolution remote sensing images,limited performance in complex scenes,a large ...In order to address the challenges associated with poor semantic segmentation results of classical semantic segmentation networks in high-resolution remote sensing images,limited performance in complex scenes,a large number of network parameters,and high training costs,this study proposes an efficient segmentation method for high-resolution remote sensing images based on an improved DeepLabv3+approach.The method focuses on three key aspects:reducing the number of network parameters,minimizing computation volume,and enhancing performance.First,the proposed method replaces the original DeepLabv3+backbone network Xception,which is computationally heavy,with the lighter MobileNetV2 network for feature extraction.This substitution helps reduce the number of network parameters while maintaining effective feature extraction.Second,a lightweight convolutional block attention module(CBAM)is added after the feature extraction module to enhance the network’s feature extraction capability.The inclusion of CBAM further reduces the number of network parameters.Last,coordinate attention is introduced after the shallow features obtained from the feature extraction module.This addition allows the network to focus more on relevant features in the image,while disregarding irrelevant background information.Experimental results demonstrate the effectiveness of the proposed method.In the segmentation task of the high-resolution image dataset,the method achieves a mean intersection over union(mIoU)of 75.33%.This result surpasses mainstream semantic segmentation networks such as SegNet,PSPNet,and U-Net by 12.49%,3.16%,and 1.62%respectively.Furthermore,the proposed model has a relatively low number of network parameters,with only 6.02×106 parameters,and a computation volume of 26.45 GFLOPs.This balance between computational efficiency and segmentation accuracy makes the model highly valuable for edge computing applications.展开更多
Microscopy imaging is fundamental in analyzing bacterial morphology and dynamics,offering critical insights into bacterial physiology and pathogenicity.Image segmentation techniques enable quantitative analysis of bac...Microscopy imaging is fundamental in analyzing bacterial morphology and dynamics,offering critical insights into bacterial physiology and pathogenicity.Image segmentation techniques enable quantitative analysis of bacterial structures,facilitating precise measurement of morphological variations and population behaviors at single-cell resolution.This paper reviews advancements in bacterial image segmentation,emphasizing the shift from traditional thresholding and watershed methods to deep learning-driven approaches.Convolutional neural networks(CNNs),U-Net architectures,and three-dimensional(3D)frameworks excel at segmenting dense biofilms and resolving antibiotic-induced morphological changes.These methods combine automated feature extraction with physics-informed postprocessing.Despite progress,challenges persist in computational efficiency,cross-species generalizability,and integration with multimodal experimental workflows.Future progress will depend on improving model robustness across species and imaging modalities,integrating multimodal data for phenotype-function mapping,and developing standard pipelines that link computational tools with clinical diagnostics.These innovations will expand microbial phenotyping beyond structural analysis,enabling deeper insights into bacterial physiology and ecological interactions.展开更多
Background:Diabetic macular edema is a prevalent retinal condition and a leading cause of visual impairment among diabetic patients’Early detection of affected areas is beneficial for effective diagnosis and treatmen...Background:Diabetic macular edema is a prevalent retinal condition and a leading cause of visual impairment among diabetic patients’Early detection of affected areas is beneficial for effective diagnosis and treatment.Traditionally,diagnosis relies on optical coherence tomography imaging technology interpreted by ophthalmologists.However,this manual image interpretation is often slow and subjective.Therefore,developing automated segmentation for macular edema images is essential to enhance to improve the diagnosis efficiency and accuracy.Methods:In order to improve clinical diagnostic efficiency and accuracy,we proposed a SegNet network structure integrated with a convolutional block attention module(CBAM).This network introduces a multi-scale input module,the CBAM attention mechanism,and jump connection.The multi-scale input module enhances the network’s perceptual capabilities,while the lightweight CBAM effectively fuses relevant features across channels and spatial dimensions,allowing for better learning of varying information levels.Results:Experimental results demonstrate that the proposed network achieves an IoU of 80.127%and an accuracy of 99.162%.Compared to the traditional segmentation network,this model has fewer parameters,faster training and testing speed,and superior performance on semantic segmentation tasks,indicating its highly practical applicability.Conclusion:The C-SegNet proposed in this study enables accurate segmentation of Diabetic macular edema lesion images,which facilitates quicker diagnosis for healthcare professionals.展开更多
Despite its remarkable performance on natural images,the segment anything model(SAM)lacks domain-specific information in medical imaging.and faces the challenge of losing local multi-scale information in the encoding ...Despite its remarkable performance on natural images,the segment anything model(SAM)lacks domain-specific information in medical imaging.and faces the challenge of losing local multi-scale information in the encoding phase.This paper presents a medical image segmentation model based on SAM with a local multi-scale feature encoder(LMSFE-SAM)to address the issues above.Firstly,based on the SAM,a local multi-scale feature encoder is introduced to improve the representation of features within local receptive field,thereby supplying the Vision Transformer(ViT)branch in SAM with enriched local multi-scale contextual information.At the same time,a multiaxial Hadamard product module(MHPM)is incorporated into the local multi-scale feature encoder in a lightweight manner to reduce the quadratic complexity and noise interference.Subsequently,a cross-branch balancing adapter is designed to balance the local and global information between the local multi-scale feature encoder and the ViT encoder in SAM.Finally,to obtain smaller input image size and to mitigate overlapping in patch embeddings,the size of the input image is reduced from 1024×1024 pixels to 256×256 pixels,and a multidimensional information adaptation component is developed,which includes feature adapters,position adapters,and channel-spatial adapters.This component effectively integrates the information from small-sized medical images into SAM,enhancing its suitability for clinical deployment.The proposed model demonstrates an average enhancement ranging from 0.0387 to 0.3191 across six objective evaluation metrics on BUSI,DDTI,and TN3K datasets compared to eight other representative image segmentation models.This significantly enhances the performance of the SAM on medical images,providing clinicians with a powerful tool in clinical diagnosis.展开更多
Medical image segmentation is of critical importance in the domain of contemporary medical imaging.However,U-Net and its variants exhibit limitations in capturing complex nonlinear patterns and global contextual infor...Medical image segmentation is of critical importance in the domain of contemporary medical imaging.However,U-Net and its variants exhibit limitations in capturing complex nonlinear patterns and global contextual information.Although the subsequent U-KAN model enhances nonlinear representation capabilities,it still faces challenges such as gradient vanishing during deep network training and spatial detail loss during feature downsampling,resulting in insufficient segmentation accuracy for edge structures and minute lesions.To address these challenges,this paper proposes the RE-UKAN model,which innovatively improves upon U-KAN.Firstly,a residual network is introduced into the encoder to effectively mitigate gradient vanishing through cross-layer identity mappings,thus enhancing modelling capabilities for complex pathological structures.Secondly,Efficient Local Attention(ELA)is integrated to suppress spatial detail loss during downsampling,thereby improving the perception of edge structures and minute lesions.Experimental results on four public datasets demonstrate that RE-UKAN outperforms existing medical image segmentation methods across multiple evaluation metrics,with particularly outstanding performance on the TN-SCUI 2020 dataset,achieving IoU of 88.18%and Dice of 93.57%.Compared to the baseline model,it achieves improvements of 3.05%and 1.72%,respectively.These results fully demonstrate RE-UKAN’s superior detail retention capability and boundary recognition accuracy in complex medical image segmentation tasks,providing a reliable solution for clinical precision segmentation.展开更多
Vision-language segmentation models (VLSMs) are effective in medical image segmentation tasks. However, a major limitation of these models is their dependence on manually crafted textual inputs. Studies have used visu...Vision-language segmentation models (VLSMs) are effective in medical image segmentation tasks. However, a major limitation of these models is their dependence on manually crafted textual inputs. Studies have used visual question answering to semiautomatically generate textual information. However, these methods encounter challenges such as error accumulation. Herein, we propose a method to learn conceptual text prompts directly from visual regions of interest (ROIs) for facilitating medical image segmentation. We extracted textual conceptual attributes from ROIs using a large multimodal model to derive coarse real-text prompts. A text latent space transformation module accepted the ROI images as input for generating fine-grained pseudo-text prompts to compensate for the lack of image detail perception in the abovementioned real-text prompts. These prompts were encoded into a unified text embedding. Thereafter, we applied a self-adding noise knowledge distillation method to transfer the knowledge from text embedding to the class token of the image encoder, enabling direct text-guided inference during testing while reducing error accumulation. Our approach minimized the need for manual prompt design by leveraging explicit discrete and implicit continuous text prompts to effectively guide visual segmentation. Extensive evaluation across 13 medical image segmentation datasets demonstrated that our model outperformed the state-of-the-art VLSMs and vision-based segmentation models, exhibiting superior segmentation accuracy.展开更多
U-Net,a fully convolutional neural network(FCNN)with U-shaped features,has demonstrated significant success in biomedical image segmentation.However,the locality of convolution operations in the U-Net limits its abili...U-Net,a fully convolutional neural network(FCNN)with U-shaped features,has demonstrated significant success in biomedical image segmentation.However,the locality of convolution operations in the U-Net limits its ability to learn long-range dependencies.Transformers,originally developed for natural language processing,have recently been adapted for image segmentation because of their global self-attention mechanisms.Inspired by the long-range feature learning capability of transformers,we propose Dense-Transformer(DenT),an architecture designed for volumetric microscopy image segmentation.DenT incorporates transformers as encoders within each convolutional layer to capture global contextual information.Additionally,dense skip connections at multiple resolutions enhance feature propagation,enabling precise localization.We evaluated DenT on mitochondrial segmentation using our confocal microscopy dataset and a public fluorescence microscope dataset from the Allen Institute for Cell Science.The experimental results demonstrate that DenT incrementally improves the segmentation of mitochondria and mitochondrial DNA substructures from transmitted light microscopy images.DenT offers a tool for visualization,measurement,and analysis of mitochondrial morphology and mitochondrial DNA in label-free microscopy.展开更多
基金This work was supported by Natural Science Foundation of China(Grant No.62071098)Sichuan Science and Technology Program(Grant Nos.2019YFG0191,2021YFG0307)Sichuan Zizhou Agricultural Science and Technology Co.,Ltd.project:Internet+smart Zanthoxylum planting weather risk warning system.
摘要Zanthoxylum bungeanum Maxim,generally called prickly ash,is widely grown in China.Zanthoxylum rust is the main disease affecting the growth and quality of Zanthoxylum.Traditional method for recognizing the degree of infection of Zanthoxylum rust mainly rely on manual experience.Due to the complex colors and shapes of rust areas,the accuracy of manual recognition is low and difficult to be quantified.In recent years,the application of artificial intelligence technology in the agricultural field has gradually increased.In this paper,based on the DeepLabV2 model,we proposed a Zanthoxylum rust image segmentation model based on the FASPP module and enhanced features of rust areas.This paper constructed a fine-grained Zanthoxylum rust image dataset.In this dataset,the Zanthoxylum rust image was segmented and labeled according to leaves,spore piles,and brown lesions.The experimental results showed that the Zanthoxylum rust image segmentation method proposed in this paper was effective.The segmentation accuracy rates of leaves,spore piles and brown lesions reached 99.66%,85.16%and 82.47%respectively.MPA reached 91.80%,and MIoU reached 84.99%.At the same time,the proposed image segmentation model also had good efficiency,which can process 22 images per minute.This article provides an intelligent method for efficiently and accurately recognizing the degree of infection of Zanthoxylum rust.
基金The Key Project of Ningxia Natural Science Foundation under grant 2025AAC020006the Regional Program of National Natural Science Foundation of China under grant 62561002The Shaanxi University of Technology Foundation Project under grant SLGRCQD2137.
摘要Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity and high memory consumption in self-attention,Mamba-UNet achieves efficient global modeling through linear-complexity state space modeling.This paper proposes TriLVM-UNet,a lightweight three-stage architecture that integrates parameter-efficient VMamba blocks and enhances cross-stage feature interaction via an improved skip-attention bridge(SAB)module inspired by UltraLight VM-UNet.The model incorporates a Lightweight Vision Mamba(LVM)layer for high-resolution feature extraction,alongside multi-scale dilated convolution(MSDC)and convolutional block attention module(CBAM)for enhanced feature fusion.Evaluated on the 3D ACDC dataset against six baseline models,TriLVM-UNet achieves 98.57%accuracy.The GitHub repository is available at:http://gffzz188fe103f8f1460asbc55wqwfwqno6ovq.ffgz.tsg.suse.edu.cn/730432ch/TriLVM-UNet.
基金supported by the Natural Science Foundation of the Anhui Higher Education Institutions of China(Grant Nos.2023AH040149 and 2024AH051915)the Anhui Provincial Natural Science Foundation(Grant No.2208085MF168)+1 种基金the Science and Technology Innovation Tackle Plan Project of Maanshan(Grant No.2024RGZN001)the Scientific Research Fund Project of Anhui Medical University(Grant No.2023xkj122).
摘要Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to the inability to effectively capture global information from images,CNNs can easily lead to loss of contours and textures in segmentation results.Notice that the transformer model can effectively capture the properties of long-range dependencies in the image,and furthermore,combining the CNN and the transformer can effectively extract local details and global contextual features of the image.Motivated by this,we propose a multi-branch and multi-scale attention network(M2ANet)for medical image segmentation,whose architecture consists of three components.Specifically,in the first component,we construct an adaptive multi-branch patch module for parallel extraction of image features to reduce information loss caused by downsampling.In the second component,we apply residual block to the well-known convolutional block attention module to enhance the network’s ability to recognize important features of images and alleviate the phenomenon of gradient vanishing.In the third component,we design a multi-scale feature fusion module,in which we adopt adaptive average pooling and position encoding to enhance contextual features,and then multi-head attention is introduced to further enrich feature representation.Finally,we validate the effectiveness and feasibility of the proposed M2ANet method through comparative experiments on four benchmark medical image segmentation datasets,particularly in the context of preserving contours and textures.
基金supported by the National Key R&D Program of China(No.2022YFC2504403)the National Natural Science Foundation of China(No.62172202)+1 种基金the Experiment Project of China Manned Space Program(No.HYZHXM01019)the Fundamental Research Funds for the Central Universities from Southeast University(No.3207032101C3)。
摘要Organoids possess immense potential for unraveling the intricate functions of human tissues and facilitating preclinical disease treatment.Their applications span from high-throughput drug screening to the modeling of complex diseases,with some even achieving clinical translation.Changes in the overall size,shape,boundary,and other morphological features of organoids provide a noninvasive method for assessing organoid drug sensitivity.However,the precise segmentation of organoids in bright-field microscopy images is made difficult by the complexity of the organoid morphology and interference,including overlapping organoids,bubbles,dust particles,and cell fragments.This paper introduces the precision organoid segmentation technique(POST),which is a deep-learning algorithm for segmenting challenging organoids under simple bright-field imaging conditions.Unlike existing methods,POST accurately segments each organoid and eliminates various artifacts encountered during organoid culturing and imaging.Furthermore,it is sensitive to and aligns with measurements of organoid activity in drug sensitivity experiments.POST is expected to be a valuable tool for drug screening using organoids owing to its capability of automatically and rapidly eliminating interfering substances and thereby streamlining the organoid analysis and drug screening process.
基金supported by the National Natural Science Foundation of China(No.42101362)the Natural Science Foundation of Henan Province(No.252300421158)+1 种基金the Shenzhen Science and Technology Program(No.JCYJ20220530162001003)the Science and Technology Development Program of Henan Province(No.242300421639),China。
摘要Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between fragmented residue and soil,compounded by variable illumination and shadows in field imagery,often lead to poor segmentation performance.To overcome these limitations,we introduce RCTUnet,a novel deep learning architecture designed for robust crop-residue-soil segmentation and precise CRC estimation.RCTUnet’s architecture synergistically integrates three key components:(1)a ResNet50 backbone for deep,multi-scale feature extraction;(2)a convolutional block attention module(CBAM)to adaptively focus on salient residue features across both channel and spatial dimensions;and(3)a transformer-based global context fusion module(GCFM)to model long-range spatial dependencies,which is critical for interpreting heterogeneous residue patterns.We evaluated RCTUnet on a dataset of 1220 field-acquired images spanning four typical crop rotations.Experimental results show that,compared to traditional models:(1)RCTUnet achieves significantly higher crop-residue-soil segmentation accuracy than classic models including Unet,Unet++,DeepLabV3,segmentation network(SegNet),and fully convolutional network(FCN),with improvements of 3.24%,3.42%,4.88%,8.28%,and 6.05%in overall accuracy,respectively;(2)RCTUnet yields superior residue-soil segmentation performance,with increases in residue recall of 7.67%,7.37%,14.09%,27.05%,and 16.91%,respectively;(3)RCTUnet shows enhanced CRC estimation accuracy,achieving a root mean square error(RMSE)of 4.875,representing a 45.5%improvement over Unet(RMSE=8.941).These results demonstrate the efficacy of our hybrid approach,which combines deep hierarchical features,dual-domain attention,and global context modeling.RCTUnet provides a robust and reliable tool for automated CRC assessment,advancing the capabilities of in-field agricultural monitoring.
基金supported by the National Natural Science Foundation of China(Grant Nos.52264006 and 52364004)Guizhou Provincial Basic Research Program(Natural Science)(Grant No.QianKeHe Basic-ZK[2025]general program 630).
摘要The stability of concrete-rock interfaces is a critical issue in underground engineering.This study investigated the strain localization mechanism and energy evolution of concrete-sandstone specimens containing single and double interfacial cracks at various inclination angles.Acoustic emission(AE)technology and energy theory were used to analyze energy evolution,whereas digital image correlation was employed to examine strain development and fracture mechanisms.A new approach combining digital image processing and custom binarization was introduced to characterize the fractal properties of crack patterns using the box-counting method.Experimental results showed that the AE cumulative energy and stress-strain curves divided the loading process into three stages:crack closure(Ⅰ),stable crack growth(Ⅱ),and rapid crack propagation(Ⅲ).Fractal dimensions were computed for both singleand double-crack specimens using the Otsu method and the proposed binarization technique.The Otsu method yielded values of 1.457,1.482,1.131,1.512,1.489,1.536,1.171,and 1.491,whereas the new method produced higher values—1.6038,1.6643,1.2713,1.6806,1.5594,1.6282,1.2239,and 1.6565—indicating enhanced fractal characteristics.Furthermore,the proposed method detected a three-phase evolution in fractal dimension before failure,which follows an initial increase,a stable period,and a finalrapid rise.These findingsprovide theoretical support for the application of the proposed method in underground engineering.
基金supported by Program for Young Talents of Basic Research in Universities of Heilongjiang Province(YQJH2024036)Collaborative Innovation Projects of“Double First-class”Disciplines in Heilongjiang Province(LJGXCG2024-P20)Outstanding Talent Cultivation Foundation of Northeast Petroleum University(SJQHB202004)。
摘要The development of oil and gas is constrained by difficulties in dynamically characterizing pore structures.Traditional methods inadequately represent the complex interactions between mineral dissolution,precipitation,and fluid flow.This study addresses these gaps by introducing a Transformer U-Neural Network(TransUNet)for computed tomography(CT)image segmentation.The integrated workflow combines conventional CT(Resolution of 5.4μm)and synchrotron radiation CT(Resolution of0.8μm)for dynamic flooding,imaging,segmentation,and precise 3D pore network extraction,overcoming resolution limits.TransUNet's strong global attention and feature extraction reduce overfitting and deliver high-accuracy segmentation of minerals,pores,and argillaceous microporous networks(AMN),achieving 74.92%intersection over union(IoU)for AMN.A porosity correction method improves conventional CT porosity accuracy to 94%of gas-measured values.Alkaline flooding experiments reveal:(1)initial clay swelling reduces small pore size by~50%as alkaline ions destabilize clay;(2)mineral dissolution,such as dolomite,creates secondary pores,increasing 80μm pores by 1.8 times;(3)silicate dissolution increases porosity and leads to a 93.7%rise in permeability.Clay reorganization enhances the AMN by 46.1%.The pore size distribution shifts to log-no rmal at steady state,and throat connectivity improves flow capacity.This work pioneers Transformer-based CT image segmentation,introduces cross-resolution prediction,and clarifies pore regulation by mineral phase changes,establishing a new paradigm for chemical flooding in sandstone reservoirs.
基金provided by the Science Research Project of Hebei Education Department under grant No.BJK2024115.
摘要High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes and wealth of spatial details pose challenges for semantic segmentation.While convolutional neural networks(CNNs)excel at capturing local features,they are limited in modeling long-range dependencies.Conversely,transformers utilize multihead self-attention to integrate global context effectively,but this approach often incurs a high computational cost.This paper proposes a global-local multiscale context network(GLMCNet)to extract both global and local multiscale contextual information from HRSIs.A detail-enhanced filtering module(DEFM)is proposed at the end of the encoder to refine the encoder outputs further,thereby enhancing the key details extracted by the encoder and effectively suppressing redundant information.In addition,a global-local multiscale transformer block(GLMTB)is proposed in the decoding stage to enable the modeling of rich multiscale global and local information.We also design a stair fusion mechanism to transmit deep semantic information from deep to shallow layers progressively.Finally,we propose the semantic awareness enhancement module(SAEM),which further enhances the representation of multiscale semantic features through spatial attention and covariance channel attention.Extensive ablation analyses and comparative experiments were conducted to evaluate the performance of the proposed method.Specifically,our method achieved a mean Intersection over Union(mIoU)of 86.89%on the ISPRS Potsdam dataset and 84.34%on the ISPRS Vaihingen dataset,outperforming existing models such as ABCNet and BANet.
基金funded by UKRI(Grants EP/W020408/1 and RS718)through Doctoral Training Centre at Swansea UniversityQilu Medical Talent Cultivation Project of Shandong Health Commission(Grant[2023]78)National Natural Science Foundation of China(Grant 82405459)。
摘要Models based on U-shaped networks have achieved widespread success in the field of medical image segmentation,but their performance is generally limited by structural bottlenecks in the network.At this stage,feature maps experience a sharp decline in spatial resolution due to continuous downsampling,resulting in significant loss of critical boundaries and structural details.Additionally,the local receptive fields of convolutions limit the effective modelling of global context.To address this core issue,we propose a novel enhanced segmentation network called FMTNet.FMTNet fundamentally enhances the expressive power of deep features by integrating an innovative composite enhancement module at the bottleneck of the U-Net.This module consists of three synergistically working submodules:the Fourier spatial fusion module,which introduces a frequency-domain perspective to compensate for and reconstruct high-frequency structural information lost in the spatial domain;the hybrid mamba–transformer module,which efficiently captures cross-regional long-range dependencies to establish global context and the multi-scale context Aggregation module,which fuses features of different scales to adapt to objects of varying sizes.We conducted extensive experiments on multiple public multi-modal datasets,including colonoscopy polyps,dermatoscopy lesions,breast ultrasound and dental X-rays.The results demonstrate that FMTNet comprehensively outperforms SOTA methods across all key metrics,showcasing exceptional segmentation accuracy and generalisation capabilities.Our research study demonstrates that by synergistically enhancing deep features across three dimensions—frequency,global,and multi-scale—FMTNet provides a general and efficient solution to address the bottleneck issues of U-Net,significantly enhancing the accuracy and robustness of medical image segmentation.The source code and pre-trained weights are available at http://gffzz188fe103f8f1460asbc55wqwfwqno6ovq.ffgz.tsg.suse.edu.cn/shiguiling0-has/FMTNet.
基金funded by the National Natural Science Foundation of China,grant numbers 52374156 and 62476005。
摘要Images taken in dim environments frequently exhibit issues like insufficient brightness,noise,color shifts,and loss of detail.These problems pose significant challenges to dark image enhancement tasks.Current approaches,while effective in global illumination modeling,often struggle to simultaneously suppress noise and preserve structural details,especially under heterogeneous lighting.Furthermore,misalignment between luminance and color channels introduces additional challenges to accurate enhancement.In response to the aforementioned difficulties,we introduce a single-stage framework,M2ATNet,using the multi-scale multi-attention and Transformer architecture.First,to address the problems of texture blurring and residual noise,we design a multi-scale multi-attention denoising module(MMAD),which is applied separately to the luminance and color channels to enhance the structural and texture modeling capabilities.Secondly,to solve the non-alignment problem of the luminance and color channels,we introduce the multi-channel feature fusion Transformer(CFFT)module,which effectively recovers the dark details and corrects the color shifts through cross-channel alignment and deep feature interaction.To guide the model to learn more stably and efficiently,we also fuse multiple types of loss functions to form a hybrid loss term.We extensively evaluate the proposed method on various standard datasets,including LOL-v1,LOL-v2,DICM,LIME,and NPE.Evaluation in terms of numerical metrics and visual quality demonstrate that M2ATNet consistently outperforms existing advanced approaches.Ablation studies further confirm the critical roles played by the MMAD and CFFT modules to detail preservation and visual fidelity under challenging illumination-deficient environments.
摘要Multilevel image segmentation is a critical task in image analysis,which imposes high requirements on the global search capability and convergence efficiency of segmentation algorithms.In this paper,an improved Artificial Protozoa Optimization algorithm,termed the two-stage Taguchi-assisted Gaussian–Levy Artificial Protozoa Optimization(TGAPO)algorithm,is proposed and applied tomultilevel image segmentation.The proposed algorithm adopts a two-stage evolutionary mechanism.In the first stage,Gaussian perturbation is introduced to enhance local search capability;in the second stage,Levy flight is incorporated to expand the global search range;and finally,the Taguchi strategy is employed to further refine the optimal solution.Consequently,the global optimization performance and robustness of the algorithm are significantly improved.To evaluate the effectiveness of the proposed TGAPO algorithm,comparative experiments are conducted with representative optimization algorithms,including the Grey Wolf Optimizer(GWO)and Particle Swarm Optimization(PSO),in the context ofmultilevel image segmentation.The segmentation quality is assessed using the minimum cross-entropy function as the performance metric.Experimental results demonstrate that the TGAPO algorithm outperforms the comparison algorithms in terms of segmentation accuracy and convergence speed,and exhibits superior stability in high-threshold segmentation tasks.Furthermore,the proposedmethod achieves excellentmulti-threshold segmentation performance for color images and shows strong potential for practical applications.
基金supported by the National Natural Science Foundation of China(No.61961037)the Industrial Support Plan of Education Department of Gansu Province(No.2021CYZC-30).
摘要More accurate segmentation of skin cancers in dermoscopy images is crucial for clinical treatment.However,the prevalence of interfering noise in dermoscopy images poses a challenge to its accurate segmentation.For this reason,this paper proposes an improved GLF-Segformer to improve segmentation.The model adds polarized self-attention(PSA)module and R-convolution and attention fusion module(R-CAFM)to the Segformer’s encoder to enhance the ability to capture local information and facilitate the effective fusion of local and global information.The decoder employs an innovative two-stage hybrid up-sampling to effectively reduce information loss.In addition,a new hybrid loss function is designed to further improve the segmentation accuracy of the model at complex boundaries.The experimental results show that GLF-Segformer achieves 90.73%and 89.85%mean intersection over union(mIoU)on two standard datasets,ISIC2017 and ISIC2018,respectively,and exhibits better segmentation performance compared to other comparison algorithms.
摘要Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts.
基金the Sichuan Science and Technology Program of China(No.2021YFG0055)the Enterprise Informatization and Internet of Things Measurement and Control Technology Key Laboratory Project of Sichuan Province Colleges and Universities(No.2022WZJ01)+1 种基金the Natural Science Foundation of Sichuan University of Science&Engineering(No.2020RC32)the 2022 Graduate InnovationFund Project of Sichuan University of Science&Engineering(No.Y2022156)。
摘要In order to address the challenges associated with poor semantic segmentation results of classical semantic segmentation networks in high-resolution remote sensing images,limited performance in complex scenes,a large number of network parameters,and high training costs,this study proposes an efficient segmentation method for high-resolution remote sensing images based on an improved DeepLabv3+approach.The method focuses on three key aspects:reducing the number of network parameters,minimizing computation volume,and enhancing performance.First,the proposed method replaces the original DeepLabv3+backbone network Xception,which is computationally heavy,with the lighter MobileNetV2 network for feature extraction.This substitution helps reduce the number of network parameters while maintaining effective feature extraction.Second,a lightweight convolutional block attention module(CBAM)is added after the feature extraction module to enhance the network’s feature extraction capability.The inclusion of CBAM further reduces the number of network parameters.Last,coordinate attention is introduced after the shallow features obtained from the feature extraction module.This addition allows the network to focus more on relevant features in the image,while disregarding irrelevant background information.Experimental results demonstrate the effectiveness of the proposed method.In the segmentation task of the high-resolution image dataset,the method achieves a mean intersection over union(mIoU)of 75.33%.This result surpasses mainstream semantic segmentation networks such as SegNet,PSPNet,and U-Net by 12.49%,3.16%,and 1.62%respectively.Furthermore,the proposed model has a relatively low number of network parameters,with only 6.02×106 parameters,and a computation volume of 26.45 GFLOPs.This balance between computational efficiency and segmentation accuracy makes the model highly valuable for edge computing applications.
基金financially supported by the Open Project Program of Wuhan National Laboratory for Optoelectronics(No.2022WNLOKF009)the National Natural Science Foundation of China(No.62475216)+2 种基金the Key Research and Development Program of Shaanxi(No.2024GH-ZDXM-37)the Fujian Provincial Natural Science Foundation of China(No.2024J01060)the Startup Program of XMU,and the Fundamental Research Funds for the Central Universities.
摘要Microscopy imaging is fundamental in analyzing bacterial morphology and dynamics,offering critical insights into bacterial physiology and pathogenicity.Image segmentation techniques enable quantitative analysis of bacterial structures,facilitating precise measurement of morphological variations and population behaviors at single-cell resolution.This paper reviews advancements in bacterial image segmentation,emphasizing the shift from traditional thresholding and watershed methods to deep learning-driven approaches.Convolutional neural networks(CNNs),U-Net architectures,and three-dimensional(3D)frameworks excel at segmenting dense biofilms and resolving antibiotic-induced morphological changes.These methods combine automated feature extraction with physics-informed postprocessing.Despite progress,challenges persist in computational efficiency,cross-species generalizability,and integration with multimodal experimental workflows.Future progress will depend on improving model robustness across species and imaging modalities,integrating multimodal data for phenotype-function mapping,and developing standard pipelines that link computational tools with clinical diagnostics.These innovations will expand microbial phenotyping beyond structural analysis,enabling deeper insights into bacterial physiology and ecological interactions.
基金supported by the Guangdong Pharmaceutical University 2024 Higher Education Research Projects(GKP202403,GMP202402)the Guangdong Pharmaceutical University College Students’Innovation and Entrepreneurship Training Programs(Grant No.202504302033,202504302034,202504302036,and 202504302244).
摘要Background:Diabetic macular edema is a prevalent retinal condition and a leading cause of visual impairment among diabetic patients’Early detection of affected areas is beneficial for effective diagnosis and treatment.Traditionally,diagnosis relies on optical coherence tomography imaging technology interpreted by ophthalmologists.However,this manual image interpretation is often slow and subjective.Therefore,developing automated segmentation for macular edema images is essential to enhance to improve the diagnosis efficiency and accuracy.Methods:In order to improve clinical diagnostic efficiency and accuracy,we proposed a SegNet network structure integrated with a convolutional block attention module(CBAM).This network introduces a multi-scale input module,the CBAM attention mechanism,and jump connection.The multi-scale input module enhances the network’s perceptual capabilities,while the lightweight CBAM effectively fuses relevant features across channels and spatial dimensions,allowing for better learning of varying information levels.Results:Experimental results demonstrate that the proposed network achieves an IoU of 80.127%and an accuracy of 99.162%.Compared to the traditional segmentation network,this model has fewer parameters,faster training and testing speed,and superior performance on semantic segmentation tasks,indicating its highly practical applicability.Conclusion:The C-SegNet proposed in this study enables accurate segmentation of Diabetic macular edema lesion images,which facilitates quicker diagnosis for healthcare professionals.
基金supported by Natural Science Foundation Programme of Gansu Province(No.24JRRA231)National Natural Science Foundation of China(No.62061023)Gansu Provincial Science and Technology Plan Key Research and Development Program Project(No.24YFFA024).
摘要Despite its remarkable performance on natural images,the segment anything model(SAM)lacks domain-specific information in medical imaging.and faces the challenge of losing local multi-scale information in the encoding phase.This paper presents a medical image segmentation model based on SAM with a local multi-scale feature encoder(LMSFE-SAM)to address the issues above.Firstly,based on the SAM,a local multi-scale feature encoder is introduced to improve the representation of features within local receptive field,thereby supplying the Vision Transformer(ViT)branch in SAM with enriched local multi-scale contextual information.At the same time,a multiaxial Hadamard product module(MHPM)is incorporated into the local multi-scale feature encoder in a lightweight manner to reduce the quadratic complexity and noise interference.Subsequently,a cross-branch balancing adapter is designed to balance the local and global information between the local multi-scale feature encoder and the ViT encoder in SAM.Finally,to obtain smaller input image size and to mitigate overlapping in patch embeddings,the size of the input image is reduced from 1024×1024 pixels to 256×256 pixels,and a multidimensional information adaptation component is developed,which includes feature adapters,position adapters,and channel-spatial adapters.This component effectively integrates the information from small-sized medical images into SAM,enhancing its suitability for clinical deployment.The proposed model demonstrates an average enhancement ranging from 0.0387 to 0.3191 across six objective evaluation metrics on BUSI,DDTI,and TN3K datasets compared to eight other representative image segmentation models.This significantly enhances the performance of the SAM on medical images,providing clinicians with a powerful tool in clinical diagnosis.
摘要Medical image segmentation is of critical importance in the domain of contemporary medical imaging.However,U-Net and its variants exhibit limitations in capturing complex nonlinear patterns and global contextual information.Although the subsequent U-KAN model enhances nonlinear representation capabilities,it still faces challenges such as gradient vanishing during deep network training and spatial detail loss during feature downsampling,resulting in insufficient segmentation accuracy for edge structures and minute lesions.To address these challenges,this paper proposes the RE-UKAN model,which innovatively improves upon U-KAN.Firstly,a residual network is introduced into the encoder to effectively mitigate gradient vanishing through cross-layer identity mappings,thus enhancing modelling capabilities for complex pathological structures.Secondly,Efficient Local Attention(ELA)is integrated to suppress spatial detail loss during downsampling,thereby improving the perception of edge structures and minute lesions.Experimental results on four public datasets demonstrate that RE-UKAN outperforms existing medical image segmentation methods across multiple evaluation metrics,with particularly outstanding performance on the TN-SCUI 2020 dataset,achieving IoU of 88.18%and Dice of 93.57%.Compared to the baseline model,it achieves improvements of 3.05%and 1.72%,respectively.These results fully demonstrate RE-UKAN’s superior detail retention capability and boundary recognition accuracy in complex medical image segmentation tasks,providing a reliable solution for clinical precision segmentation.
摘要Vision-language segmentation models (VLSMs) are effective in medical image segmentation tasks. However, a major limitation of these models is their dependence on manually crafted textual inputs. Studies have used visual question answering to semiautomatically generate textual information. However, these methods encounter challenges such as error accumulation. Herein, we propose a method to learn conceptual text prompts directly from visual regions of interest (ROIs) for facilitating medical image segmentation. We extracted textual conceptual attributes from ROIs using a large multimodal model to derive coarse real-text prompts. A text latent space transformation module accepted the ROI images as input for generating fine-grained pseudo-text prompts to compensate for the lack of image detail perception in the abovementioned real-text prompts. These prompts were encoded into a unified text embedding. Thereafter, we applied a self-adding noise knowledge distillation method to transfer the knowledge from text embedding to the class token of the image encoder, enabling direct text-guided inference during testing while reducing error accumulation. Our approach minimized the need for manual prompt design by leveraging explicit discrete and implicit continuous text prompts to effectively guide visual segmentation. Extensive evaluation across 13 medical image segmentation datasets demonstrated that our model outperformed the state-of-the-art VLSMs and vision-based segmentation models, exhibiting superior segmentation accuracy.
基金Taiwan University Center for Advanced Computing and Imaging in Biomedicine(NTU-114L900701).
摘要U-Net,a fully convolutional neural network(FCNN)with U-shaped features,has demonstrated significant success in biomedical image segmentation.However,the locality of convolution operations in the U-Net limits its ability to learn long-range dependencies.Transformers,originally developed for natural language processing,have recently been adapted for image segmentation because of their global self-attention mechanisms.Inspired by the long-range feature learning capability of transformers,we propose Dense-Transformer(DenT),an architecture designed for volumetric microscopy image segmentation.DenT incorporates transformers as encoders within each convolutional layer to capture global contextual information.Additionally,dense skip connections at multiple resolutions enhance feature propagation,enabling precise localization.We evaluated DenT on mitochondrial segmentation using our confocal microscopy dataset and a public fluorescence microscope dataset from the Allen Institute for Cell Science.The experimental results demonstrate that DenT incrementally improves the segmentation of mitochondria and mitochondrial DNA substructures from transmitted light microscopy images.DenT offers a tool for visualization,measurement,and analysis of mitochondrial morphology and mitochondrial DNA in label-free microscopy.