Accurate segmentation of digital rock images is essential for characterizing pore-matrix systems and predicting petrophysical properties.However,the diversity of rock textures across different lithologies poses a sign...Accurate segmentation of digital rock images is essential for characterizing pore-matrix systems and predicting petrophysical properties.However,the diversity of rock textures across different lithologies poses a significant challenge for conventional segmentation networks,especially under limited training data.To address this,we introduce DRI-SAM(Digital Rock Image-Segment Anything Model),a hybrid segmentation framework that leverages the powerful visual prior of the Segment Anything Model(SAM)and adapts it to the digital rock domain.Specifically,we apply LoRA-based fine-tuning to SAM’s image encoder to better capture rock-specific microstructures,while U-Net is employed to generate prompt points,guiding SAM toward accurate pore-matrix delineation.This approach retains the encoder’s representational power while allowing domain-specific adaptation via LoRA,enabling effective cross-domain generalization under limited training data.The model is trained exclusively on 200 annotated images of Bentheimer sandstone,covering two distinct voxel resolutions,and is evaluated on digital rock images of varying lithologies,resolutions and imaging modalities.The results confirm that DRI-SAM achieves accurate segmentation on both sandstone and more challenging carbonate samples,including synthetic and SEM images,without additional retraining or parameter adjustments.Compared to DeepLabV3+and the only LoRA-tuned SAM,DRI-SAM demonstrates superior performance under limited supervision,highlighting its strong generalization and practical value in digital rock image analysis.Moreover,the findings suggest that foundation models like SAM,when properly adapted,also hold great promise for broader geoscientific imaging tasks.展开更多
Rock fragment size distribution(FSD)plays an important role in various engineering applications,such as mining,tunnelling,and other underground construction scenarios.While vision-based deep learning approaches have b...Rock fragment size distribution(FSD)plays an important role in various engineering applications,such as mining,tunnelling,and other underground construction scenarios.While vision-based deep learning approaches have been increasingly applied to FSD analysis,they are often case-specific,showing limited cross-site generalization despite their accuracy.To address these challenges,FragSAM,an end-to-end,fully automated framework is proposed for near real-time rock fragment segmentation and FSD analysis across diverse engineering environments.FragSAM integrates the generalization power of Segment Anything Model(SAM)with a context-aware prompting mechanism and lightweight architecture for efficient dense fragment segmentation.In Stage 1,an enhanced SAM automatically generates high-quality annotations,which are used to train a modified CenterNet for precise centroid prediction.In Stage 2,these centroids serve as prompts for EdgeSAM,a lightweight SAM variant optimized for real-time inference.This two-stage design eliminates dense grid prompting and reduces reliance on heavy postprocessing,enabling efficient and scalable segmentation.Experimental results show that FragSAM achieves competitive segmentation performance with significantly lower latency and model complexity compared to existing SAM-based methods.In comparison with supervised learning approaches,it also demonstrates superior generalization and performs better in low-quality or unseen scenarios.Furthermore,case studies on blasting fragmentation,TBM muck,and coastal rock surfaces confirm its robustness and seamless cross-site adaptability,requiring no tuning or retraining,making it highly practical for on-site applications.展开更多
Locked segments are high-strength structural elements in fault zones that release significantseismic energy during earthquakes.In fracture mechanics,they act as high-stress concentration patches(asperities)where ruptu...Locked segments are high-strength structural elements in fault zones that release significantseismic energy during earthquakes.In fracture mechanics,they act as high-stress concentration patches(asperities)where rupture initiates.The progressive failure of locked segments along faults plays a crucial role in the energy partition of earthquakes.The impact of locked segments on the near-fielddeformation and nucleation of faults,however,remains poorly understood.In this study,rock-like materials with pre-manufactured strike-slip faults containing various locked segments lengths under uniaxial stress.The mechanical properties,local deformation fields,and slip displacement rates during the uniaxial loading of the models were quantified.Results indicate that the uniaxial compressive strength and elastic modulus of the system peak once the ratio of locked segment to fault length is approximately 0.6.Meanwhile,the resistance of the models to deformation increased,and the failure mode transformed from shear failure to tensile failure.Under loading,compression and dilatation quadrants were formed on both sides of the fault.Large-scale fractures dominate the dilatation quadrants,and the degree of deformation disturbance in this region was significantly higher than that in the compression quadrants.With increasing locked segment length,the amplitude of deformation perturbations decreased after the peak strength.Shorter locked segments were more susceptible to deformation and failure.In the fracture evolution process,a relationship between the stress deflectionangle and the displacement rate was found,which is empirically described by an exponential function.These findings clarify geological structures failure mechanisms and support seismic hazard assessment for strike-slip earthquake regions.展开更多
The northern segment of the North-South Seismic Belt is characterized by intense crustal deformation,well-developed active tectonics,and frequent occurrences of strong earthquakes.Therefore,conducting a Probabilistic ...The northern segment of the North-South Seismic Belt is characterized by intense crustal deformation,well-developed active tectonics,and frequent occurrences of strong earthquakes.Therefore,conducting a Probabilistic Seismic Hazard Analysis(PSHA)for this region is of significant importance for supporting seismic fortification in major engineering projects and formulating disaster prevention and mitigation policies.In this study,a composite seismic source model was constructed by integrating data on historical earthquakes,active faults,and paleoseismicity.Furthermore,a logic tree framework was employed to quantify epistemic uncertainties,enabling a systematic seismic hazard assessment of the region.To more accurately characterize the spatial heterogeneity of seismic activity,improvements were made to both the Circular Spatial Smoothing Model(CSSM)with a fixed radius and the Adaptive Spatial Smoothing Model(ASSM),with full consideration given to the spatiotemporal completeness of historical earthquake magnitudes.Regarding the CSSM,for scenarios involving small sample sizes in earthquake catalogs,the cross-validation method proposed in this study demonstrated higher robustness than the maximum likelihood method in determining the optimal correlation distance.Performance evaluation results indicate that while both models effectively characterize seismic activity,the ASSM exhibits superior overall predictive performance compared to the CSSM,owing to its ability to adaptively adjust the smoothing radius according to seismic density.Significant discrepancies were observed in the Peak Ground Acceleration(PGA)results calculated with a 10%probability of exceedance in 50 years across different combinations of seismic source models.The single spatially smoothed point-source model yielded a maximum PGA of approximately 0.52 g,with high-value areas concentrated near historical epicenters,thereby significantly underestimating the hazard associated with major fault zones.When combined with the simple fault-source model,the maximum PGA increased to 0.8 g,with high-value zones exhibiting a zonal distribution along faults;however,the risk remained underestimated for faults with low slip rates that are nevertheless approaching their recurrence cycles.Following the introduction of the time-dependent characteristic fault-source model,local PGA values for faults in the middle-to-late stages of their recurrence cycles increased by a factor of 2 to 7 compared to the single model.These results demonstrate that the characteristic fault-source model reasonably delineates the time-dependence of large earthquake recurrence,thereby providing a more accurate assessment of imminent seismic risks.By comprehensively applying the improved spatially smoothed pointsource model,the simple fault-source model,and the characteristic fault-source model,the following faults within the region were identified as having high seismic hazard:the Huangxianggou,Zhangxian,and Tianshui segments of the Xiqinling northern edge fault;the Maqin-Maqu segment of the Dongkunlun fault;the Longriqu fault;the Maoergai fault;the Elashan fault;the Riyueshan fault;the eastern segment of the Lenglongling fault;the Maxianshan segment of the Maxianshan northern Margin fault;and the Maomaoshan-Jinqianghe segment of the Laohushan-Maomaoshan fault.As these faults are located within seismic gaps or are approaching the recurrence periods of large earthquakes,they should be prioritized for current and future seismic monitoring as well as disaster prevention and mitigation efforts.展开更多
The unprecedented developments in generalist segmentation foundation models have become a dominant focus in the field of computer vision,introducing a multitude of previously unexplored capabilities in a wide range of...The unprecedented developments in generalist segmentation foundation models have become a dominant focus in the field of computer vision,introducing a multitude of previously unexplored capabilities in a wide range of natural image and video analysis tasks.From the pioneering segment anything model(SAM)that revolutionized prompt-driven image segmentation to the recent SAM2 which enables streaming video with robust spatiotemporal consistency,these models have demonstrated effective adaptability in natural scenarios and show strong potential for biomedical applications.In this paper,we present a comprehensive and in-depth review of the development,adaptation,and application of generalist segmentation foundation models in biomedical domains.We first contextualize the evolution of key models and their core mechanisms,highlighting their potential for bridging the gap between general vision and specialized biomedical tasks.We then systematically examine the challenges in applying these models to biomedical data,including domain shift,ambiguous boundaries,and dimensional gaps for 3D medical images.Finally,we articulate our perspectives on the future research directions.This review aims to provide a roadmap for researchers,facilitating the translation of generalist segmentation capabilities into effective biomedical solutions.展开更多
Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity ...Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity and high memory consumption in self-attention,Mamba-UNet achieves efficient global modeling through linear-complexity state space modeling.This paper proposes TriLVM-UNet,a lightweight three-stage architecture that integrates parameter-efficient VMamba blocks and enhances cross-stage feature interaction via an improved skip-attention bridge(SAB)module inspired by UltraLight VM-UNet.The model incorporates a Lightweight Vision Mamba(LVM)layer for high-resolution feature extraction,alongside multi-scale dilated convolution(MSDC)and convolutional block attention module(CBAM)for enhanced feature fusion.Evaluated on the 3D ACDC dataset against six baseline models,TriLVM-UNet achieves 98.57%accuracy.The GitHub repository is available at:http://gffzz188fe103f8f1460asxo6ox0uwbkoo6uxk.ffgz.tsg.suse.edu.cn/730432ch/TriLVM-UNet.展开更多
Colorectal cancer(CRC)is a prevalent disease,with polyps serving as its precursors.Accurate polyp segmentation is crucial for early CRC prevention.However,due to different sizes of the polyps,the boundaries are not cl...Colorectal cancer(CRC)is a prevalent disease,with polyps serving as its precursors.Accurate polyp segmentation is crucial for early CRC prevention.However,due to different sizes of the polyps,the boundaries are not clear.Therefore,accurate segmentation of polyps is a challenging task.This paper proposes vision Mamba attention feature fusion UNet(VMA-UNet),a U-shaped asymmetric codec structure model grounded in the state space model(SSM).The VMA-UNet incorporates attention feature fusion(AFF)in order to enhance the feature representation of small polyps.A new IUD loss function,namely combining intersection over union(IoU)loss function and Dice loss function,is proposed to address both large polyps and small polyps,and to mitigate the issue of data imbalance.When applied to multiple datasets,VMA-UNet demonstrates robust performance,particularly in small polyp segmentation,showcasing its practical value.The network proposed in this paper overcomes the inherent shortcomings of convolutional neural network(CNN)and transformers,not only performing well in remote interaction modeling,but also maintaining linear computational complexity.Our study introduces a new method for polyp segmentation based on SSM and advances the field.展开更多
Accurate segmentation of breast cancer in mammogram images plays a critical role in early diagnosis and treatment planning.As research in this domain continues to expand,various segmentation techniques have been propo...Accurate segmentation of breast cancer in mammogram images plays a critical role in early diagnosis and treatment planning.As research in this domain continues to expand,various segmentation techniques have been proposed across classical image processing,machine learning(ML),deep learning(DL),and hybrid/ensemble models.This study conducts a systematic literature review using the PRISMA methodology,analyzing 57 selected articles to explore how these methods have evolved and been applied.The review highlights the strengths and limitations of each approach,identifies commonly used public datasets,and observes emerging trends in model integration and clinical relevance.By synthesizing current findings,this work provides a structured overview of segmentation strategies and outlines key considerations for developing more adaptable and explainable tools for breast cancer detection.Overall,our synthesis suggests that classical and ML methods are suitable for limited labels and computing resources,while DL models are preferable when pixel-level annotations and resources are available,and hybrid pipelines are most appropriate when fine-grained clinical precision is required.展开更多
In this editorial we comment on the article by Chan et al.The study presents the most comprehensive comparative evaluation to date of deep learning models for multi-class segmentation of upper gastrointestinal disease...In this editorial we comment on the article by Chan et al.The study presents the most comprehensive comparative evaluation to date of deep learning models for multi-class segmentation of upper gastrointestinal diseases,leveraging a novel 3313-image,nine-class clinical dataset alongside the public EDD2020 benchmark.Their results demonstrate that hierarchical,pre-trained encoders(notably Swin-UMamba-D)deliver the highest segmentation accuracy,while SegFormer balances accuracy with computational efficiency,an important consideration for clinical deployment.Beyond raw performance metrics,the work confronts core translational barriers:Limited and biased datasets,lighting and imaging variability,boundary ambiguity,and multi-label complexity.This Editorial argues that the manuscript marks a pivotal shift from isolated technical advances toward clinically-minded validation of segmentation systems,and proposes a concrete agenda for the field to accelerate safe,generalizable,and ethically responsible adoption of automated endoscopic assistance.展开更多
The quantitative analysis of dispersed phases(bubbles,droplets,and particles)in multiphase flow systems represents a persistent technological challenge in petroleum engineering applications,including CO2-enhanced o...The quantitative analysis of dispersed phases(bubbles,droplets,and particles)in multiphase flow systems represents a persistent technological challenge in petroleum engineering applications,including CO2-enhanced oil recovery,foam flooding,and unconventional reservoir development.Current characterization methods remain constrained by labor-intensive manual workflows and limited dynamic analysis capabilities,particularly for processing large-scale microscopy data and video sequences that capture critical transient behavior like gas cluster migration and droplet coalescence.These limitations hinder the establishment of robust correlations between pore-scale flow patterns and reservoir-scale production performance.This study introduces a novel computer vision framework that integrates foundation models with lightweight neural networks to address these industry challenges.Leveraging the segment anything model's zero-shot learning capability,we developed an automated workflow that achieves an efficiency improvement of approximately 29 times in bubble labeling compared to manual methods while maintaining less than 2%deviation from expert annotations.Engineering-oriented optimization ensures lightweight deployment with 94%segmentation accuracy,while the integrated quantification system precisely resolves gas saturation,shape factors,and interfacial dynamics,parameters critical for optimizing gas injection strategies and predicting phase redistribution patterns.Validated through microfluidic gas-liquid displacement experiments for discontinuous phase segmentation accuracy,this methodology enables precise bubble morphology quantification with broad application potential in multiphase systems,including emulsion droplet dynamics characterization and particle transport behavior analysis.This work bridges the critical gap between pore-scale dynamics characterization and reservoir-scale simulation requirements,providing a foundational framework for intelligent flow diagnostics and predictive modeling in next-generation digital oilfield systems.展开更多
Despite its remarkable performance on natural images,the segment anything model(SAM)lacks domain-specific information in medical imaging.and faces the challenge of losing local multi-scale information in the encoding ...Despite its remarkable performance on natural images,the segment anything model(SAM)lacks domain-specific information in medical imaging.and faces the challenge of losing local multi-scale information in the encoding phase.This paper presents a medical image segmentation model based on SAM with a local multi-scale feature encoder(LMSFE-SAM)to address the issues above.Firstly,based on the SAM,a local multi-scale feature encoder is introduced to improve the representation of features within local receptive field,thereby supplying the Vision Transformer(ViT)branch in SAM with enriched local multi-scale contextual information.At the same time,a multiaxial Hadamard product module(MHPM)is incorporated into the local multi-scale feature encoder in a lightweight manner to reduce the quadratic complexity and noise interference.Subsequently,a cross-branch balancing adapter is designed to balance the local and global information between the local multi-scale feature encoder and the ViT encoder in SAM.Finally,to obtain smaller input image size and to mitigate overlapping in patch embeddings,the size of the input image is reduced from 1024×1024 pixels to 256×256 pixels,and a multidimensional information adaptation component is developed,which includes feature adapters,position adapters,and channel-spatial adapters.This component effectively integrates the information from small-sized medical images into SAM,enhancing its suitability for clinical deployment.The proposed model demonstrates an average enhancement ranging from 0.0387 to 0.3191 across six objective evaluation metrics on BUSI,DDTI,and TN3K datasets compared to eight other representative image segmentation models.This significantly enhances the performance of the SAM on medical images,providing clinicians with a powerful tool in clinical diagnosis.展开更多
Efficient segmentation of oiled pixels in optical remotely sensed images is the precondition of optical identification and classification of different spilled oils,which remains one of the keys to optical remote sensi...Efficient segmentation of oiled pixels in optical remotely sensed images is the precondition of optical identification and classification of different spilled oils,which remains one of the keys to optical remote sensing of oil spills.Optical remotely sensed images of oil spills are inherently multidimensional and embedded with a complex knowledge framework.This complexity often hinders the effectiveness of mechanistic algorithms across varied scenarios.Although optical remote-sensing theory for oil spills has advanced,the scarcity of curated datasets and the difficulty of collecting them limit their usefulness for training deep learning models.This study introduces a data expansion strategy that utilizes the Segment Anything Model(SAM),effectively bridging the gap between traditional mechanism algorithms and emergent self-adaptive deep learning models.Optical dimension reduction is achieved through standardized preprocessing processes that address the decipherable properties of the input image.After preprocessing,SAM can swiftly and accurately segment spilled oil in images.The unified AI-based workflow significantly accelerates labeled-dataset creation and has proven effective for both rapid emergency intelligence during spill incidents and the rapid mapping and classification of oil footprints across China’s coastal waters.Our results show that coupling a remote sensing mechanism with a foundation model enables near-real-time,large-scale monitoring of complex surface slicks and offers guidance for the next generation of detection and quantification algorithms.展开更多
Segmenting skin lesions is critical for early skin cancer detection.Existing CNN and Transformer-based methods face challenges such as high computational complexity and limited adaptability to variations in lesion siz...Segmenting skin lesions is critical for early skin cancer detection.Existing CNN and Transformer-based methods face challenges such as high computational complexity and limited adaptability to variations in lesion sizes.To overcome these limitations,we introduce MSAMamba-UNet,a lightweight model that integrates two novel architectures:Multi-Scale Mamba(MSMamba)and Adaptive Dynamic Gating Block(ADGB).MSMamba utilizes multi-scale decomposition and a parallel hierarchical structure to enhance the delineation of irregular lesion boundaries and sensitivity to small targets.ADGB dynamically selects convolutional kernels with varying receptive fields based on input features,improving the model’s capacity to accommodate diverse lesion textures and scales.Additionally,we introduce a Mix Attention Fusion Block(MAF)to enhance shallow feature representation by integrating parallel channel and pixel attention mechanisms.Extensive evaluation of MSAMamba-UNet on the ISIC 2016,ISIC 2017,and ISIC 2018 datasets demonstrates competitive segmentation accuracy with only 0.056 M parameters and 0.069 GFLOPs.Our experiments revealed that MSAMamba-UNet achieved IoU scores of 85.53%,85.47%,and 82.22%,as well as DSC scores of 92.20%,92.17%,and 90.24%,respectively.These results underscore the lightweight design and effectiveness of MSAMamba-UNet.展开更多
Objective This study aimed to explore a novel method that integrates the segmentation guidance classification and the dif-fusion model augmentation to realize the automatic classification for tibial plateau fractures(...Objective This study aimed to explore a novel method that integrates the segmentation guidance classification and the dif-fusion model augmentation to realize the automatic classification for tibial plateau fractures(TPFs).Methods YOLOv8n-cls was used to construct a baseline model on the data of 3781 patients from the Orthopedic Trauma Center of Wuhan Union Hospital.Additionally,a segmentation-guided classification approach was proposed.To enhance the dataset,a diffusion model was further demonstrated for data augmentation.Results The novel method that integrated the segmentation-guided classification and diffusion model augmentation sig-nificantly improved the accuracy and robustness of fracture classification.The average accuracy of classification for TPFs rose from 0.844 to 0.896.The comprehensive performance of the dual-stream model was also significantly enhanced after many rounds of training,with both the macro-area under the curve(AUC)and the micro-AUC increasing from 0.94 to 0.97.By utilizing diffusion model augmentation and segmentation map integration,the model demonstrated superior efficacy in identifying SchatzkerⅠ,achieving an accuracy of 0.880.It yielded an accuracy of 0.898 for SchatzkerⅡandⅢand 0.913 for SchatzkerⅣ;for SchatzkerⅤandⅥ,the accuracy was 0.887;and for intercondylar ridge fracture,the accuracy was 0.923.Conclusion The dual-stream attention-based classification network,which has been verified by many experiments,exhibited great potential in predicting the classification of TPFs.This method facilitates automatic TPF assessment and may assist surgeons in the rapid formulation of surgical plans.展开更多
This systematic review aims to comprehensively examine and compare deep learning methods for brain tumor segmentation and classification using MRI and other imaging modalities,focusing on recent trends from 2022 to 20...This systematic review aims to comprehensively examine and compare deep learning methods for brain tumor segmentation and classification using MRI and other imaging modalities,focusing on recent trends from 2022 to 2025.The primary objective is to evaluate methodological advancements,model performance,dataset usage,and existing challenges in developing clinically robust AI systems.We included peer-reviewed journal articles and highimpact conference papers published between 2022 and 2025,written in English,that proposed or evaluated deep learning methods for brain tumor segmentation and/or classification.Excluded were non-open-access publications,books,and non-English articles.A structured search was conducted across Scopus,Google Scholar,Wiley,and Taylor&Francis,with the last search performed in August 2025.Risk of bias was not formally quantified but considered during full-text screening based on dataset diversity,validation methods,and availability of performance metrics.We used narrative synthesis and tabular benchmarking to compare performance metrics(e.g.,accuracy,Dice score)across model types(CNN,Transformer,Hybrid),imaging modalities,and datasets.A total of 49 studies were included(43 journal articles and 6 conference papers).These studies spanned over 9 public datasets(e.g.,BraTS,Figshare,REMBRANDT,MOLAB)and utilized a range of imaging modalities,predominantly MRI.Hybrid models,especially ResViT and UNetFormer,consistently achieved high performance,with classification accuracy exceeding 98%and segmentation Dice scores above 0.90 across multiple studies.Transformers and hybrid architectures showed increasing adoption post2023.Many studies lacked external validation and were evaluated only on a few benchmark datasets,raising concerns about generalizability and dataset bias.Few studies addressed clinical interpretability or uncertainty quantification.Despite promising results,particularly for hybrid deep learning models,widespread clinical adoption remains limited due to lack of validation,interpretability concerns,and real-world deployment barriers.展开更多
AIM:To construct an intelligent segmentation scheme for precise localization of central serous chorioretinopathy(CSC)leakage points,thereby enabling ophthalmologists to deliver accurate laser treatment without navigat...AIM:To construct an intelligent segmentation scheme for precise localization of central serous chorioretinopathy(CSC)leakage points,thereby enabling ophthalmologists to deliver accurate laser treatment without navigational laser equipment.METHODS:A dataset with dual labels(point-level and pixel-level)was first established based on fundus fluorescein angiography(FFA)images of CSC and subsequently divided into training(102 images),validation(40 images),and test(40 images)datasets.An intelligent segmentation method was then developed,based on the You Only Look Once version 8 Pose Estimation(YOLOv8-Pose)model and segment anything model(SAM),to segment CSC leakage points.Next,the YOLOv8-Pose model was trained for 200 epochs,and the best-performing model was selected to form the optimal combination with SAM.Additionally,the classic five types of U-Net series models[i.e.,U-Net,recurrent residual U-Net(R2U-Net),attention U-Net(AttU-Net),recurrent residual attention U-Net(R2AttUNet),and nested U-Net(UNet++)]were initialized with three random seeds and trained for 200 epochs,resulting in a total of 15 baseline models for comparison.Finally,based on the metrics including Dice similarity coefficient(DICE),intersection over union(IoU),precision,recall,precisionrecall(PR)curve,and receiver operating characteristic(ROC)curve,the proposed method was compared with baseline models through quantitative and qualitative experiments for leakage point segmentation,thereby demonstrating its effectiveness.RESULTS:With the increase of training epochs,the mAP50-95,Recall,and precision of the YOLOv8-Pose model showed a significant increase and tended to stabilize,and it achieved a preliminary localization success rate of 90%(i.e.,36 images)for CSC leakage points in 40 test images.Using manually expert-annotated pixel-level labels as the ground truth,the proposed method achieved outcomes with a DICE of 57.13%,an IoU of 45.31%,a precision of 45.91%,a recall of 93.57%,an area under the PR curve(AUC-PR)of 0.78 and an area under the ROC curve(AUC-ROC)of 0.97,which enables more accurate segmentation of CSC leakage points.CONCLUSION:By combining the precise localization capability of the YOLOv8-Pose model with the robust and flexible segmentation ability of SAM,the proposed method not only demonstrates the effectiveness of the YOLOv8-Pose model in detecting keypoint coordinates of CSC leakage points from the perspective of application innovation but also establishes a novel approach for accurate segmentation of CSC leakage points through the“detect-then-segment”strategy,thereby providing a potential auxiliary means for the automatic and precise realtime localization of leakage points during traditional laser photocoagulation for CSC.展开更多
Real-time identification of rock chip size and shape distributions from muck images plays a critical role in intelligently optimizing cutterhead thrust and torque parameters for tunnel boring machines(TBM).However,com...Real-time identification of rock chip size and shape distributions from muck images plays a critical role in intelligently optimizing cutterhead thrust and torque parameters for tunnel boring machines(TBM).However,complex light environments in field images are difficult to recognize via traditional methods.This paper proposes a U-Net-SAM framework integrating semantic segmentation and the vision foundation model—Segment Anything Model(SAM),combined with dropout-based uncertainty analysis,achieving efficient rock chip segmentation and parameter quantification.First,a U-Net is trained to identify the rock mass centroid as an automatic SAM prompt.Next,an overlap region optimization strategy based on Intersection over Union(IoU)and a noise filtering method is employed to tackle boundary blurring and particle adhesion.Finally,a Dropout layer is added to implement the committee-based uncertainty analysis model and quantify predictive uncertainty.Results show that:(1)U-Net-SAM improves mean F1-score and PA by 9.1%and 7.8%over U-Net;(2)A strong correlation between prediction standard deviation(SD)and error rate validates the proposed uncertainty quantification strategy.This framework provides reliable rock chip perception for intelligent TBM tunneling,with potential applications in other engineering scenarios.展开更多
基金National Natural Science Foundation of China(42325403)National Science and Technology Major Project of China(2024ZD1004201)support from China Scholarship Council(CSC202306450071).
摘要Accurate segmentation of digital rock images is essential for characterizing pore-matrix systems and predicting petrophysical properties.However,the diversity of rock textures across different lithologies poses a significant challenge for conventional segmentation networks,especially under limited training data.To address this,we introduce DRI-SAM(Digital Rock Image-Segment Anything Model),a hybrid segmentation framework that leverages the powerful visual prior of the Segment Anything Model(SAM)and adapts it to the digital rock domain.Specifically,we apply LoRA-based fine-tuning to SAM’s image encoder to better capture rock-specific microstructures,while U-Net is employed to generate prompt points,guiding SAM toward accurate pore-matrix delineation.This approach retains the encoder’s representational power while allowing domain-specific adaptation via LoRA,enabling effective cross-domain generalization under limited training data.The model is trained exclusively on 200 annotated images of Bentheimer sandstone,covering two distinct voxel resolutions,and is evaluated on digital rock images of varying lithologies,resolutions and imaging modalities.The results confirm that DRI-SAM achieves accurate segmentation on both sandstone and more challenging carbonate samples,including synthetic and SEM images,without additional retraining or parameter adjustments.Compared to DeepLabV3+and the only LoRA-tuned SAM,DRI-SAM demonstrates superior performance under limited supervision,highlighting its strong generalization and practical value in digital rock image analysis.Moreover,the findings suggest that foundation models like SAM,when properly adapted,also hold great promise for broader geoscientific imaging tasks.
基金financially supported by the China Scholarship Council(No.202208320010).
摘要Rock fragment size distribution(FSD)plays an important role in various engineering applications,such as mining,tunnelling,and other underground construction scenarios.While vision-based deep learning approaches have been increasingly applied to FSD analysis,they are often case-specific,showing limited cross-site generalization despite their accuracy.To address these challenges,FragSAM,an end-to-end,fully automated framework is proposed for near real-time rock fragment segmentation and FSD analysis across diverse engineering environments.FragSAM integrates the generalization power of Segment Anything Model(SAM)with a context-aware prompting mechanism and lightweight architecture for efficient dense fragment segmentation.In Stage 1,an enhanced SAM automatically generates high-quality annotations,which are used to train a modified CenterNet for precise centroid prediction.In Stage 2,these centroids serve as prompts for EdgeSAM,a lightweight SAM variant optimized for real-time inference.This two-stage design eliminates dense grid prompting and reduces reliance on heavy postprocessing,enabling efficient and scalable segmentation.Experimental results show that FragSAM achieves competitive segmentation performance with significantly lower latency and model complexity compared to existing SAM-based methods.In comparison with supervised learning approaches,it also demonstrates superior generalization and performs better in low-quality or unseen scenarios.Furthermore,case studies on blasting fragmentation,TBM muck,and coastal rock surfaces confirm its robustness and seamless cross-site adaptability,requiring no tuning or retraining,making it highly practical for on-site applications.
基金supported by the National Natural Science Foundation of China(Grant No.42372322)Science and Technology Innovation Program for Postgraduate students in IDP subsidized by Fundamental Research Funds for the Central Universities(Grant No.ZY20240314)the research project of Spark Plan for Earthquake Science and Technology(Grant No.XH24060A).
摘要Locked segments are high-strength structural elements in fault zones that release significantseismic energy during earthquakes.In fracture mechanics,they act as high-stress concentration patches(asperities)where rupture initiates.The progressive failure of locked segments along faults plays a crucial role in the energy partition of earthquakes.The impact of locked segments on the near-fielddeformation and nucleation of faults,however,remains poorly understood.In this study,rock-like materials with pre-manufactured strike-slip faults containing various locked segments lengths under uniaxial stress.The mechanical properties,local deformation fields,and slip displacement rates during the uniaxial loading of the models were quantified.Results indicate that the uniaxial compressive strength and elastic modulus of the system peak once the ratio of locked segment to fault length is approximately 0.6.Meanwhile,the resistance of the models to deformation increased,and the failure mode transformed from shear failure to tensile failure.Under loading,compression and dilatation quadrants were formed on both sides of the fault.Large-scale fractures dominate the dilatation quadrants,and the degree of deformation disturbance in this region was significantly higher than that in the compression quadrants.With increasing locked segment length,the amplitude of deformation perturbations decreased after the peak strength.Shorter locked segments were more susceptible to deformation and failure.In the fracture evolution process,a relationship between the stress deflectionangle and the displacement rate was found,which is empirically described by an exponential function.These findings clarify geological structures failure mechanisms and support seismic hazard assessment for strike-slip earthquake regions.
基金supported by the National Key R&D Program of China(No.2022YFC3003502).
摘要The northern segment of the North-South Seismic Belt is characterized by intense crustal deformation,well-developed active tectonics,and frequent occurrences of strong earthquakes.Therefore,conducting a Probabilistic Seismic Hazard Analysis(PSHA)for this region is of significant importance for supporting seismic fortification in major engineering projects and formulating disaster prevention and mitigation policies.In this study,a composite seismic source model was constructed by integrating data on historical earthquakes,active faults,and paleoseismicity.Furthermore,a logic tree framework was employed to quantify epistemic uncertainties,enabling a systematic seismic hazard assessment of the region.To more accurately characterize the spatial heterogeneity of seismic activity,improvements were made to both the Circular Spatial Smoothing Model(CSSM)with a fixed radius and the Adaptive Spatial Smoothing Model(ASSM),with full consideration given to the spatiotemporal completeness of historical earthquake magnitudes.Regarding the CSSM,for scenarios involving small sample sizes in earthquake catalogs,the cross-validation method proposed in this study demonstrated higher robustness than the maximum likelihood method in determining the optimal correlation distance.Performance evaluation results indicate that while both models effectively characterize seismic activity,the ASSM exhibits superior overall predictive performance compared to the CSSM,owing to its ability to adaptively adjust the smoothing radius according to seismic density.Significant discrepancies were observed in the Peak Ground Acceleration(PGA)results calculated with a 10%probability of exceedance in 50 years across different combinations of seismic source models.The single spatially smoothed point-source model yielded a maximum PGA of approximately 0.52 g,with high-value areas concentrated near historical epicenters,thereby significantly underestimating the hazard associated with major fault zones.When combined with the simple fault-source model,the maximum PGA increased to 0.8 g,with high-value zones exhibiting a zonal distribution along faults;however,the risk remained underestimated for faults with low slip rates that are nevertheless approaching their recurrence cycles.Following the introduction of the time-dependent characteristic fault-source model,local PGA values for faults in the middle-to-late stages of their recurrence cycles increased by a factor of 2 to 7 compared to the single model.These results demonstrate that the characteristic fault-source model reasonably delineates the time-dependence of large earthquake recurrence,thereby providing a more accurate assessment of imminent seismic risks.By comprehensively applying the improved spatially smoothed pointsource model,the simple fault-source model,and the characteristic fault-source model,the following faults within the region were identified as having high seismic hazard:the Huangxianggou,Zhangxian,and Tianshui segments of the Xiqinling northern edge fault;the Maqin-Maqu segment of the Dongkunlun fault;the Longriqu fault;the Maoergai fault;the Elashan fault;the Riyueshan fault;the eastern segment of the Lenglongling fault;the Maxianshan segment of the Maxianshan northern Margin fault;and the Maomaoshan-Jinqianghe segment of the Laohushan-Maomaoshan fault.As these faults are located within seismic gaps or are approaching the recurrence periods of large earthquakes,they should be prioritized for current and future seismic monitoring as well as disaster prevention and mitigation efforts.
基金supported by the National Natural Science Foundation of China(Grants 82394432 and 92249302)Shanghai Municipal Science and Technology Major Project(Grant 2023SHZDZX02).
摘要The unprecedented developments in generalist segmentation foundation models have become a dominant focus in the field of computer vision,introducing a multitude of previously unexplored capabilities in a wide range of natural image and video analysis tasks.From the pioneering segment anything model(SAM)that revolutionized prompt-driven image segmentation to the recent SAM2 which enables streaming video with robust spatiotemporal consistency,these models have demonstrated effective adaptability in natural scenarios and show strong potential for biomedical applications.In this paper,we present a comprehensive and in-depth review of the development,adaptation,and application of generalist segmentation foundation models in biomedical domains.We first contextualize the evolution of key models and their core mechanisms,highlighting their potential for bridging the gap between general vision and specialized biomedical tasks.We then systematically examine the challenges in applying these models to biomedical data,including domain shift,ambiguous boundaries,and dimensional gaps for 3D medical images.Finally,we articulate our perspectives on the future research directions.This review aims to provide a roadmap for researchers,facilitating the translation of generalist segmentation capabilities into effective biomedical solutions.
基金The Key Project of Ningxia Natural Science Foundation under grant 2025AAC020006the Regional Program of National Natural Science Foundation of China under grant 62561002The Shaanxi University of Technology Foundation Project under grant SLGRCQD2137.
摘要Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity and high memory consumption in self-attention,Mamba-UNet achieves efficient global modeling through linear-complexity state space modeling.This paper proposes TriLVM-UNet,a lightweight three-stage architecture that integrates parameter-efficient VMamba blocks and enhances cross-stage feature interaction via an improved skip-attention bridge(SAB)module inspired by UltraLight VM-UNet.The model incorporates a Lightweight Vision Mamba(LVM)layer for high-resolution feature extraction,alongside multi-scale dilated convolution(MSDC)and convolutional block attention module(CBAM)for enhanced feature fusion.Evaluated on the 3D ACDC dataset against six baseline models,TriLVM-UNet achieves 98.57%accuracy.The GitHub repository is available at:http://gffzz188fe103f8f1460asxo6ox0uwbkoo6uxk.ffgz.tsg.suse.edu.cn/730432ch/TriLVM-UNet.
基金supported by the Natural Science Research Project of Tianjin Education Commission(No.2020KJ124)the National Natural Science Foundation of China(No.11601372)the National Key Research and Development Program of China(No.2022YFF0706003)。
摘要Colorectal cancer(CRC)is a prevalent disease,with polyps serving as its precursors.Accurate polyp segmentation is crucial for early CRC prevention.However,due to different sizes of the polyps,the boundaries are not clear.Therefore,accurate segmentation of polyps is a challenging task.This paper proposes vision Mamba attention feature fusion UNet(VMA-UNet),a U-shaped asymmetric codec structure model grounded in the state space model(SSM).The VMA-UNet incorporates attention feature fusion(AFF)in order to enhance the feature representation of small polyps.A new IUD loss function,namely combining intersection over union(IoU)loss function and Dice loss function,is proposed to address both large polyps and small polyps,and to mitigate the issue of data imbalance.When applied to multiple datasets,VMA-UNet demonstrates robust performance,particularly in small polyp segmentation,showcasing its practical value.The network proposed in this paper overcomes the inherent shortcomings of convolutional neural network(CNN)and transformers,not only performing well in remote interaction modeling,but also maintaining linear computational complexity.Our study introduces a new method for polyp segmentation based on SSM and advances the field.
基金funded by BK21 FOUR(Fostering Outstanding Universities for Research)(No.:5199990914048).
摘要Accurate segmentation of breast cancer in mammogram images plays a critical role in early diagnosis and treatment planning.As research in this domain continues to expand,various segmentation techniques have been proposed across classical image processing,machine learning(ML),deep learning(DL),and hybrid/ensemble models.This study conducts a systematic literature review using the PRISMA methodology,analyzing 57 selected articles to explore how these methods have evolved and been applied.The review highlights the strengths and limitations of each approach,identifies commonly used public datasets,and observes emerging trends in model integration and clinical relevance.By synthesizing current findings,this work provides a structured overview of segmentation strategies and outlines key considerations for developing more adaptable and explainable tools for breast cancer detection.Overall,our synthesis suggests that classical and ML methods are suitable for limited labels and computing resources,while DL models are preferable when pixel-level annotations and resources are available,and hybrid pipelines are most appropriate when fine-grained clinical precision is required.
摘要In this editorial we comment on the article by Chan et al.The study presents the most comprehensive comparative evaluation to date of deep learning models for multi-class segmentation of upper gastrointestinal diseases,leveraging a novel 3313-image,nine-class clinical dataset alongside the public EDD2020 benchmark.Their results demonstrate that hierarchical,pre-trained encoders(notably Swin-UMamba-D)deliver the highest segmentation accuracy,while SegFormer balances accuracy with computational efficiency,an important consideration for clinical deployment.Beyond raw performance metrics,the work confronts core translational barriers:Limited and biased datasets,lighting and imaging variability,boundary ambiguity,and multi-label complexity.This Editorial argues that the manuscript marks a pivotal shift from isolated technical advances toward clinically-minded validation of segmentation systems,and proposes a concrete agenda for the field to accelerate safe,generalizable,and ethically responsible adoption of automated endoscopic assistance.
基金supported by Sichuan Province Outstanding Young Scientist Fund(Grant No.2025NSFJQ0009)Sichuan Regional Innovation Cooperation Fund(Grant No.2025YFHZ0270)。
摘要The quantitative analysis of dispersed phases(bubbles,droplets,and particles)in multiphase flow systems represents a persistent technological challenge in petroleum engineering applications,including CO2-enhanced oil recovery,foam flooding,and unconventional reservoir development.Current characterization methods remain constrained by labor-intensive manual workflows and limited dynamic analysis capabilities,particularly for processing large-scale microscopy data and video sequences that capture critical transient behavior like gas cluster migration and droplet coalescence.These limitations hinder the establishment of robust correlations between pore-scale flow patterns and reservoir-scale production performance.This study introduces a novel computer vision framework that integrates foundation models with lightweight neural networks to address these industry challenges.Leveraging the segment anything model's zero-shot learning capability,we developed an automated workflow that achieves an efficiency improvement of approximately 29 times in bubble labeling compared to manual methods while maintaining less than 2%deviation from expert annotations.Engineering-oriented optimization ensures lightweight deployment with 94%segmentation accuracy,while the integrated quantification system precisely resolves gas saturation,shape factors,and interfacial dynamics,parameters critical for optimizing gas injection strategies and predicting phase redistribution patterns.Validated through microfluidic gas-liquid displacement experiments for discontinuous phase segmentation accuracy,this methodology enables precise bubble morphology quantification with broad application potential in multiphase systems,including emulsion droplet dynamics characterization and particle transport behavior analysis.This work bridges the critical gap between pore-scale dynamics characterization and reservoir-scale simulation requirements,providing a foundational framework for intelligent flow diagnostics and predictive modeling in next-generation digital oilfield systems.
基金supported by Natural Science Foundation Programme of Gansu Province(No.24JRRA231)National Natural Science Foundation of China(No.62061023)Gansu Provincial Science and Technology Plan Key Research and Development Program Project(No.24YFFA024).
摘要Despite its remarkable performance on natural images,the segment anything model(SAM)lacks domain-specific information in medical imaging.and faces the challenge of losing local multi-scale information in the encoding phase.This paper presents a medical image segmentation model based on SAM with a local multi-scale feature encoder(LMSFE-SAM)to address the issues above.Firstly,based on the SAM,a local multi-scale feature encoder is introduced to improve the representation of features within local receptive field,thereby supplying the Vision Transformer(ViT)branch in SAM with enriched local multi-scale contextual information.At the same time,a multiaxial Hadamard product module(MHPM)is incorporated into the local multi-scale feature encoder in a lightweight manner to reduce the quadratic complexity and noise interference.Subsequently,a cross-branch balancing adapter is designed to balance the local and global information between the local multi-scale feature encoder and the ViT encoder in SAM.Finally,to obtain smaller input image size and to mitigate overlapping in patch embeddings,the size of the input image is reduced from 1024×1024 pixels to 256×256 pixels,and a multidimensional information adaptation component is developed,which includes feature adapters,position adapters,and channel-spatial adapters.This component effectively integrates the information from small-sized medical images into SAM,enhancing its suitability for clinical deployment.The proposed model demonstrates an average enhancement ranging from 0.0387 to 0.3191 across six objective evaluation metrics on BUSI,DDTI,and TN3K datasets compared to eight other representative image segmentation models.This significantly enhances the performance of the SAM on medical images,providing clinicians with a powerful tool in clinical diagnosis.
基金The National Natural Science Foundation of China under contract No.42371380the National Key Research and Development Program of China under contract No.2023YFC2811800the Fundamental Research Funds for the Central Universities under contract No.0904-14380035.
摘要Efficient segmentation of oiled pixels in optical remotely sensed images is the precondition of optical identification and classification of different spilled oils,which remains one of the keys to optical remote sensing of oil spills.Optical remotely sensed images of oil spills are inherently multidimensional and embedded with a complex knowledge framework.This complexity often hinders the effectiveness of mechanistic algorithms across varied scenarios.Although optical remote-sensing theory for oil spills has advanced,the scarcity of curated datasets and the difficulty of collecting them limit their usefulness for training deep learning models.This study introduces a data expansion strategy that utilizes the Segment Anything Model(SAM),effectively bridging the gap between traditional mechanism algorithms and emergent self-adaptive deep learning models.Optical dimension reduction is achieved through standardized preprocessing processes that address the decipherable properties of the input image.After preprocessing,SAM can swiftly and accurately segment spilled oil in images.The unified AI-based workflow significantly accelerates labeled-dataset creation and has proven effective for both rapid emergency intelligence during spill incidents and the rapid mapping and classification of oil footprints across China’s coastal waters.Our results show that coupling a remote sensing mechanism with a foundation model enables near-real-time,large-scale monitoring of complex surface slicks and offers guidance for the next generation of detection and quantification algorithms.
基金supported in part by the National Natural Science Foundation of China under Grant 62201201the Foundation of Henan Educational Committee under Grant 242102211042.
摘要Segmenting skin lesions is critical for early skin cancer detection.Existing CNN and Transformer-based methods face challenges such as high computational complexity and limited adaptability to variations in lesion sizes.To overcome these limitations,we introduce MSAMamba-UNet,a lightweight model that integrates two novel architectures:Multi-Scale Mamba(MSMamba)and Adaptive Dynamic Gating Block(ADGB).MSMamba utilizes multi-scale decomposition and a parallel hierarchical structure to enhance the delineation of irregular lesion boundaries and sensitivity to small targets.ADGB dynamically selects convolutional kernels with varying receptive fields based on input features,improving the model’s capacity to accommodate diverse lesion textures and scales.Additionally,we introduce a Mix Attention Fusion Block(MAF)to enhance shallow feature representation by integrating parallel channel and pixel attention mechanisms.Extensive evaluation of MSAMamba-UNet on the ISIC 2016,ISIC 2017,and ISIC 2018 datasets demonstrates competitive segmentation accuracy with only 0.056 M parameters and 0.069 GFLOPs.Our experiments revealed that MSAMamba-UNet achieved IoU scores of 85.53%,85.47%,and 82.22%,as well as DSC scores of 92.20%,92.17%,and 90.24%,respectively.These results underscore the lightweight design and effectiveness of MSAMamba-UNet.
基金supported by the National Natural Science Foundation of China(Nos.81974355 and 82172524)Key Research and Development Program of Hubei Province(No.2021BEA161)+2 种基金National Innovation Platform Development Program(No.2020021105012440)Open Project Funding of the Hubei Key Laboratory of Big Data Intelligent Analysis and Application,Hubei University(No.2024BDIAA03)Free Innovation Preliminary Research Fund of Wuhan Union Hospital(No.2024XHYN047).
摘要Objective This study aimed to explore a novel method that integrates the segmentation guidance classification and the dif-fusion model augmentation to realize the automatic classification for tibial plateau fractures(TPFs).Methods YOLOv8n-cls was used to construct a baseline model on the data of 3781 patients from the Orthopedic Trauma Center of Wuhan Union Hospital.Additionally,a segmentation-guided classification approach was proposed.To enhance the dataset,a diffusion model was further demonstrated for data augmentation.Results The novel method that integrated the segmentation-guided classification and diffusion model augmentation sig-nificantly improved the accuracy and robustness of fracture classification.The average accuracy of classification for TPFs rose from 0.844 to 0.896.The comprehensive performance of the dual-stream model was also significantly enhanced after many rounds of training,with both the macro-area under the curve(AUC)and the micro-AUC increasing from 0.94 to 0.97.By utilizing diffusion model augmentation and segmentation map integration,the model demonstrated superior efficacy in identifying SchatzkerⅠ,achieving an accuracy of 0.880.It yielded an accuracy of 0.898 for SchatzkerⅡandⅢand 0.913 for SchatzkerⅣ;for SchatzkerⅤandⅥ,the accuracy was 0.887;and for intercondylar ridge fracture,the accuracy was 0.923.Conclusion The dual-stream attention-based classification network,which has been verified by many experiments,exhibited great potential in predicting the classification of TPFs.This method facilitates automatic TPF assessment and may assist surgeons in the rapid formulation of surgical plans.
摘要This systematic review aims to comprehensively examine and compare deep learning methods for brain tumor segmentation and classification using MRI and other imaging modalities,focusing on recent trends from 2022 to 2025.The primary objective is to evaluate methodological advancements,model performance,dataset usage,and existing challenges in developing clinically robust AI systems.We included peer-reviewed journal articles and highimpact conference papers published between 2022 and 2025,written in English,that proposed or evaluated deep learning methods for brain tumor segmentation and/or classification.Excluded were non-open-access publications,books,and non-English articles.A structured search was conducted across Scopus,Google Scholar,Wiley,and Taylor&Francis,with the last search performed in August 2025.Risk of bias was not formally quantified but considered during full-text screening based on dataset diversity,validation methods,and availability of performance metrics.We used narrative synthesis and tabular benchmarking to compare performance metrics(e.g.,accuracy,Dice score)across model types(CNN,Transformer,Hybrid),imaging modalities,and datasets.A total of 49 studies were included(43 journal articles and 6 conference papers).These studies spanned over 9 public datasets(e.g.,BraTS,Figshare,REMBRANDT,MOLAB)and utilized a range of imaging modalities,predominantly MRI.Hybrid models,especially ResViT and UNetFormer,consistently achieved high performance,with classification accuracy exceeding 98%and segmentation Dice scores above 0.90 across multiple studies.Transformers and hybrid architectures showed increasing adoption post2023.Many studies lacked external validation and were evaluated only on a few benchmark datasets,raising concerns about generalizability and dataset bias.Few studies addressed clinical interpretability or uncertainty quantification.Despite promising results,particularly for hybrid deep learning models,widespread clinical adoption remains limited due to lack of validation,interpretability concerns,and real-world deployment barriers.
基金Supported by the Shenzhen Science and Technology Program(No.JCYJ20240813152704006)the National Natural Science Foundation of China(No.62401259)+2 种基金the Fundamental Research Funds for the Central Universities(No.NZ2024036)the Postdoctoral Fellowship Program of CPSF(No.GZC20242228)High Performance Computing Platform of Nanjing University of Aeronautics and Astronautics。
摘要AIM:To construct an intelligent segmentation scheme for precise localization of central serous chorioretinopathy(CSC)leakage points,thereby enabling ophthalmologists to deliver accurate laser treatment without navigational laser equipment.METHODS:A dataset with dual labels(point-level and pixel-level)was first established based on fundus fluorescein angiography(FFA)images of CSC and subsequently divided into training(102 images),validation(40 images),and test(40 images)datasets.An intelligent segmentation method was then developed,based on the You Only Look Once version 8 Pose Estimation(YOLOv8-Pose)model and segment anything model(SAM),to segment CSC leakage points.Next,the YOLOv8-Pose model was trained for 200 epochs,and the best-performing model was selected to form the optimal combination with SAM.Additionally,the classic five types of U-Net series models[i.e.,U-Net,recurrent residual U-Net(R2U-Net),attention U-Net(AttU-Net),recurrent residual attention U-Net(R2AttUNet),and nested U-Net(UNet++)]were initialized with three random seeds and trained for 200 epochs,resulting in a total of 15 baseline models for comparison.Finally,based on the metrics including Dice similarity coefficient(DICE),intersection over union(IoU),precision,recall,precisionrecall(PR)curve,and receiver operating characteristic(ROC)curve,the proposed method was compared with baseline models through quantitative and qualitative experiments for leakage point segmentation,thereby demonstrating its effectiveness.RESULTS:With the increase of training epochs,the mAP50-95,Recall,and precision of the YOLOv8-Pose model showed a significant increase and tended to stabilize,and it achieved a preliminary localization success rate of 90%(i.e.,36 images)for CSC leakage points in 40 test images.Using manually expert-annotated pixel-level labels as the ground truth,the proposed method achieved outcomes with a DICE of 57.13%,an IoU of 45.31%,a precision of 45.91%,a recall of 93.57%,an area under the PR curve(AUC-PR)of 0.78 and an area under the ROC curve(AUC-ROC)of 0.97,which enables more accurate segmentation of CSC leakage points.CONCLUSION:By combining the precise localization capability of the YOLOv8-Pose model with the robust and flexible segmentation ability of SAM,the proposed method not only demonstrates the effectiveness of the YOLOv8-Pose model in detecting keypoint coordinates of CSC leakage points from the perspective of application innovation but also establishes a novel approach for accurate segmentation of CSC leakage points through the“detect-then-segment”strategy,thereby providing a potential auxiliary means for the automatic and precise realtime localization of leakage points during traditional laser photocoagulation for CSC.
基金financial support of National Natural Science Foundation of China(Grant No.52008039)the Natural Science Foundation of Hunan Province(Grant No.2021JJ40592)support from the Research Grants Council of Hong Kong(Grant No.GRF#16208224).
摘要Real-time identification of rock chip size and shape distributions from muck images plays a critical role in intelligently optimizing cutterhead thrust and torque parameters for tunnel boring machines(TBM).However,complex light environments in field images are difficult to recognize via traditional methods.This paper proposes a U-Net-SAM framework integrating semantic segmentation and the vision foundation model—Segment Anything Model(SAM),combined with dropout-based uncertainty analysis,achieving efficient rock chip segmentation and parameter quantification.First,a U-Net is trained to identify the rock mass centroid as an automatic SAM prompt.Next,an overlap region optimization strategy based on Intersection over Union(IoU)and a noise filtering method is employed to tackle boundary blurring and particle adhesion.Finally,a Dropout layer is added to implement the committee-based uncertainty analysis model and quantify predictive uncertainty.Results show that:(1)U-Net-SAM improves mean F1-score and PA by 9.1%and 7.8%over U-Net;(2)A strong correlation between prediction standard deviation(SD)and error rate validates the proposed uncertainty quantification strategy.This framework provides reliable rock chip perception for intelligent TBM tunneling,with potential applications in other engineering scenarios.