摘要
为解决传统路面裂缝分割精度差、完整度不高的问题,提出一种基于Transformer与CNN架构的路面裂缝语义分割模型TCU-Net.该模型以Transformer架构的FastViT作为主干提取全局语义表征,引入EVC ...展开更多
为解决传统路面裂缝分割精度差、完整度不高的问题,提出一种基于Transformer与CNN架构的路面裂缝语义分割模型TCU-Net.该模型以Transformer架构的FastViT作为主干提取全局语义表征,引入EVC Block增强局部纹理与边界响应,并设计融合坐标注意力的CAFusion模块实现全局与局部特征的高效协同;训练阶段采用Focal Dice Loss作为损失函数,避免样本不均衡带来的影响.通过与UNet、SegNet、DcsNet、DTrc-Net等模型的对比实验及消融实验验证了方法的有效性.结果表明:TCU-Net在Pavementscapes和Crack500两个路面裂缝数据集上均取得了最优性能,mIoU分别达77.62%和73.28%,mPA分别达86.2%和82.73%,F1-score分别为85.70%和82.19%,显著优于现有主流语义分割模型,体现了其在裂缝检测任务中的强大表达能力与泛化性能.FastViT作为主干网络,显著提升了模型在复杂图像分割任务中的表现,各项指标均有大幅提升;EVC Block能够提高模型对局部细节的感知能力,进一步提升了模型的精度;CAFusion通过自适应地为各个通道分配权重,充分融合低维局部信息和高维全局信息,较传统融合方法显著提升了模型对边缘细节和结构连续性的表征能力.基于Transformer与CNN架构的TCU-Net能够实现完整准确的路面裂缝分割,为道路养护工作提供精准的决策依据.收起
To address the issues of low accuracy and poor integrity in traditional pavement crack segmentation,this study proposes a semantic segmentation model for pavement cracks—TCU-Net—based on a hybrid Transformer and CNN architecture.This model uses FastViT,based on the Tran...MORE
To address the issues of low accuracy and poor integrity in traditional pavement crack segmentation,this study proposes a semantic segmentation model for pavement cracks—TCU-Net—based on a hybrid Transformer and CNN architecture.This model uses FastViT,based on the Transformer architecture,as the backbone to extract global semantic representations,introduces the EVC Block to enhance local texture and boundary responses,and designs a CAFusion module with integrated coordinate attention to achieve efficient collaboration between global and local features.During the training phase,Focal Dice Loss is used as the loss function to mitigate the impact of sample imbalance.Comparative experiments with models such as UNet,SegNet,DcsNet,and DTrc-Net,as well as ablation studies,validate the effectiveness of the proposed method.Results demonstrate that TCU-Net achieves state-of-the-art performance on two pavement crack datasets:Pavementscapes and Crack500.Specifically,its mIoU reaches 77.62% and 73.28%,mPA reaches 86.2% and 82.73%,and F1-score reaches 85.70% and 82.19%,respectively—outperforming existing mainstream semantic segmentation models.This reflects its strong expressive capability and generalization performance in crack detection tasks.As the backbone network,FastViT significantly enhances the model's performance in complex image segmentation,with substantial improvements in all metrics.The EVC Block strengthens the model's perception of local details,further boosting segmentation accuracy.Compared with traditional fusion methods,CAFusion adaptively assigns weights to each channel,fully integrates low-dimensional local information and high-dimensional global information,and notably improves the model's ability to represent edge details and structural continuity.TCU-Net,based on the Transformer-CNN hybrid architecture,enables complete and accurate pavement crack segmentation,providing a precise decision-making basis for road maintenance.FEWER
作者
赵全满
马志浩
葛鲁民
刘朝晖
朱琳琳
刘继法
郭桂宏
ZHAO Quanman;MA Zhihao;GE Lumin;LIU Zhaohui;ZHU Linlin;LIU Jifa;GUO Guihong(School of Traffic Engineering,Shandong Jianzhu University,Jinan 250101,China;Shandong High-speed Engineering Inspection Co.,Ltd.,Jinan 250002,China;Shandong Taishan Traffic Planning and Design Consulting Co.,Ltd.,Tai’an 271000,Shandong,China)
出处
《昆明理工大学学报(自然科学版)》
CAS
北大核心
2026年第2期144-156,共13页
Journal of Kunming University of Science and Technology(Natural Science)
基金
山东省交通运输科技计划项目(2024B39,2023B32).