Image inpainting is a crucial research area in computer vision.Despite significant advancements with deep learning methods,challenges such as information loss and weak adaptability remain.This paper introduces a Trans...Image inpainting is a crucial research area in computer vision.Despite significant advancements with deep learning methods,challenges such as information loss and weak adaptability remain.This paper introduces a Transformer-based image inpainting method named Swin2FII,which integrates SwinV2 Transformer and fast Fourier convolution structure to address information loss and bottleneck issues,significantly enhancing inpainting accuracy and expanding its application scope.Swin2FII incorporates a super-resolution model,enhancing feature extraction and information transmission through efficient reconstruction,thereby improving detail recovery and stability.We employ the Charbonnier loss function to address gradient explosion,accurately estimating low-frequency signals and enhancing the precision of detail and texture reconstruction.Furthermore,combining mixed-precision training and data augmentation significantly boosts the model's adaptability and generalization ability.Experimental results show that our Swin2FII method outperforms the existing techniques on multiple public datasets.Notably,it exhibits excellent generalization and performance in a variety of scenarios and mask scales.In addition,Swin2FII also demonstrates strong capabilities in fluid image inpainting and mural image inpainting tasks.展开更多
基金the National Natural Science Founda-tion of China(No.62162027)the Jiangxi Provincial Graduate Innovation Project(No.YC2023-S533)the Science and Technology Project of Jiangxi Provin-cial Department of Education(No.GJJ210655)。
摘要Image inpainting is a crucial research area in computer vision.Despite significant advancements with deep learning methods,challenges such as information loss and weak adaptability remain.This paper introduces a Transformer-based image inpainting method named Swin2FII,which integrates SwinV2 Transformer and fast Fourier convolution structure to address information loss and bottleneck issues,significantly enhancing inpainting accuracy and expanding its application scope.Swin2FII incorporates a super-resolution model,enhancing feature extraction and information transmission through efficient reconstruction,thereby improving detail recovery and stability.We employ the Charbonnier loss function to address gradient explosion,accurately estimating low-frequency signals and enhancing the precision of detail and texture reconstruction.Furthermore,combining mixed-precision training and data augmentation significantly boosts the model's adaptability and generalization ability.Experimental results show that our Swin2FII method outperforms the existing techniques on multiple public datasets.Notably,it exhibits excellent generalization and performance in a variety of scenarios and mask scales.In addition,Swin2FII also demonstrates strong capabilities in fluid image inpainting and mural image inpainting tasks.