Vision Transformers(ViTs)have achieved remarkable success across various artificial intelligence-based computer vision applications.However,their demanding computational and memory requirements pose significant challe...Vision Transformers(ViTs)have achieved remarkable success across various artificial intelligence-based computer vision applications.However,their demanding computational and memory requirements pose significant challenges for de-ployment on resource-constrained edge devices.Although post-training quantization(PTQ)provides a promising solution by reducing model precision with minimal calibration data,aggressive low-bit quantization typically leads to substantial perfor-mance degradation.To address this challenge,we present the truncated uniform-log2 quantizer and progressive bit-decline reconstruction method for vision Transformer quantization(TP-ViT).It is an innovative PTQ framework specifically designed for ViTs,featuring two key technical contributions:(1)truncated uniform-log2 quantizer,a novel quantization approach which effectively handles outlier values in post-Softmax activations,significantly reducing quantization errors;(2)bit-decline optimiza-tion strategy,which employs transition weights to gradually reduce bit precision while maintaining model performance under extreme quantization conditions.Comprehensive experiments on image classification,object detection,and instance segmenta-tion tasks demonstrate TP-ViT’s superior performance compared to state-of-the-art PTQ methods,particularly in challenging 3-bit quantization scenarios.Our framework achieves a notable 6.18 percentage points improvement in top-1 accuracy for ViT-small under 3-bit quantization.These results validate TP-ViT’s robustness and general applicability,paving the way for more efficient deployment of ViT models in computer vision applications on edge hardware.展开更多
The Jiamusi Block in Northeast Asia provides a critical window for investigating the subduction of the Pacific Plate beneath the Western Pacific Orogenic Belt during the Mesozoic and Cenozoic.To constrain the structur...The Jiamusi Block in Northeast Asia provides a critical window for investigating the subduction of the Pacific Plate beneath the Western Pacific Orogenic Belt during the Mesozoic and Cenozoic.To constrain the structure and evolution of the lithosphere-asthenosphere system,this study integrates regional gravity and magnetic surveys with two NW-SE-trending magnetotelluric(MT)profiles of lengths 250 and 150 km across the block.Two-dimensional nonlinear conjugate gradient inversion derived the electrical structure to a depth of 200 km.The inversion models reveal distinct crustal and mantle low-resistivity bodies(labelled C1-C8)that spatially correlate with abrupt positive gravity and aeromagnetic anomalies.These low-resistivity features are interpreted as signatures of mantle-derived fluids and partial melts,indicative of asthenospheric upwelling and lithospheric thinning.We attribute these processes to the westward subduction of the Pacific Plate,which induced significant crust-mantle interaction and fluid metasomatism.Furthermore,a comparative analysis reveals that the lithospheric thickness of the Jiamusi Block(70-95 km)is intermediate between that of the Songnen Block(65-80 km)and the Xing'an Block(100-120 km).This variation suggests that lithospheric thickness in the region is primarily controlled by the geometry of the subducting slab's mantle wedge and the intensity of the associated fluid metasomatism.展开更多
It is well known that a block discrete cosine transform compressed image exhibits visually annoying blocking artifacts at low-bit-rate. A new post-processing deblocking algorithm in wavelet domain is proposed. The alg...It is well known that a block discrete cosine transform compressed image exhibits visually annoying blocking artifacts at low-bit-rate. A new post-processing deblocking algorithm in wavelet domain is proposed. The algorithm exploits blocking-artifact features shown in wavelet domain. The energy of blocking artifacts is concentrated into some lines to form annoying visual effects after wavelet transform. The aim of reducing blocking artifacts is to capture excessive energy on the block boundary effectively and reduce it below the visual scope. Adaptive operators for different subbands are computed based on the wavelet coefficients. The operators are made adaptive to different images and characteristics of blocking artifacts. Experimental results show that the proposed method can significantly improve the visual quality and also increase the peak signal-noise-ratio(PSNR) in the output image.展开更多
Convolutional neural networks(CNN)based on U-shaped structures and skip connections play a pivotal role in various image segmentation tasks.Recently,Transformer starts to lead new trends in the image segmentation task...Convolutional neural networks(CNN)based on U-shaped structures and skip connections play a pivotal role in various image segmentation tasks.Recently,Transformer starts to lead new trends in the image segmentation task.Transformer layer can construct the relationship between all pixels,and the two parties can complement each other well.On the basis of these characteristics,we try to combine Transformer pipeline and convolutional neural network pipeline to gain the advantages of both.The image is put into the U-shaped encoder-decoder architecture based on empirical combination of self-attention and convolution,in which skip connections are utilized for localglobal semantic feature learning.At the same time,the image is also put into the convolutional neural network architecture.The final segmentation result will be formed by Mix block which combines both.The mixture model of the convolutional neural network and the Transformer network for road segmentation(MCTNet)can achieve effective segmentation results on KITTI dataset and Unstructured Road Scene(URS)dataset built by ourselves.Codes,self-built datasets and trainable models will be available on http://gffzz188fe103f8f1460asb5qn0ckvbow66v0w.ffgz.tsg.suse.edu.cn/xflxfl1992/MCTNet.展开更多
基金supported by the National Natural Science Foundation of China(Nos.62301092 and 62301093).
摘要Vision Transformers(ViTs)have achieved remarkable success across various artificial intelligence-based computer vision applications.However,their demanding computational and memory requirements pose significant challenges for de-ployment on resource-constrained edge devices.Although post-training quantization(PTQ)provides a promising solution by reducing model precision with minimal calibration data,aggressive low-bit quantization typically leads to substantial perfor-mance degradation.To address this challenge,we present the truncated uniform-log2 quantizer and progressive bit-decline reconstruction method for vision Transformer quantization(TP-ViT).It is an innovative PTQ framework specifically designed for ViTs,featuring two key technical contributions:(1)truncated uniform-log2 quantizer,a novel quantization approach which effectively handles outlier values in post-Softmax activations,significantly reducing quantization errors;(2)bit-decline optimiza-tion strategy,which employs transition weights to gradually reduce bit precision while maintaining model performance under extreme quantization conditions.Comprehensive experiments on image classification,object detection,and instance segmenta-tion tasks demonstrate TP-ViT’s superior performance compared to state-of-the-art PTQ methods,particularly in challenging 3-bit quantization scenarios.Our framework achieves a notable 6.18 percentage points improvement in top-1 accuracy for ViT-small under 3-bit quantization.These results validate TP-ViT’s robustness and general applicability,paving the way for more efficient deployment of ViT models in computer vision applications on edge hardware.
基金supported by the Young Faculty Research Capacity Enhancement Program of Northwest Normal University(NWNU-LKQN2025-27)the China Geological Survey(1212115001601).
摘要The Jiamusi Block in Northeast Asia provides a critical window for investigating the subduction of the Pacific Plate beneath the Western Pacific Orogenic Belt during the Mesozoic and Cenozoic.To constrain the structure and evolution of the lithosphere-asthenosphere system,this study integrates regional gravity and magnetic surveys with two NW-SE-trending magnetotelluric(MT)profiles of lengths 250 and 150 km across the block.Two-dimensional nonlinear conjugate gradient inversion derived the electrical structure to a depth of 200 km.The inversion models reveal distinct crustal and mantle low-resistivity bodies(labelled C1-C8)that spatially correlate with abrupt positive gravity and aeromagnetic anomalies.These low-resistivity features are interpreted as signatures of mantle-derived fluids and partial melts,indicative of asthenospheric upwelling and lithospheric thinning.We attribute these processes to the westward subduction of the Pacific Plate,which induced significant crust-mantle interaction and fluid metasomatism.Furthermore,a comparative analysis reveals that the lithospheric thickness of the Jiamusi Block(70-95 km)is intermediate between that of the Songnen Block(65-80 km)and the Xing'an Block(100-120 km).This variation suggests that lithospheric thickness in the region is primarily controlled by the geometry of the subducting slab's mantle wedge and the intensity of the associated fluid metasomatism.
基金Science and Technology Project of Guangdong Province(2006A10201003)2005 Startup Project of Jinan University(51205067)Soft Science Project of Guangdong Province(2006B70103011)
摘要It is well known that a block discrete cosine transform compressed image exhibits visually annoying blocking artifacts at low-bit-rate. A new post-processing deblocking algorithm in wavelet domain is proposed. The algorithm exploits blocking-artifact features shown in wavelet domain. The energy of blocking artifacts is concentrated into some lines to form annoying visual effects after wavelet transform. The aim of reducing blocking artifacts is to capture excessive energy on the block boundary effectively and reduce it below the visual scope. Adaptive operators for different subbands are computed based on the wavelet coefficients. The operators are made adaptive to different images and characteristics of blocking artifacts. Experimental results show that the proposed method can significantly improve the visual quality and also increase the peak signal-noise-ratio(PSNR) in the output image.
基金supported by the Postgraduate Research&Practice Innovation Program of Jiangsu Province (SJCX21_1427)General Program of Natural Science Research in Jiangsu Universities (21KJB520019).
摘要Convolutional neural networks(CNN)based on U-shaped structures and skip connections play a pivotal role in various image segmentation tasks.Recently,Transformer starts to lead new trends in the image segmentation task.Transformer layer can construct the relationship between all pixels,and the two parties can complement each other well.On the basis of these characteristics,we try to combine Transformer pipeline and convolutional neural network pipeline to gain the advantages of both.The image is put into the U-shaped encoder-decoder architecture based on empirical combination of self-attention and convolution,in which skip connections are utilized for localglobal semantic feature learning.At the same time,the image is also put into the convolutional neural network architecture.The final segmentation result will be formed by Mix block which combines both.The mixture model of the convolutional neural network and the Transformer network for road segmentation(MCTNet)can achieve effective segmentation results on KITTI dataset and Unstructured Road Scene(URS)dataset built by ourselves.Codes,self-built datasets and trainable models will be available on http://gffzz188fe103f8f1460asb5qn0ckvbow66v0w.ffgz.tsg.suse.edu.cn/xflxfl1992/MCTNet.
摘要针对现有方法在腹部中小器官图像分割性能方面存在的不足,提出一种基于局部和全局并行编码的网络模型用于腹部多器官图像分割.首先,设计一种提取多尺度特征信息的局部编码分支;其次,全局特征编码分支采用分块Transformer,通过块内Transformer和块间Transformer的组合,既捕获了全局的长距离依赖信息又降低了计算量;再次,设计特征融合模块,以融合来自两条编码分支的上下文信息;最后,设计解码模块,实现全局信息与局部上下文信息的交互,更好地补偿解码阶段的信息损失.在Synapse多器官CT数据集上进行实验,与目前9种先进方法相比,在平均Dice相似系数(DSC)和Hausdorff距离(HD)指标上都达到了最佳性能,分别为83.10%和17.80 mm.