To address the difficulty in recognizing subtle differences in facial biomarkers in children with autism,a learnable positional encoding enhancement(LPEE)module was combined with the adaptive token aggregation(ATA)mod...To address the difficulty in recognizing subtle differences in facial biomarkers in children with autism,a learnable positional encoding enhancement(LPEE)module was combined with the adaptive token aggregation(ATA)module.The vision transformer with learnable positional encoding and adaptive token aggregation(Vi T-LPATA),a predictive model for autism,was proposed.The model leverages the LPEE module to dynamically capture facial geometric deformation features and integrates the ATA module to enhance the feature representation capability of pathological regions,thereby establishing precise mappings of biomarker differences.Experiments on a publicly available autism facial dataset demonstrated that the Vi T-LPATA achieved optimal performance,with 99.2%accuracy and an area under the curve(AUC)value of 0.940.展开更多
基金supported by the Scientific Research Project of Hunan Provincial Education DepartmentChina(No.23A0423)+1 种基金the Natural Science Foundation of Hunan ProvinceChina(No.2025JJ70029)。
摘要To address the difficulty in recognizing subtle differences in facial biomarkers in children with autism,a learnable positional encoding enhancement(LPEE)module was combined with the adaptive token aggregation(ATA)module.The vision transformer with learnable positional encoding and adaptive token aggregation(Vi T-LPATA),a predictive model for autism,was proposed.The model leverages the LPEE module to dynamically capture facial geometric deformation features and integrates the ATA module to enhance the feature representation capability of pathological regions,thereby establishing precise mappings of biomarker differences.Experiments on a publicly available autism facial dataset demonstrated that the Vi T-LPATA achieved optimal performance,with 99.2%accuracy and an area under the curve(AUC)value of 0.940.