Driven by the high penetration of renewable energy,the inherent intermittency of photovoltaic(PV)generation poses severe challenges to grid stability.To manage this volatility and ensure reliable grid integration,prec...Driven by the high penetration of renewable energy,the inherent intermittency of photovoltaic(PV)generation poses severe challenges to grid stability.To manage this volatility and ensure reliable grid integration,precise PV system modeling and power forecasting have emerged as critical solutions.However,existing research predominantly focuses on algorithmic innovations and model architectures,frequently overlooking the foundational role of dataset selection.Because capturing the complex spatiotemporal dynamics of solar generation increasingly requires the integration of diverse data types,understanding how to select and fuse these multimodal sources is crucial for determining the upper bound of predictive performance.To address the persistent fragmentation of data resources in PV predictive modeling,this paper delivers a comprehensive taxonomy of publicly available benchmark datasets,establishing a roadmap for future data-driven research.We categorize these valuable resources into three core pillars:1)meteorological datasets(encompassing observational,synthetic,hybrid,and reanalysis types);2)PV generation datasets(grouped by temporal resolution);and 3)static system parameters(including plant-level geospatial data and module-level physical properties).Building upon this categorization,this review thoroughly examines multimodal data fusion strategies across various forecasting horizons and elucidates the specific data dependencies of persistence,physical,and data-driven modeling paradigms.Furthermore,we critically analyze key challenges in multi-source data fusion,particularly spatiotemporal misalignment and the lack of standardized quality control flags.Ultimately,this work provides researchers with an authoritative guide for robust data selection and model construction.展开更多
Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using d...Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using deep learning algorithms further enhances performance;nevertheless,it is challenging due to the nature of small-scale and imbalanced PD datasets.This paper proposed a convolutional neural network-based deep support vector machine(CNN-DSVM)to automate the feature extraction process using CNN and extend the conventional SVM to a DSVM for better classification performance in small-scale PD datasets.A customized kernel function reduces the impact of biased classification towards the majority class(healthy candidates in our consideration).An improved generative adversarial network(IGAN)was designed to generate additional training data to enhance the model’s performance.For performance evaluation,the proposed algorithm achieves a sensitivity of 97.6%and a specificity of 97.3%.The performance comparison is evaluated from five perspectives,including comparisons with different data generation algorithms,feature extraction techniques,kernel functions,and existing works.Results reveal the effectiveness of the IGAN algorithm,which improves the sensitivity and specificity by 4.05%–4.72%and 4.96%–5.86%,respectively;and the effectiveness of the CNN-DSVM algorithm,which improves the sensitivity by 1.24%–57.4%and specificity by 1.04%–163%and reduces biased detection towards the majority class.The ablation experiments confirm the effectiveness of individual components.Two future research directions have also been suggested.展开更多
Malware has evolved from the early Creeper virus into highly sophisticated and organized cyber threats.Over time,it grew in sophistication,adopting advanced techniques,stealth tactics,and autonomous propagation.Modern...Malware has evolved from the early Creeper virus into highly sophisticated and organized cyber threats.Over time,it grew in sophistication,adopting advanced techniques,stealth tactics,and autonomous propagation.Modern malware leverages encryption,obfuscation,zero-day exploits,and AI-assisted techniques to conduct stealthy and persistent attacks.Classification of its exact family is the end goal to defend and mitigate the latest attacks.Researchers have contributed significantly and introduced many techniques to tackle malware threats.Binary detection is performed at a large scale,but very little in multi-class classification.In this research,a hybrid technique is proposed by combining a sandbox with AI models to extract hidden patterns and classify its category and family with high accuracy.A dataset(AU-PEMAL-2025)is prepared,which includes 10,839 records of 26 malware families.Five ML and three DL models are trained on the newly created dataset to validate its effectiveness.The ML classifiers achieved the highest accuracies of 0.9945,0.9788,and 0.9485,while the DL models achieved 0.9932,0.9591,and 0.9286 accuracies with minimal losses in detection and multi-class classification of category and family,respectively.Our findings reveal that the proposed approach can efficiently detect the obfuscated malware variants and safeguard organizations from unseen malware threats.展开更多
Data-driven autonomous driving is a hot topic in academic and industry research due to its impressive performance,flexible mobility,and reduced human intervention.However,the development of this technology relies heav...Data-driven autonomous driving is a hot topic in academic and industry research due to its impressive performance,flexible mobility,and reduced human intervention.However,the development of this technology relies heavily on large datasets that contain accurately annotated data,obtained through artificial or semi-automated strategies.Consequently,datasets play a crucial role in autonomous driving,and their characteristics significantly impact the effectiveness of algorithms.Currently,there are several diverse datasets available,such as KITTI and City Scape,that cover various tasks.However,researchers often overlook the unique features,similarities,and specificities of these datasets.Furthermore,to the best of our knowledge,there is a lack of survey articles focusing on special metrics and benchmark performance on different datasets in autonomous driving.Therefore,the purpose of this article is to analyze autonomous driving datasets,guide researchers on collecting and utilizing relevant datasets,summarize evaluation strategies,analyze benchmark performance,and provide future research points to enrich the autonomous driving community.We believe that this work will assist researchers in evaluating their data using suitable metrics and offer a fresh perspective on autonomous driving.展开更多
In subsalt hydrocarbon exploration,the strong velocity contrast associated with salt structures poses significant challenges to conventional full waveform inversion(FWI)method.While direct envelope inversion can inver...In subsalt hydrocarbon exploration,the strong velocity contrast associated with salt structures poses significant challenges to conventional full waveform inversion(FWI)method.While direct envelope inversion can invert large-scale salt dome,it fails to invert the salt-bottom velocity structures due to the absence of waveform phase information.To solve this problem,we first use a sliding Gaussian window to decompose seismic data into the local scale waveform.Subsequently,we combine the local scale envelope signal with waveform instantaneous phase to obtain the polarity envelope.The resulting polarity envelope can invert low-frequency components while preserving the phase characteristics of seismic data,enabling more accurate low-wavenumber velocity structures.Based on this,we propose a phase-based polarized direct envelope inversion with total variation regularization for simultaneous source seismic data(simultaneous source TV-PDEI)to improve the accuracy of velocity inversion,remove the crosstalk noise,and enhance the computational efficiency.Numerical experiments on salt models and Chevron blind dataset test demonstrate that the simultaneous source TV-PDEI efficiently provides a robust initial velocity model for FWI.展开更多
This paper presents a systematic survey of machine vision-based surface defect detection technologies,focusing on five core challenges in the field:interference from complex backgrounds,small object detection,class im...This paper presents a systematic survey of machine vision-based surface defect detection technologies,focusing on five core challenges in the field:interference from complex backgrounds,small object detection,class imbalance,dynamic scene modeling,and cross-scenario generalization.It reviews key technical approaches corresponding to these challenges over the past five years.Furthermore,a dataset characterization analysis framework is established around these challenges,summarizing and comparing the characteristics of over 40 publicly available datasets across more than ten scenarios,including PCB,photovoltaic,metal,and pavement surfaces.Quantitative selection metrics(such as the small target coefficient and texture complexity)are proposed for challenges like small target detection and complex backgrounds,offering a methodological guide for aligning research questions with benchmark data.Finally,the paper summarizes current limitations and provides an outlook on new paradigms driven by large-scale models and the construction of high-quality benchmark datasets,aiming to offer valuable references for both research and engineering practices in this field.展开更多
Monitoring concrete cracks for structural health in civil engineering presents a significant challenge.This is primarily due to the reliance on manual investigation methods,impacts of global climatic shifts stress,and...Monitoring concrete cracks for structural health in civil engineering presents a significant challenge.This is primarily due to the reliance on manual investigation methods,impacts of global climatic shifts stress,and geohazard threats to engineering structures.To cope with this challenge,state-of-the-art Deep Learning(DL)models are utilized to predict concrete cracks and accurately identify subtle variations in crack patterns and sizes,which lighting conditions and surface textures can influence.Previous studies indicate that model accuracy may decrease when faced with obscured concrete cracks,irregular shapes,or limited datasets for real-world problem scenarios.Feature fusion enhances model performance by combining complementary information,resulting in more accurate predictions,but may increase complexity and potential information redundancy.The study presents the Fractur Encoder to Decoder(FractED)block,a novel architecture consisting of three sub-blocks:the inner block(Encoder),intermediate block(Intermediate block),and outer block(Decoder).This approach integrates fused features into the model without additional fine-tuning steps,allowing for comprehensive feature refinement and enhancement,ultimately optimizing model performance.The study investigates a DL methodology on three datasets,demonstrating its effectiveness in handling complex classification scenarios in civil engineering.The model achieved high accuracy rates,with 88.41%for multiclass(Deck,Pavement,and Walls)classification tasks,91.94%on the Pillow Dam Borehole image binary dataset,and 99.77%on the Surface Crack binary dataset.The FractED block integration ensures adaptability and scalability,making it valuable for various Artificial Intelligence(AI)applications in civil engineering.The research also provides a scientific foundation for automatizing civil engineering inspection instruments for the future.展开更多
Medical data has specificity compared to other fields of data,and the description of medical data characteristics is still in a qualitative stage.This study included 293 sub-datasets of 138 independent datasets.First,...Medical data has specificity compared to other fields of data,and the description of medical data characteristics is still in a qualitative stage.This study included 293 sub-datasets of 138 independent datasets.First,data preprocessing was performed using methods such as incomplete data removal,inconsistent data normalization,and data integration.Then,the characteristics of 293 research datasets were quantified using 26 indicators in three categories:simple indicators,statistical indicators,and informational indicators.Furthermore,statistical analysis was performed on the above-mentioned quantitative characteristics,and stepwise regression and decision tree methods were used for modeling learning.The characteristics of the biological and medical datasets in the study were compared with those of other fields’datasets.By comparing the results of statistical analysis and learning modeling,the study found that the sample size of medical datasets included in the UCI database analyzed in this paper is small,most within 1000.The harmonic mean or geometric mean of continuous variables is significantly higher than the data from other fields.That is to say,the scope of the continuous variable range is large.This study uses quantitative indicators to describe the characteristics of medical datasets to avoid the decrease in credibility caused by subjective analysis,and lays a foundation for further algorithm applicability research.展开更多
With the deep integration of smart manufacturing and IoT technologies,higher demands are placed on the intelligence and real-time performance of industrial equipment fault detection.For industrial fans,base bolt loose...With the deep integration of smart manufacturing and IoT technologies,higher demands are placed on the intelligence and real-time performance of industrial equipment fault detection.For industrial fans,base bolt loosening faults are difficult to identify through conventional spectrum analysis,and the extreme scarcity of fault data leads to limited training datasets,making traditional deep learning methods inaccurate in fault identification and incapable of detecting loosening severity.This paper employs Bayesian Learning by training on a small fault dataset collected from the actual operation of axial-flow fans in a factory to obtain posterior distribution.This method proposes specific data processing approaches and a configuration of Bayesian Convolutional Neural Network(BCNN).It can effectively improve the model’s generalization ability.Experimental results demonstrate high detection accuracy and alignment with real-world applications,offering practical significance and reference value for industrial fan bolt loosening detection under data-limited conditions.展开更多
With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the mod...With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the model in identifying non-standard semantics due to its semantic complexity and high-risk features.However,existing fine-tuning methods rely heavily on static data selection strategies,making it difficult to adapt to the dynamic evolution of model capabilities,resulting in low training efficiency.This article proposes ADS(Adaptive Dataset Selection),an adaptive framework for selecting data in anomaly text detection.ADS performs model-aware data selection prior to fine-tuning,adapting the initial state of pre-trained language models by selecting samples that are most informative for the target anomaly detection task.Empirical results on mainstream large language model architectures show that ADS significantly compresses data size while still outperforming existing static strategies and mainstream compression methods.When using only 1000 fine-tuning samples,ADS achieves a 92%F1 score,with an accuracy improvement of over 22%compared to the baseline,demonstrating excellent performance.This study proposes an efficient data selection mechanism from the perspective of model capability and dynamic adaptation of data,providing theoretical support and a practical path for fine-tuning large models in low-resource scenarios.展开更多
Cyclohexene is an important raw material for nylon production,and the selective hydrogenation of benzene is a key route for preparing cyclohexene.To promote data sharing and reuse in this field,we collected and standa...Cyclohexene is an important raw material for nylon production,and the selective hydrogenation of benzene is a key route for preparing cyclohexene.To promote data sharing and reuse in this field,we collected and standardized experimental data on the hydrogenation of benzene to cyclohexene from publicly available literature and constructed a comprehensive dataset containing catalyst composition,reaction conditions,and reaction results(conversion,selectivity and yield).This data descriptor details the source,field definitions,generation and processing workflow,quality control,sharing approach and usage recommendations of the dataset,with the aim of providing a reusable data foundation for subsequent statistical analysis,machine learning modeling,experimental design,and catalyst screening.展开更多
Dear Editor,This letter presents techniques to simplify dataset generation for instance segmentation of raw meat products,a critical step toward automating food production lines.Accurate segmentation is essential for ...Dear Editor,This letter presents techniques to simplify dataset generation for instance segmentation of raw meat products,a critical step toward automating food production lines.Accurate segmentation is essential for addressing challenges such as occlusions,indistinct edges,and stacked configurations,which demand large,diverse datasets.To meet these demands,we propose two complementary approaches:a semi-automatic annotation interface using tools like the segment anything model(SAM)and GrabCut and a synthetic data generation pipeline leveraging 3D-scanned models.These methods reduce reliance on real meat,mitigate food waste,and improve scalability.Experimental results demonstrate that incorporating synthetic data enhances segmentation model performance and,when combined with real data,further boosts accuracy,paving the way for more efficient automation in the food industry.展开更多
Accurate recognition of visually similar pest species remains a major challenge in agricultural vision,given that existing datasets often lack sufficient taxonomic structure,confusable categories,and quantitative anal...Accurate recognition of visually similar pest species remains a major challenge in agricultural vision,given that existing datasets often lack sufficient taxonomic structure,confusable categories,and quantitative analysis of class-level visual difficulty.To address these limitations,we present AP60,a taxonomy-guided benchmark dataset for fine-grained pest recognition,comprising 62,091 images from 60 pest categories and organized according to insect taxonomy.A distinctive characteristic of AP60 is the deliberate inclusion of morphologically confusable taxa,which enables more realistic evaluation of recognition models under biologically meaningful fine-grained settings.Beyond dataset construction,we introduce a feature-level confusion analysis framework to characterize the intrinsic visual structure of AP60 from two complementary aspects:intra-class consistency and inter-class overlap.Using ResNet-34 features and cosine similarity,we quantify class-wise representation similarity and relate it to downstream recognition difficulty.Benchmark evaluations were conducted under two complementary settings.In the closed-set setting,12 supervised models achieved an average accuracy of 85.8%and an average F1-score of 85.1%,indicating that AP60 is a challenging yet stable benchmark for standard pest recognition.In the class-disjoint few-shot setting,three representative few-shot methods were evaluated on unseen pest categories,with FLoR achieving the best accuracy of 74.4%under the 5-way 5-shot protocol.These results suggest that AP60 supports both conventional supervised classification and data-efficient recognition of unseen pest categories with limited labeled samples.Further analysis shows that higher intra-class similarity is associated with better class-level accuracy,whereas lower inter-class separability is associated with increased misclassification.Validation on two additional related pest datasets shows that the same relationships remain stable after data expansion,indicating that the proposed analysis is useful not only for performance interpretation but also for identifying classes that may benefit most from targeted dataset refinement.Overall,AP60 serves as both a benchmark dataset for fine-grained pest recognition and a data-centric resource for diagnosing feature confusion in agricultural image classification.展开更多
Gastrointestinal polyps are well-known precursors to colorectal cancer(CRC),making their accurate detection and segmentation during colonoscopy essential for early diagnosis and cancer prevention.Deep learning-based s...Gastrointestinal polyps are well-known precursors to colorectal cancer(CRC),making their accurate detection and segmentation during colonoscopy essential for early diagnosis and cancer prevention.Deep learning-based segmentation models trained on publicly available datasets such as Kvasir-SEG have demonstrated promising performance;however,two key challenges remain:limited robustness across diverse polyp morphologies and endoscopic imaging conditions,and the lack of interpretable decision-making mechanisms that support clinical trust and validation.Many existing centralized segmentation approaches are primarily optimized using overlap-based metrics such as the Dice coefficient and intersection over union(IoU),without adequately analyzing challenging cases such as small,flat,or low-contrast polyps or providing insight into the visual cues influencing model predictions.This study presents an explainable centralized deep learning segmentation model for gastrointestinal polyp segmentation using the Kvasir-SEG dataset.The approach integrates a ResUNet++-Lite encoder-decoder segmentation model with Grad-CAM and masked Grad-CAM visualizations to analyze the spatial regions influencing segmentation predictions.The study focuses on establishing a reproducible and interpretable experimental model that combines systematic preprocessing,data augmentation,centralized training,and explainability analysis.Experimental evaluation on an 80:20 train-test split of the Kvasir-SEG dataset,where data augmentation was applied after splitting,demonstrates stable training behavior and competitive segmentation performance,achieving a pixel accuracy of 0.964,a Dice coefficient of 0.858,and an IoU of 0.791 on the held-out test set.Qualitative explainability results further indicate that the model consistently focuses on anatomically relevant polyp regions.Overall,the study illustrates how segmentation performance and explainable AI techniques can be integrated to support the development of clinically interpretable AI-assisted colonoscopy systems.展开更多
Accurate purchase prediction in e-commerce critically depends on the quality of behavioral features.This paper proposes a layered and interpretable feature engineering framework that organizes user signals into three ...Accurate purchase prediction in e-commerce critically depends on the quality of behavioral features.This paper proposes a layered and interpretable feature engineering framework that organizes user signals into three layers:Basic,Conversion&Stability(efficiency and volatility across actions),and Advanced Interactions&Activity(crossbehavior synergies and intensity).Using real Taobao(Alibaba’s primary e-commerce platform)logs(57,976 records for 10,203 users;25 November–03 December 2017),we conducted a hierarchical,layer-wise evaluation that holds data splits and hyperparameters fixed while varying only the feature set to quantify each layer’s marginal contribution.Across logistic regression(LR),decision tree,random forest,XGBoost,and CatBoost models with stratified 5-fold cross-validation,the performance improvedmonotonically fromBasic to Conversion&Stability to Advanced features.With LR,F1 increased from 0.613(Basic)to 0.962(Advanced);boosted models achieved high discrimination(0.995 AUC Score)and an F1 score up to 0.983.Calibration and precision–recall analyses indicated strong ranking quality and acknowledged potential dataset and period biases given the short(9-day)window.By making feature contributions measurable and reproducible,the framework complements model-centric advances and offers a transparent blueprint for production-grade behavioralmodeling.The code and processed artifacts are publicly available,and future work will extend the validation to longer,seasonal datasets and hybrid approaches that combine automated feature learning with domain-driven design.展开更多
This dataset compiles the HER performance data of 203 non-noble transition metal phosphide(TMP)catalysts,covering detailed information on catalyst preparation(e.g.,phosphating temperature,precursor,synthesis method),c...This dataset compiles the HER performance data of 203 non-noble transition metal phosphide(TMP)catalysts,covering detailed information on catalyst preparation(e.g.,phosphating temperature,precursor,synthesis method),chemical composition(mass fractions of elements such as Ni,Co,Fe,P,Mo,W and Zn),and testing conditions(e.g.,electrolyte type and concentration,electrode substrate).The key parameters for catalytic performance include the overpotential at 10 mA/cm2(η10)and the Tafel slope.This dataset has been rigorously extracted,cleaned,and standardized to ensure a high degree of structure and machine readability.This provides a reliable data foundation for data-driven methods,such as machine learning and statistical modeling,enabling rapid screening and design of high-performance HER catalysts,supporting performance prediction,in-depth structure-activity analysis and the rational development of novel catalysts.展开更多
Metal oxide catalysts have emerged as highly promising materials for the CO2 cycloaddition reaction,owing to their tunable composition,facile separation,reusability and low cost.Previous studies have identified tha...Metal oxide catalysts have emerged as highly promising materials for the CO2 cycloaddition reaction,owing to their tunable composition,facile separation,reusability and low cost.Previous studies have identified that the type and ratio of metal dopants,surface defect characteristics and crystal plane orientation are critical factors affecting catalytic performance.Despite this potential,systematic investigations into metal oxide catalysts for CO2 cycloaddition remain limited and a comprehensive understanding of the underlying reaction mechanisms is hindered by the lack of extensive,well-curated datasets.To address this gap,this study establishes a systematic and comprehensive dataset of metal oxides including layered double hydroxide(LDH)and ZnO catalysts,encompassing variations in metal dopant type and ratio,defect characteristic and crystal plane orientation.Through high-throughput calculations,we have generated a robust multi-dimensional dataset containing elementary reaction energies,vibrational frequencies,Bader charges and density of states.A rigorous two-tiered quality control protocol is applied to both computational parameter settings and output results,ensuring the integrity and reliability of the data.This dataset,providing complete raw calculation files,offers a reliable foundation for exploring catalytic performance,structure-performance relationships and reaction mechanisms of metal oxide catalysts in CO2 cycloaddition.展开更多
Small datasets are often challenging due to their limited sample size.This research introduces a novel solution to these problems:average linkage virtual sample generation(ALVSG).ALVSG leverages the underlying data st...Small datasets are often challenging due to their limited sample size.This research introduces a novel solution to these problems:average linkage virtual sample generation(ALVSG).ALVSG leverages the underlying data structure to create virtual samples,which can be used to augment the original dataset.The ALVSG process consists of two steps.First,an average-linkage clustering technique is applied to the dataset to create a dendrogram.The dendrogram represents the hierarchical structure of the dataset,with each merging operation regarded as a linkage.Next,the linkages are combined into an average-based dataset,which serves as a new representation of the dataset.The second step in the ALVSG process involves generating virtual samples using the average-based dataset.The research project generates a set of 100 virtual samples by uniformly distributing them within the provided boundary.These virtual samples are then added to the original dataset,creating a more extensive dataset with improved generalization performance.The efficacy of the ALVSG approach is validated through resampling experiments and t-tests conducted on two small real-world datasets.The experiments are conducted on three forecasting models:the support vector machine for regression(SVR),the deep learning model(DL),and XGBoost.The results show that the ALVSG approach outperforms the baseline methods in terms of mean square error(MSE),root mean square error(RMSE),and mean absolute error(MAE).展开更多
Reliable vehicle detection in urban traffic environments remains challenging,particularly for fixed-view CCTV systems deployed in Southeast Asian cities,where heterogeneous traffic composition,high traffic density,fre...Reliable vehicle detection in urban traffic environments remains challenging,particularly for fixed-view CCTV systems deployed in Southeast Asian cities,where heterogeneous traffic composition,high traffic density,frequent occlusions,and complex visual conditions are prevalent.The absence of large-scale datasets tailored to such mixed-traffic environments poses a significant limitation to the performance and generalization capability of existing object detection models.To address this gap,this paper presents a large-scale traffic image dataset for real-time vehicle detection in Vietnamese urban environments.The proposed dataset comprises 23,364 images collected from fixed-view CCTV traffic cameras deployed across Da Nang City,a representative urban area exhibiting mixed-traffic patterns commonly observed in Southeast Asian cities.The data cover diverse temporal periods,weather conditions,and traffic density levels encountered in real-world traffic monitoring scenarios.To comprehensively characterize these conditions,over 1.1 million instances are annotated across multiple traffic-related categories,including pedestrians,bicycles,motorbikes,cars,buses,trucks,and traffic lights with explicit signal-state labels.Such fine-grained,multi-class annotations support not only object-level detection but also higher-level traffic scene analysis relevant to intelligent transportation system(ITS)applications,such as traffic flow analysis and signal control.To balance annotation accuracy and scalability,a semi-automatic labeling pipeline is employed.Initial object annotations are generated using a pretrained YOLOv11m model and subsequently refined through systematic manual verification using the CVAT platform.Comprehensive experiments are conducted under the same experimental protocol,using the same YOLOv11m architecture,comprising a pretrained baseline and a version fine-tuned on the proposed dataset with domain-specific data augmentation and optimized hyperparameter settings tailored to fixed-view CCTV conditions.Under the same evaluation setting,the pretrained YOLOv11m achieves a mean Average Precision(mAP)of 0.409;in contrast,fine-tuning on the proposed dataset improves the mAP to 0.788.These results underscore the necessity of localized,context-aware datasets such as the one presented in this work for robust real-time traffic perception in Vietnam and similar Southeast Asian urban contexts.展开更多
Objective To investigate methods for constructing a high-quality instructional dataset for traditional Chinese medicine(TCM)mental disorders and to validate its efficacy.Methods We proposed the Fine-Med-Mental-T&P...Objective To investigate methods for constructing a high-quality instructional dataset for traditional Chinese medicine(TCM)mental disorders and to validate its efficacy.Methods We proposed the Fine-Med-Mental-T&P methodology for constructing high-quality instruction datasets in TCM mental disorders.This approach integrates theoretical knowledge and practical case studies through a dual-track strategy.(i)Theoretical track:textbooks and guidelines on TCM mental disorders were manually segmented.Initial responses were generated using DeepSeek-V3,followed by refinement by the Qwen3-32B model to align the expression with human preferences.A screening algorithm was then applied to select 16000 high-quality instruction pairs.(ii)Practical track:starting from over 600 real clinical case seeds,diagnostic and therapeutic instruction pairs were generated using DeepSeek-V3 and subsequently screened through manual evaluation,resulting in 4000 high-quality practiceoriented instruction pairs.The integration of both tracks yielded the Med-Mental-Instruct-T&P dataset,comprising a total of 20000 instruction pairs.To validate the dataset’s effectiveness,three experimental evaluations(both manual and automated)were conducted:(i)comparative studies to compare the performance of models fine-tuned on different datasets;(ii)benchmarking to compare against mainstream TCM-specific large language models(LLMs);(iii)data ablation study to investigate the relationship between data volume and model performance.Results Experimental results demonstrate the superior performance of T&P-model finetuned on the Med-Mental-Instruct-T&P dataset.In the comparative study,the T&P-model significantly outperformed the baseline models trained solely on self-generated or purely human-curated baseline data.This superiority was evident in both automated metrics(ROUGEL>0.55)and expert manual evaluations(scoring above 7/10 across accuracy).In benchmark comparisons,the T&P-model also excelled against existing mainstream TCM LLMs(e.g.,HuatuoGPT and ZuoyiGPT).It showed particularly strong capabilities in handling diverse clinical presentations,including challenging disorders such as insomnia and coma,showcasing its robustness and versatility.Data ablation studies showed that T&P-model performance had an overall upward trend with minor fluctuations when training data increased from 10%to 50%;beyond 50%,performance improvement slowed significantly,with metrics plateauing and approaching a saturation point.展开更多
基金supported by Smart Grid-National Science and Technology Major Project(No.2025ZD0803600,2025ZD0803601)National Natural Science Foundation of China(No.52307133)+1 种基金Tianjin Metrology Science and Technology Project(No.2024TJMT028)Tianjin Transportation Technology Project(No.2025-76).
摘要Driven by the high penetration of renewable energy,the inherent intermittency of photovoltaic(PV)generation poses severe challenges to grid stability.To manage this volatility and ensure reliable grid integration,precise PV system modeling and power forecasting have emerged as critical solutions.However,existing research predominantly focuses on algorithmic innovations and model architectures,frequently overlooking the foundational role of dataset selection.Because capturing the complex spatiotemporal dynamics of solar generation increasingly requires the integration of diverse data types,understanding how to select and fuse these multimodal sources is crucial for determining the upper bound of predictive performance.To address the persistent fragmentation of data resources in PV predictive modeling,this paper delivers a comprehensive taxonomy of publicly available benchmark datasets,establishing a roadmap for future data-driven research.We categorize these valuable resources into three core pillars:1)meteorological datasets(encompassing observational,synthetic,hybrid,and reanalysis types);2)PV generation datasets(grouped by temporal resolution);and 3)static system parameters(including plant-level geospatial data and module-level physical properties).Building upon this categorization,this review thoroughly examines multimodal data fusion strategies across various forecasting horizons and elucidates the specific data dependencies of persistence,physical,and data-driven modeling paradigms.Furthermore,we critically analyze key challenges in multi-source data fusion,particularly spatiotemporal misalignment and the lack of standardized quality control flags.Ultimately,this work provides researchers with an authoritative guide for robust data selection and model construction.
基金The work described in this paper was fully supported by a grant from Hong Kong Metropolitan University(RIF/2021/05).
摘要Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using deep learning algorithms further enhances performance;nevertheless,it is challenging due to the nature of small-scale and imbalanced PD datasets.This paper proposed a convolutional neural network-based deep support vector machine(CNN-DSVM)to automate the feature extraction process using CNN and extend the conventional SVM to a DSVM for better classification performance in small-scale PD datasets.A customized kernel function reduces the impact of biased classification towards the majority class(healthy candidates in our consideration).An improved generative adversarial network(IGAN)was designed to generate additional training data to enhance the model’s performance.For performance evaluation,the proposed algorithm achieves a sensitivity of 97.6%and a specificity of 97.3%.The performance comparison is evaluated from five perspectives,including comparisons with different data generation algorithms,feature extraction techniques,kernel functions,and existing works.Results reveal the effectiveness of the IGAN algorithm,which improves the sensitivity and specificity by 4.05%–4.72%and 4.96%–5.86%,respectively;and the effectiveness of the CNN-DSVM algorithm,which improves the sensitivity by 1.24%–57.4%and specificity by 1.04%–163%and reduces biased detection towards the majority class.The ablation experiments confirm the effectiveness of individual components.Two future research directions have also been suggested.
基金funded by Ongoing Research Funding Program(ORF-2026-947)King Saud University,Riyadh,Saudi Arabia,National Science and Technology Council under grant number NSTC 113-2622-8-029-004-IE.
摘要Malware has evolved from the early Creeper virus into highly sophisticated and organized cyber threats.Over time,it grew in sophistication,adopting advanced techniques,stealth tactics,and autonomous propagation.Modern malware leverages encryption,obfuscation,zero-day exploits,and AI-assisted techniques to conduct stealthy and persistent attacks.Classification of its exact family is the end goal to defend and mitigate the latest attacks.Researchers have contributed significantly and introduced many techniques to tackle malware threats.Binary detection is performed at a large scale,but very little in multi-class classification.In this research,a hybrid technique is proposed by combining a sandbox with AI models to extract hidden patterns and classify its category and family with high accuracy.A dataset(AU-PEMAL-2025)is prepared,which includes 10,839 records of 26 malware families.Five ML and three DL models are trained on the newly created dataset to validate its effectiveness.The ML classifiers achieved the highest accuracies of 0.9945,0.9788,and 0.9485,while the DL models achieved 0.9932,0.9591,and 0.9286 accuracies with minimal losses in detection and multi-class classification of category and family,respectively.Our findings reveal that the proposed approach can efficiently detect the obfuscated malware variants and safeguard organizations from unseen malware threats.
基金supported by the Joint Funds of the National Natural Science Foundation of China(U24B20162)the National Natural Science Foundation of China(62373356)。
摘要Data-driven autonomous driving is a hot topic in academic and industry research due to its impressive performance,flexible mobility,and reduced human intervention.However,the development of this technology relies heavily on large datasets that contain accurately annotated data,obtained through artificial or semi-automated strategies.Consequently,datasets play a crucial role in autonomous driving,and their characteristics significantly impact the effectiveness of algorithms.Currently,there are several diverse datasets available,such as KITTI and City Scape,that cover various tasks.However,researchers often overlook the unique features,similarities,and specificities of these datasets.Furthermore,to the best of our knowledge,there is a lack of survey articles focusing on special metrics and benchmark performance on different datasets in autonomous driving.Therefore,the purpose of this article is to analyze autonomous driving datasets,guide researchers on collecting and utilizing relevant datasets,summarize evaluation strategies,analyze benchmark performance,and provide future research points to enrich the autonomous driving community.We believe that this work will assist researchers in evaluating their data using suitable metrics and offer a fresh perspective on autonomous driving.
基金supported by the National Natural Science Foundation of China(Grant Nos.42104116,42230803)。
摘要In subsalt hydrocarbon exploration,the strong velocity contrast associated with salt structures poses significant challenges to conventional full waveform inversion(FWI)method.While direct envelope inversion can invert large-scale salt dome,it fails to invert the salt-bottom velocity structures due to the absence of waveform phase information.To solve this problem,we first use a sliding Gaussian window to decompose seismic data into the local scale waveform.Subsequently,we combine the local scale envelope signal with waveform instantaneous phase to obtain the polarity envelope.The resulting polarity envelope can invert low-frequency components while preserving the phase characteristics of seismic data,enabling more accurate low-wavenumber velocity structures.Based on this,we propose a phase-based polarized direct envelope inversion with total variation regularization for simultaneous source seismic data(simultaneous source TV-PDEI)to improve the accuracy of velocity inversion,remove the crosstalk noise,and enhance the computational efficiency.Numerical experiments on salt models and Chevron blind dataset test demonstrate that the simultaneous source TV-PDEI efficiently provides a robust initial velocity model for FWI.
基金supported in part by the Natural Science Foundation of Shaanxi Province of China under Grant 2024JC-YBQN-0695.
摘要This paper presents a systematic survey of machine vision-based surface defect detection technologies,focusing on five core challenges in the field:interference from complex backgrounds,small object detection,class imbalance,dynamic scene modeling,and cross-scenario generalization.It reviews key technical approaches corresponding to these challenges over the past five years.Furthermore,a dataset characterization analysis framework is established around these challenges,summarizing and comparing the characteristics of over 40 publicly available datasets across more than ten scenarios,including PCB,photovoltaic,metal,and pavement surfaces.Quantitative selection metrics(such as the small target coefficient and texture complexity)are proposed for challenges like small target detection and complex backgrounds,offering a methodological guide for aligning research questions with benchmark data.Finally,the paper summarizes current limitations and provides an outlook on new paradigms driven by large-scale models and the construction of high-quality benchmark datasets,aiming to offer valuable references for both research and engineering practices in this field.
基金supported by Fundamental Research Funds of Central University Research Grant no:B240201122-Muhammad Ishfaque under the Post-Doctoral Research Program of Hohai University,Nanjing,Jiangsu Province of China.
摘要Monitoring concrete cracks for structural health in civil engineering presents a significant challenge.This is primarily due to the reliance on manual investigation methods,impacts of global climatic shifts stress,and geohazard threats to engineering structures.To cope with this challenge,state-of-the-art Deep Learning(DL)models are utilized to predict concrete cracks and accurately identify subtle variations in crack patterns and sizes,which lighting conditions and surface textures can influence.Previous studies indicate that model accuracy may decrease when faced with obscured concrete cracks,irregular shapes,or limited datasets for real-world problem scenarios.Feature fusion enhances model performance by combining complementary information,resulting in more accurate predictions,but may increase complexity and potential information redundancy.The study presents the Fractur Encoder to Decoder(FractED)block,a novel architecture consisting of three sub-blocks:the inner block(Encoder),intermediate block(Intermediate block),and outer block(Decoder).This approach integrates fused features into the model without additional fine-tuning steps,allowing for comprehensive feature refinement and enhancement,ultimately optimizing model performance.The study investigates a DL methodology on three datasets,demonstrating its effectiveness in handling complex classification scenarios in civil engineering.The model achieved high accuracy rates,with 88.41%for multiclass(Deck,Pavement,and Walls)classification tasks,91.94%on the Pillow Dam Borehole image binary dataset,and 99.77%on the Surface Crack binary dataset.The FractED block integration ensures adaptability and scalability,making it valuable for various Artificial Intelligence(AI)applications in civil engineering.The research also provides a scientific foundation for automatizing civil engineering inspection instruments for the future.
基金funded by the Qingdao Huanghai University Doctoral Research Foundation Project,grant number 2023boshi02,and Qingdao Huanghai University scientific research project,grant number KYH2025001.
摘要Medical data has specificity compared to other fields of data,and the description of medical data characteristics is still in a qualitative stage.This study included 293 sub-datasets of 138 independent datasets.First,data preprocessing was performed using methods such as incomplete data removal,inconsistent data normalization,and data integration.Then,the characteristics of 293 research datasets were quantified using 26 indicators in three categories:simple indicators,statistical indicators,and informational indicators.Furthermore,statistical analysis was performed on the above-mentioned quantitative characteristics,and stepwise regression and decision tree methods were used for modeling learning.The characteristics of the biological and medical datasets in the study were compared with those of other fields’datasets.By comparing the results of statistical analysis and learning modeling,the study found that the sample size of medical datasets included in the UCI database analyzed in this paper is small,most within 1000.The harmonic mean or geometric mean of continuous variables is significantly higher than the data from other fields.That is to say,the scope of the continuous variable range is large.This study uses quantitative indicators to describe the characteristics of medical datasets to avoid the decrease in credibility caused by subjective analysis,and lays a foundation for further algorithm applicability research.
基金funded by the Zhejiang Provincial Key Science and Technology“LingYan”Project Foundation,grant number 2023C01145Zhejiang Gongshang University Higher Education Research Projects,grant number Xgy22028.
摘要With the deep integration of smart manufacturing and IoT technologies,higher demands are placed on the intelligence and real-time performance of industrial equipment fault detection.For industrial fans,base bolt loosening faults are difficult to identify through conventional spectrum analysis,and the extreme scarcity of fault data leads to limited training datasets,making traditional deep learning methods inaccurate in fault identification and incapable of detecting loosening severity.This paper employs Bayesian Learning by training on a small fault dataset collected from the actual operation of axial-flow fans in a factory to obtain posterior distribution.This method proposes specific data processing approaches and a configuration of Bayesian Convolutional Neural Network(BCNN).It can effectively improve the model’s generalization ability.Experimental results demonstrate high detection accuracy and alignment with real-world applications,offering practical significance and reference value for industrial fan bolt loosening detection under data-limited conditions.
摘要With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the model in identifying non-standard semantics due to its semantic complexity and high-risk features.However,existing fine-tuning methods rely heavily on static data selection strategies,making it difficult to adapt to the dynamic evolution of model capabilities,resulting in low training efficiency.This article proposes ADS(Adaptive Dataset Selection),an adaptive framework for selecting data in anomaly text detection.ADS performs model-aware data selection prior to fine-tuning,adapting the initial state of pre-trained language models by selecting samples that are most informative for the target anomaly detection task.Empirical results on mainstream large language model architectures show that ADS significantly compresses data size while still outperforming existing static strategies and mainstream compression methods.When using only 1000 fine-tuning samples,ADS achieves a 92%F1 score,with an accuracy improvement of over 22%compared to the baseline,demonstrating excellent performance.This study proposes an efficient data selection mechanism from the perspective of model capability and dynamic adaptation of data,providing theoretical support and a practical path for fine-tuning large models in low-resource scenarios.
基金Supported by Strategic Priority Research Program of the Chinese Academy of Sciences(XDB 1190000)the General Program of the National Natural Science Foundation of China(22572209)。
摘要Cyclohexene is an important raw material for nylon production,and the selective hydrogenation of benzene is a key route for preparing cyclohexene.To promote data sharing and reuse in this field,we collected and standardized experimental data on the hydrogenation of benzene to cyclohexene from publicly available literature and constructed a comprehensive dataset containing catalyst composition,reaction conditions,and reaction results(conversion,selectivity and yield).This data descriptor details the source,field definitions,generation and processing workflow,quality control,sharing approach and usage recommendations of the dataset,with the aim of providing a reusable data foundation for subsequent statistical analysis,machine learning modeling,experimental design,and catalyst screening.
基金supported by European Union’s Horizon Europe research and innovation programme,project AGILEHAND(Smart Grading,Handling and Packaging Solutions for Soft and Deformable Products in Agile and Reconfigurable Lines)(101092043).
摘要Dear Editor,This letter presents techniques to simplify dataset generation for instance segmentation of raw meat products,a critical step toward automating food production lines.Accurate segmentation is essential for addressing challenges such as occlusions,indistinct edges,and stacked configurations,which demand large,diverse datasets.To meet these demands,we propose two complementary approaches:a semi-automatic annotation interface using tools like the segment anything model(SAM)and GrabCut and a synthetic data generation pipeline leveraging 3D-scanned models.These methods reduce reliance on real meat,mitigate food waste,and improve scalability.Experimental results demonstrate that incorporating synthetic data enhances segmentation model performance and,when combined with real data,further boosts accuracy,paving the way for more efficient automation in the food industry.
基金financially supported by Zhejiang Provincial Natural Science Foundation of China(ZCLZ24F0201)National Natural Science Foundation of China(Project Nos.:41901268,62276086)National Key R&D Program of China(2022YFD2000100).
摘要Accurate recognition of visually similar pest species remains a major challenge in agricultural vision,given that existing datasets often lack sufficient taxonomic structure,confusable categories,and quantitative analysis of class-level visual difficulty.To address these limitations,we present AP60,a taxonomy-guided benchmark dataset for fine-grained pest recognition,comprising 62,091 images from 60 pest categories and organized according to insect taxonomy.A distinctive characteristic of AP60 is the deliberate inclusion of morphologically confusable taxa,which enables more realistic evaluation of recognition models under biologically meaningful fine-grained settings.Beyond dataset construction,we introduce a feature-level confusion analysis framework to characterize the intrinsic visual structure of AP60 from two complementary aspects:intra-class consistency and inter-class overlap.Using ResNet-34 features and cosine similarity,we quantify class-wise representation similarity and relate it to downstream recognition difficulty.Benchmark evaluations were conducted under two complementary settings.In the closed-set setting,12 supervised models achieved an average accuracy of 85.8%and an average F1-score of 85.1%,indicating that AP60 is a challenging yet stable benchmark for standard pest recognition.In the class-disjoint few-shot setting,three representative few-shot methods were evaluated on unseen pest categories,with FLoR achieving the best accuracy of 74.4%under the 5-way 5-shot protocol.These results suggest that AP60 supports both conventional supervised classification and data-efficient recognition of unseen pest categories with limited labeled samples.Further analysis shows that higher intra-class similarity is associated with better class-level accuracy,whereas lower inter-class separability is associated with increased misclassification.Validation on two additional related pest datasets shows that the same relationships remain stable after data expansion,indicating that the proposed analysis is useful not only for performance interpretation but also for identifying classes that may benefit most from targeted dataset refinement.Overall,AP60 serves as both a benchmark dataset for fine-grained pest recognition and a data-centric resource for diagnosing feature confusion in agricultural image classification.
摘要Gastrointestinal polyps are well-known precursors to colorectal cancer(CRC),making their accurate detection and segmentation during colonoscopy essential for early diagnosis and cancer prevention.Deep learning-based segmentation models trained on publicly available datasets such as Kvasir-SEG have demonstrated promising performance;however,two key challenges remain:limited robustness across diverse polyp morphologies and endoscopic imaging conditions,and the lack of interpretable decision-making mechanisms that support clinical trust and validation.Many existing centralized segmentation approaches are primarily optimized using overlap-based metrics such as the Dice coefficient and intersection over union(IoU),without adequately analyzing challenging cases such as small,flat,or low-contrast polyps or providing insight into the visual cues influencing model predictions.This study presents an explainable centralized deep learning segmentation model for gastrointestinal polyp segmentation using the Kvasir-SEG dataset.The approach integrates a ResUNet++-Lite encoder-decoder segmentation model with Grad-CAM and masked Grad-CAM visualizations to analyze the spatial regions influencing segmentation predictions.The study focuses on establishing a reproducible and interpretable experimental model that combines systematic preprocessing,data augmentation,centralized training,and explainability analysis.Experimental evaluation on an 80:20 train-test split of the Kvasir-SEG dataset,where data augmentation was applied after splitting,demonstrates stable training behavior and competitive segmentation performance,achieving a pixel accuracy of 0.964,a Dice coefficient of 0.858,and an IoU of 0.791 on the held-out test set.Qualitative explainability results further indicate that the model consistently focuses on anatomically relevant polyp regions.Overall,the study illustrates how segmentation performance and explainable AI techniques can be integrated to support the development of clinically interpretable AI-assisted colonoscopy systems.
基金supported by the research fund of Hanyang University(HY-202500000001616).
摘要Accurate purchase prediction in e-commerce critically depends on the quality of behavioral features.This paper proposes a layered and interpretable feature engineering framework that organizes user signals into three layers:Basic,Conversion&Stability(efficiency and volatility across actions),and Advanced Interactions&Activity(crossbehavior synergies and intensity).Using real Taobao(Alibaba’s primary e-commerce platform)logs(57,976 records for 10,203 users;25 November–03 December 2017),we conducted a hierarchical,layer-wise evaluation that holds data splits and hyperparameters fixed while varying only the feature set to quantify each layer’s marginal contribution.Across logistic regression(LR),decision tree,random forest,XGBoost,and CatBoost models with stratified 5-fold cross-validation,the performance improvedmonotonically fromBasic to Conversion&Stability to Advanced features.With LR,F1 increased from 0.613(Basic)to 0.962(Advanced);boosted models achieved high discrimination(0.995 AUC Score)and an F1 score up to 0.983.Calibration and precision–recall analyses indicated strong ranking quality and acknowledged potential dataset and period biases given the short(9-day)window.By making feature contributions measurable and reproducible,the framework complements model-centric advances and offers a transparent blueprint for production-grade behavioralmodeling.The code and processed artifacts are publicly available,and future work will extend the validation to longer,seasonal datasets and hybrid approaches that combine automated feature learning with domain-driven design.
基金Supported by the Fundamental Research Funds for the Central Universities(the special project of the doctoral research innovation plan of the China People's Police University(XJ2024002203))。
摘要This dataset compiles the HER performance data of 203 non-noble transition metal phosphide(TMP)catalysts,covering detailed information on catalyst preparation(e.g.,phosphating temperature,precursor,synthesis method),chemical composition(mass fractions of elements such as Ni,Co,Fe,P,Mo,W and Zn),and testing conditions(e.g.,electrolyte type and concentration,electrode substrate).The key parameters for catalytic performance include the overpotential at 10 mA/cm2(η10)and the Tafel slope.This dataset has been rigorously extracted,cleaned,and standardized to ensure a high degree of structure and machine readability.This provides a reliable data foundation for data-driven methods,such as machine learning and statistical modeling,enabling rapid screening and design of high-performance HER catalysts,supporting performance prediction,in-depth structure-activity analysis and the rational development of novel catalysts.
基金Supported by the National Natural Science Foundation of China(U23A20132)。
摘要Metal oxide catalysts have emerged as highly promising materials for the CO2 cycloaddition reaction,owing to their tunable composition,facile separation,reusability and low cost.Previous studies have identified that the type and ratio of metal dopants,surface defect characteristics and crystal plane orientation are critical factors affecting catalytic performance.Despite this potential,systematic investigations into metal oxide catalysts for CO2 cycloaddition remain limited and a comprehensive understanding of the underlying reaction mechanisms is hindered by the lack of extensive,well-curated datasets.To address this gap,this study establishes a systematic and comprehensive dataset of metal oxides including layered double hydroxide(LDH)and ZnO catalysts,encompassing variations in metal dopant type and ratio,defect characteristic and crystal plane orientation.Through high-throughput calculations,we have generated a robust multi-dimensional dataset containing elementary reaction energies,vibrational frequencies,Bader charges and density of states.A rigorous two-tiered quality control protocol is applied to both computational parameter settings and output results,ensuring the integrity and reliability of the data.This dataset,providing complete raw calculation files,offers a reliable foundation for exploring catalytic performance,structure-performance relationships and reaction mechanisms of metal oxide catalysts in CO2 cycloaddition.
摘要Small datasets are often challenging due to their limited sample size.This research introduces a novel solution to these problems:average linkage virtual sample generation(ALVSG).ALVSG leverages the underlying data structure to create virtual samples,which can be used to augment the original dataset.The ALVSG process consists of two steps.First,an average-linkage clustering technique is applied to the dataset to create a dendrogram.The dendrogram represents the hierarchical structure of the dataset,with each merging operation regarded as a linkage.Next,the linkages are combined into an average-based dataset,which serves as a new representation of the dataset.The second step in the ALVSG process involves generating virtual samples using the average-based dataset.The research project generates a set of 100 virtual samples by uniformly distributing them within the provided boundary.These virtual samples are then added to the original dataset,creating a more extensive dataset with improved generalization performance.The efficacy of the ALVSG approach is validated through resampling experiments and t-tests conducted on two small real-world datasets.The experiments are conducted on three forecasting models:the support vector machine for regression(SVR),the deep learning model(DL),and XGBoost.The results show that the ALVSG approach outperforms the baseline methods in terms of mean square error(MSE),root mean square error(RMSE),and mean absolute error(MAE).
基金supported in part by the Ministry of Science and Technology(MOST)The University of Danang—University of TechnologyEducation,and the School of Computer Science,Duy Tan University,Da Nang City,Vietnam.
摘要Reliable vehicle detection in urban traffic environments remains challenging,particularly for fixed-view CCTV systems deployed in Southeast Asian cities,where heterogeneous traffic composition,high traffic density,frequent occlusions,and complex visual conditions are prevalent.The absence of large-scale datasets tailored to such mixed-traffic environments poses a significant limitation to the performance and generalization capability of existing object detection models.To address this gap,this paper presents a large-scale traffic image dataset for real-time vehicle detection in Vietnamese urban environments.The proposed dataset comprises 23,364 images collected from fixed-view CCTV traffic cameras deployed across Da Nang City,a representative urban area exhibiting mixed-traffic patterns commonly observed in Southeast Asian cities.The data cover diverse temporal periods,weather conditions,and traffic density levels encountered in real-world traffic monitoring scenarios.To comprehensively characterize these conditions,over 1.1 million instances are annotated across multiple traffic-related categories,including pedestrians,bicycles,motorbikes,cars,buses,trucks,and traffic lights with explicit signal-state labels.Such fine-grained,multi-class annotations support not only object-level detection but also higher-level traffic scene analysis relevant to intelligent transportation system(ITS)applications,such as traffic flow analysis and signal control.To balance annotation accuracy and scalability,a semi-automatic labeling pipeline is employed.Initial object annotations are generated using a pretrained YOLOv11m model and subsequently refined through systematic manual verification using the CVAT platform.Comprehensive experiments are conducted under the same experimental protocol,using the same YOLOv11m architecture,comprising a pretrained baseline and a version fine-tuned on the proposed dataset with domain-specific data augmentation and optimized hyperparameter settings tailored to fixed-view CCTV conditions.Under the same evaluation setting,the pretrained YOLOv11m achieves a mean Average Precision(mAP)of 0.409;in contrast,fine-tuning on the proposed dataset improves the mAP to 0.788.These results underscore the necessity of localized,context-aware datasets such as the one presented in this work for robust real-time traffic perception in Vietnam and similar Southeast Asian urban contexts.
基金Key Scientific Research Project of the Hunan Provincial Department of Education(23A312).
摘要Objective To investigate methods for constructing a high-quality instructional dataset for traditional Chinese medicine(TCM)mental disorders and to validate its efficacy.Methods We proposed the Fine-Med-Mental-T&P methodology for constructing high-quality instruction datasets in TCM mental disorders.This approach integrates theoretical knowledge and practical case studies through a dual-track strategy.(i)Theoretical track:textbooks and guidelines on TCM mental disorders were manually segmented.Initial responses were generated using DeepSeek-V3,followed by refinement by the Qwen3-32B model to align the expression with human preferences.A screening algorithm was then applied to select 16000 high-quality instruction pairs.(ii)Practical track:starting from over 600 real clinical case seeds,diagnostic and therapeutic instruction pairs were generated using DeepSeek-V3 and subsequently screened through manual evaluation,resulting in 4000 high-quality practiceoriented instruction pairs.The integration of both tracks yielded the Med-Mental-Instruct-T&P dataset,comprising a total of 20000 instruction pairs.To validate the dataset’s effectiveness,three experimental evaluations(both manual and automated)were conducted:(i)comparative studies to compare the performance of models fine-tuned on different datasets;(ii)benchmarking to compare against mainstream TCM-specific large language models(LLMs);(iii)data ablation study to investigate the relationship between data volume and model performance.Results Experimental results demonstrate the superior performance of T&P-model finetuned on the Med-Mental-Instruct-T&P dataset.In the comparative study,the T&P-model significantly outperformed the baseline models trained solely on self-generated or purely human-curated baseline data.This superiority was evident in both automated metrics(ROUGEL>0.55)and expert manual evaluations(scoring above 7/10 across accuracy).In benchmark comparisons,the T&P-model also excelled against existing mainstream TCM LLMs(e.g.,HuatuoGPT and ZuoyiGPT).It showed particularly strong capabilities in handling diverse clinical presentations,including challenging disorders such as insomnia and coma,showcasing its robustness and versatility.Data ablation studies showed that T&P-model performance had an overall upward trend with minor fluctuations when training data increased from 10%to 50%;beyond 50%,performance improvement slowed significantly,with metrics plateauing and approaching a saturation point.