期刊文献+
共找到150篇文章
< 1 2 8 >
每页显示 20 50 100
EGAIN: Enhanced Generative Adversarial Networks for Imputing Missing Values 认领 引用
1
作者 Abolfazl Saghafi Soodeh Moallemian +1 位作者 Miray Budak Rutvik Deshpande 《Computers, Materials & Continua》 SCIE EI 2026年第8期2241-2255,共15页
Missing data remain a persistent challenge in statistical analysis and machine learning because many predictive methods require complete observations.Generative Adversarial Imputation Networks(GAIN)offer a flexible de... Missing data remain a persistent challenge in statistical analysis and machine learning because many predictive methods require complete observations.Generative Adversarial Imputation Networks(GAIN)offer a flexible deep-learning approach for missing value imputation,but their practical use is limited by convergence instability,sensitivity to hyperparameter selection,and dependence on outdated software implementations.To address these limitations,we propose Enhanced Generative Adversarial Imputation Networks(EGAIN),a modernized extension of GAIN implemented in TensorFlow 2.x.EGAIN incorporates convolution-based generator and discriminator networks,a channel-stacked representation of the data and mask,and checkpoint-based training diagnostics to improve stability and usability.EGAIN was evaluated on five benchmark datasets under multiple Missing Completely At Random(MCAR)settings and compared with the original GAIN implementation and median imputation.Across most evaluated conditions,EGAIN achieved lower root mean squared error(RMSE)and showed greater robustness,particularly when missingness was concentrated in a subset of variables.These results indicate that EGAIN provides a more stable and reproducible framework for missing data imputation in tabular datasets. 展开更多
关键词 Missing value imputation generative adversarial network tabular data imputation missing completely at random convolutional architectures training stability
暂未订购 下载PDF
Imputing missing values using cumulative linear regression 认领 引用 被引量:4
2
作者 Samih M. Mostafa 《CAAI Transactions on Intelligence Technology》 SCIE EI 2019年第3期182-200,共19页
The concept of missing data is important to apply statistical methods on the dataset. Statisticians and researchers may end up to an inaccurate illation about the data if the missing data are not handled properly. Of ... The concept of missing data is important to apply statistical methods on the dataset. Statisticians and researchers may end up to an inaccurate illation about the data if the missing data are not handled properly. Of late, Python and R provide diverse packages for handling missing data. In this study, an imputation algorithm, cumulative linear regression, is proposed. The proposed algorithm depends on the linear regression technique. It differs from the existing methods, in that it cumulates the imputed variables;those variables will be incorporated in the linear regression equation to filling in the missing values in the next incomplete variable. The author performed a comparative study of the proposed method and those packages. The performance was measured in terms of imputation time, root-mean-square error, mean absolute error, and coefficient of determination (R^2). On analysing on five datasets with different missing values generated from different mechanisms, it was observed that the performances vary depending on the size, missing percentage, and the missingness mechanism. The results showed that the performance of the proposed method is slightly better. 展开更多
关键词 Imputing missing values cumulative linear regression statistical methods
暂未订购 下载PDF
Imputing the long-term missing heating load data using a generative network 认领 引用
3
作者 Mengbo Yu Alexander Neubauer +2 位作者 Pedram Babakhani Stefan Brandt Martin Kriegel 《Energy and AI》 EI CSCD 2025年第4期453-472,共20页
Accurately filling in missing heating data is essential for ensuring data quality in applications such as energy management optimization and building efficiency analysis.Traditional machine learning methods use histor... Accurately filling in missing heating data is essential for ensuring data quality in applications such as energy management optimization and building efficiency analysis.Traditional machine learning methods use historical heating data as an input feature to predict the following missing data.However,when the duration of missing data is long,previous estimated values are inevitably used for further imputation,leading to error accumulation and a growing deviation from true values.To overcome this problem,this paper proposes a generative network that can fill missing data solely based on weather and temporal data,without using previous imputed values for further imputation.Our method outperformed the state of the art such as Seq2seq and Transformer,achieving relative normalized root mean square error(NRMSE)reductions of 1.65%to 41.38%,0.30%to 66.43%,and 14.84%to 50.22%across three different data sources.In addition,with our proposed method,the effect of selecting different weather variables on model performance,and the benefits of transfer learning under limited data were also demonstrated.The relative NRMSE reduction is between 3.88%to 15.85%in cold months and from 7.49%to 12.29%in warm months when applying transfer learning. 展开更多
关键词 Generative network Heating load data Missing data imputation Transfer learning
Imputing not available values in single-cell DNA methylation data using the median is straightforward and effective 认领 引用
4
作者 Songming Tang Siyu Li Shengquan Chen 《Quantitative Biology》 CAS CSCD 2025年第3期127-130,共4页
Recent advances in single-cell DNA methylation have provided unprecedented opportunities to explore cellular epigenetic differences with maximal resolution.A common workflow for single-cell DNA methylation analysis is... Recent advances in single-cell DNA methylation have provided unprecedented opportunities to explore cellular epigenetic differences with maximal resolution.A common workflow for single-cell DNA methylation analysis is binning the genome into multiple regions and computing the average methylation level within each region.In this process,imputing not available(NA)values which are caused by the limited number of captured methylation sites is a necessary preprocessing step for downstream analyses.Existing studies have employed several simple imputation methods(such as zeros imputation or means imputation),however,there is a lack of theoretical studies or benchmark tests of these approaches.Through both experiments and theoretical analysis,we found that using the medians to impute NA values can effectively and simply reflect the methylation state of the NA values,providing an accurate foundation for downstream analyses. 展开更多
关键词 data imputation single-cell DNA methylation
Sequence-based genome-wide association study reveals genetic and metabolic mechanisms underlying feed efficiency-related traits in beef cattle 认领 引用
5
作者 Leonardo M.Arikawa Lucio F.M.Mota +11 位作者 Larissa F.S.Fonseca Gerardo A.Fernandes Júnior Bruna M.Salatta Gabriela B.Frezarim Patricia I.Schmidt Sindy L.C.Nasner Julia P.S.Valente Amalia M.Pelaez Roberta C.Canesin Josineudson A.Ⅱ.Ⅴ.Silva Maria Eugênia Z.Mercadante Lucia G.Albuquerque 《Journal of Animal Science and Biotechnology》 SCIE CAS CSCD 2026年第3期1264-1283,共20页
Background Efficiency is characterized by maximum productivity with lower inputs and minimal waste,resulting in greater output with the same or even fewer resources.In livestock,more efficient animals in converting fo... Background Efficiency is characterized by maximum productivity with lower inputs and minimal waste,resulting in greater output with the same or even fewer resources.In livestock,more efficient animals in converting food into protein may improve the economic efficiency of production systems,as feed costs represent a significant expense in beef production.Thus,the present study aimed to use imputed whole-genome sequencing(WGS)data to perform a genome-wide association study(GWAS)in order to identify genomic regions and potential candidate genes involved in the biological processes and metabolic pathways associated with feed efficiency-related traits(RFI:residual feed intake,DMI:dry matter intake,FE:feed efficiency,FC:feed conversion,and RWG:residual weight gain)in Nellore cattle.Results The GWAS identified significant SNPs associated with feed efficiency traits in Nellore cattle.A total of 42 SNPs were detected for RFI,10 for DMI,99 for FC,15 for FE,and 3 for RWG,distributed in different autosomes.Annotation analysis identified several candidate genes,and the prioritization highlighted 21,9,68,23,and 8 key genes for RFI,DMI,FC,FE,and RWG,respectively.The prioritized candidate genes are involved in muscle development,lipid metabolism,response to oxidative stress,nutrient metabolism,neurotransmission,and oxidative phosphorylation.Additionally,enrichment analysis indicated that these genes act in several signaling pathways related to signal transduction,the nervous system,the endocrine system,energy metabolism,the digestive system,and nutrient metabolism.Conclusion The use of imputed WGS data in GWAS analyses enabled the broad identification of regions and candidate genes throughout the genome that regulate expression of feed efficiency-related traits in Nellore cattle.Our results provide new perspectives into the molecular mechanisms underlying feed efficiency in Nellore cattle,offering a genetic basis to guide the breeding of efficient animals,thereby optimizing resource utilization and the profitability of production systems. 展开更多
关键词 Bos indicus Energy metabolism Imputed WGS Muscle development Oxidative phosphorylation Residual feed intake
暂未订购 下载PDF
Optimization of Genotype Phasing and Imputation Strategies for Cost-Effective Genomics in Sebastes schlegelii 认领 引用
6
作者 LIU Zeyu LUO Ziyang +9 位作者 SONG Weihao LI Yi YUE Xinlu YANG Ruiyan ZHANG Fengyan GAO Xiangyu SONG Zongcheng QI Jie ZHANG Quanqi HE Yan 《Journal of Ocean University of China》 SCIE CAS CSCD 2026年第4期1457-1468,I0060-I0062,共12页
The growing application of genome-wide selection and genome-wide association studies(GWAS)based on single-nucleotide polymorphisms(SNPs)has necessitated the design of cost-effective strategies.High-depth sequencing,wh... The growing application of genome-wide selection and genome-wide association studies(GWAS)based on single-nucleotide polymorphisms(SNPs)has necessitated the design of cost-effective strategies.High-depth sequencing,while accurate,remains prohibitively expensive.Genotype phasing and imputation technologies provide a viable solution by elevating low-depth sequencing data to high-depth levels with remarkable accuracy.Sebastes schlegelii,a marine species with great economic value,holds significant potential for in-depth genetic research and breeding programs.This study sampled four S.schlegelii populations—Rongcheng-cultured,Rongcheng-wild,Wendeng-cultured,and Yantai-cultured-sampled from the coast of the Shandong Province,China.High-depth genome resequencing in 160 individuals identified>7000000 high-quality SNPs.We assessed the hidden Markov chain model-based tools,phasing accuracies of SHAPEIT2,Eagle2,and Beagle5,and the imputation performances of Beagle5 and GLIMPSE2 across varying sequencing depths and chromosomes.The results indicated comparable phasing accuracies among these tools,with SHAPEIT2 being slightly superior,while GLIMPSE2 consistently outperformed Beagle5 in imputation accuracy.Consequently,an optimal strategy combining SHAPEIT2 for phasing and GLIMPSE2 for imputation was established,achieving an accuracy of about 93%at 1×sequencing depth.Validation using imputed data from an additional 80 individuals sequenced at 1×depth,together with the originally-selected 160 individuals,was down-sampled to 1×using the seqtk tool and imputed.The data were consistent with high-depth(10×)results obtained during population structure analysis.These affirmed the reliability of this cost-effective approach for SNP-based studies in S.schlegelii.The methods presented in this research markedly reduced the costs associated with genomic selection in S.schlegelii,paving the way for accelerating breeding programs to enhance economically valuable traits in this species. 展开更多
关键词 Sebastes schlegelii genotype phasing and imputation low-coverage sequencing
暂未订购 下载PDF
A Composite Loss-Based Autoencoder for Accurate and Scalable Missing Data Imputation 认领 引用
7
作者 Thierry Mugenzi Cahit Perkgoz 《Computers, Materials & Continua》 SCIE EI 2026年第1期1985-2005,共21页
Missing data presents a crucial challenge in data analysis,especially in high-dimensional datasets,where missing data often leads to biased conclusions and degraded model performance.In this study,we present a novel a... Missing data presents a crucial challenge in data analysis,especially in high-dimensional datasets,where missing data often leads to biased conclusions and degraded model performance.In this study,we present a novel autoencoder-based imputation framework that integrates a composite loss function to enhance robustness and precision.The proposed loss combines(i)a guided,masked mean squared error focusing on missing entries;(ii)a noise-aware regularization term to improve resilience against data corruption;and(iii)a variance penalty to encourage expressive yet stable reconstructions.We evaluate the proposed model across four missingness mechanisms,such as Missing Completely at Random,Missing at Random,Missing Not at Random,and Missing Not at Random with quantile censorship,under systematically varied feature counts,sample sizes,and missingness ratios ranging from 5%to 60%.Four publicly available real-world datasets(Stroke Prediction,Pima Indians Diabetes,Cardiovascular Disease,and Framingham Heart Study)were used,and the obtained results show that our proposed model consistently outperforms baseline methods,including traditional and deep learning-based techniques.An ablation study reveals the additive value of each component in the loss function.Additionally,we assessed the downstream utility of imputed data through classification tasks,where datasets imputed by the proposed method yielded the highest receiver operating characteristic area under the curve scores across all scenarios.The model demonstrates strong scalability and robustness,improving performance with larger datasets and higher feature counts.These results underscore the capacity of the proposed method to produce not only numerically accurate but also semantically useful imputations,making it a promising solution for robust data recovery in clinical applications. 展开更多
关键词 Missing data imputation autoencoder deep learning missing mechanisms
暂未订购 下载PDF
Development of a machine learning-based model for predicting postoperative survival in gastric cancer 认领 引用 被引量:1
8
作者 Ya-Na Lü Dong Liu +3 位作者 Shuai Tao Ju Wu Shu-Juan Yu Hui-Ling Yuan 《World Journal of Gastrointestinal Surgery》 SCIE 2026年第2期265-278,共14页
BACKGROUND Accurate prediction of postoperative survival is crucial for the personalized management of gastric cancer.However,the development of robust predictive models is often constrained by incomplete clinical dat... BACKGROUND Accurate prediction of postoperative survival is crucial for the personalized management of gastric cancer.However,the development of robust predictive models is often constrained by incomplete clinical data,while their clinical utility is limited by poor interpretability and the absence of practical applications.AIM To develop an interpretable machine learning model for predicting 3-year survival following gastric cancer surgery.A novel data imputation method was proposed to handle missing values,and a user-friendly online tool was developed to facilitate clinical decision-making.METHODS A retrospective analysis was conducted on a group of 304 patients with gastric adenocarcinoma.A hybrid imputation method(HDI-MF-Gower)was developed and compared against conventional techniques.Key prognostic factors were identified by integrating least absolute shrinkage and selection operator regression with the Boruta algorithm.Subsequently,ten machine learning models were trained and validated.RESULTS The proposed HDI-MF-Gower method demonstrated superior imputation accuracy.Seven features were selected for the final model.The extra trees classifier achieved the best performance on the independent validation set,with an area under the curve of 0.853 and an accuracy of 0.772.The optimal model was interpreted using SHapley Additive exPlanations analysis and deployed as an online prediction tool.CONCLUSION A robust and interpretable predictive model integrating advanced data imputation was successfully developed.The deployed tool facilitates individualized prognostic assessment and shows potential for enhancing personalized treatment planning in gastric cancer. 展开更多
关键词 Gastric cancer Machine learning Survival prediction Missing data imputation Extra trees
暂未订购 下载PDF
Systematic assessment of mixed imputation methods and explainable machine learning 认领 引用
9
作者 Jia-Yi Li Yan Zhao 《World Journal of Gastrointestinal Surgery》 SCIE 2026年第7期14-23,共10页
The convergence of artificial intelligence and precision oncology is frequently hampered by the quality of real-world clinical data,particularly the pervasive challenge of missing values.This opinion review critically... The convergence of artificial intelligence and precision oncology is frequently hampered by the quality of real-world clinical data,particularly the pervasive challenge of missing values.This opinion review critically appraises the methodology and evidentiary framework of the study,which proposes a hybrid imputation architecture,HDI-MF-Gower,integrated with an extra trees classifier and Shaply Additive exPlanation interpretability for predicting survival outcomes following curative gastrectomy.We deconstruct the pivotal assumptions and potential sensitivities of their adaptive weighted similarity initialization.This design is engineered to provide a“warm start”aligned with the underlying data structure for iterative imputation,theoretically mitigating the risks of distributional distortion associated with simplistic initialization strategies.However,a primary boundary of the current evidence lies in the validation hierarchy;the reported validation relies predominantly on random splitting within a singlecenter cohort,lacking the robustness of temporal extrapolation or genuine external validation.Furthermore,statistical comparisons suggest that the performance differences between the proposed model and several robust ensemble baselines are not consistently distinguishable,making it difficult to attribute performance gains solely to the specific choice of the learner.We conclude that future research must construct a more rigorous evidence chain within multicenter and multimodal frameworks.Crucially,adherence to transparent reporting of a multivariable prediction model for individual prognosis or diagnosis+artificial intelligence guidelines-specifically regarding missing data mechanisms,sensitivity analyses,calibration and net benefit assessments,and the availability of reproducible materials-is essential to substantiate generalizable clinical utility. 展开更多
关键词 Gastric cancer Machine learning Survival prediction Missing data imputation Extra trees
暂未订购 下载PDF
Imputing single-cell RNA-seq data by considering cell heterogeneity and prior expression of dropouts 认领 引用 被引量:2
10
作者 Lihua Zhang Shihua Zhang 《Journal of Molecular Cell Biology》 SCIE CAS CSCD 2021年第1期29-40,共12页
Single-cell RNA sequencing(scRNA-seq)provides a powerful tool to determine expression patterns of thousands of individual cells.However,the analysis of scRNA-seq data remains a computational challenge due to the high ... Single-cell RNA sequencing(scRNA-seq)provides a powerful tool to determine expression patterns of thousands of individual cells.However,the analysis of scRNA-seq data remains a computational challenge due to the high technical noise such as the presence of dropout events that lead to a large proportion of zeros for expressed genes.Taking into account the cell heterogeneity and the relationship between dropout rate and expected expression level,we present a cell sub-population based bounded low-rank(PBLR)method to impute the dropouts of scRNA-seq data.Through application to both simulated and real scRNA-seq datasets,PBLR is shown to be effective in recovering dropout events,and it can dramaimprove the low・dimensional representation and the recovery of gene-gene relationships masked by dropout events compared to several state-of-the-art methods・Moreover,PBLR also detects accurate and robust cell sub-populations automatically,shedding light on its flexibility and generality for scRNA-seq data analysis. 展开更多
关键词 single-cell RNA-seq dropout imputation low订ank systems biology
暂未订购 下载PDF
Imputing DNA Methylation by Transferred Learning Based Neural Network 认领 引用
11
作者 Xin-Feng Wang Xiang Zhou +2 位作者 Jia-Hua Rao Zhu-Jin Zhang Yue-Dong Yang 《Journal of Computer Science & Technology》 SCIE EI CSCD 2022年第2期320-329,共10页
DNA methylation is one important epigenetic type to play a vital role in many diseases including cancers.With the development of the high-throughput sequencing technology,there is much progress to disclose the relatio... DNA methylation is one important epigenetic type to play a vital role in many diseases including cancers.With the development of the high-throughput sequencing technology,there is much progress to disclose the relations of DNA methylation with diseases.However,the analyses of DNA methylation data are challenging due to the missing values caused by the limitations of current techniques.While many methods have been developed to impute the missing values,these methods are mostly based on the correlations between individual samples,and thus are limited for the abnormal samples in cancers.In this study,we present a novel transfer learning based neural network to impute missing DNA methylation data,namely the TDimpute-DNAmeth method.The method learns common relations between DNA methylation from pan-cancer samples,and then fine-tunes the learned relations over each specific cancer type for imputing the missing data.Tested on 16 cancer datasets,our method was shown to outperform other commonly-used methods.Further analyses indicated that DNA methylation is related to cancer survival and thus can be used as a biomarker of cancer prognosis. 展开更多
关键词 neural network transfer learning DNA methylation data imputation survival analysis
暂未订购 下载PDF
GNTI:Gaussian noise-based trajectory imputation via self-supervised learning 认领 引用
12
作者 Siqi Liu Penghao Zhao +3 位作者 Lei Dong Na Dong Liyuan Ding Guangshi Pei 《Journal of Highway and Transportation Research and Development(English Edition)》 2026年第1期28-34,共7页
Trajectory imputation aims to reconstruct complete movement sequences from noisy or incomplete GPS data,crucial for intelligent transportation systems(ITS).This study proposes GNTI(Gaussian Noise-based Trajectory Impu... Trajectory imputation aims to reconstruct complete movement sequences from noisy or incomplete GPS data,crucial for intelligent transportation systems(ITS).This study proposes GNTI(Gaussian Noise-based Trajectory Imputation),a self-supervised learning framework that introduces Gaussian noise during model training to simulate real-world GPS errors.By perturbing the input trajectories with probabilistic Gaussian noise,GNTI enables the model to learn robust trajectory representations without relying on labeled datasets.A transformer-based BERT encoder is employed to capture complex spatial-temporal dependencies,while a simple multilayer perceptron(MLP)decoder predicts corrected trajectory points based on contextualized embeddings.Extensive experiments were conducted on two large-scale real-world datasets,Chengdu and Porto.Comparative results show that GNTI outperforms traditional Seq2Seq-based models(gated recurrent unit(GRU),long short-term memory(LSTM))and recent transformer-based models(Transformer,ST-BerImp),achieving the highest Micro-F1 scores across all settings.Specifically,GNTI improves Micro-F1 scores by 3%–5%over ST-BerImp.Ablation studies demonstrate that Gaussian noise augmentation improves model robustness by approximately 5%compared to models trained without augmentation.GNTI offers a practical and scalable solution for trajectory imputation tasks,enhancing robustness to GPS inaccuracies and reducing the need for complex multi-task objectives.Future work may explore extending the method to denser urban environments and optimizing it for real-time deployment. 展开更多
关键词 Transportation engineering ITS trajectory imputation Gaussian noise self-supervised learning GPS noise transformer
暂未订购 下载PDF
A Modified Deep Residual-Convolutional Neural Network for Accurate Imputation of Missing Data 认领 引用 被引量:1
13
作者 Firdaus Firdaus Siti Nurmaini +8 位作者 Anggun Islami Annisa Darmawahyuni Ade Iriani Sapitri Muhammad Naufal Rachmatullah Bambang Tutuko Akhiar Wista Arum Muhammad Irfan Karim Yultrien Yultrien Ramadhana Noor Salassa Wandya 《Computers, Materials & Continua》 SCIE EI 2025年第2期3419-3441,共23页
Handling missing data accurately is critical in clinical research, where data quality directly impacts decision-making and patient outcomes. While deep learning (DL) techniques for data imputation have gained attentio... Handling missing data accurately is critical in clinical research, where data quality directly impacts decision-making and patient outcomes. While deep learning (DL) techniques for data imputation have gained attention, challenges remain, especially when dealing with diverse data types. In this study, we introduce a novel data imputation method based on a modified convolutional neural network, specifically, a Deep Residual-Convolutional Neural Network (DRes-CNN) architecture designed to handle missing values across various datasets. Our approach demonstrates substantial improvements over existing imputation techniques by leveraging residual connections and optimized convolutional layers to capture complex data patterns. We evaluated the model on publicly available datasets, including Medical Information Mart for Intensive Care (MIMIC-III and MIMIC-IV), which contain critical care patient data, and the Beijing Multi-Site Air Quality dataset, which measures environmental air quality. The proposed DRes-CNN method achieved a root mean square error (RMSE) of 0.00006, highlighting its high accuracy and robustness. We also compared with Low Light-Convolutional Neural Network (LL-CNN) and U-Net methods, which had RMSE values of 0.00075 and 0.00073, respectively. This represented an improvement of approximately 92% over LL-CNN and 91% over U-Net. The results showed that this DRes-CNN-based imputation method outperforms current state-of-the-art models. These results established DRes-CNN as a reliable solution for addressing missing data. 展开更多
关键词 Data imputation missing data deep learning deep residual convolutional neural network
暂未订购 下载PDF
Handling missing data in large-scale TBM datasets:Methods,strategies,and applications 认领 引用 被引量:1
14
作者 Haohan Xiao Ruilang Cao +5 位作者 Zuyu Chen Chengyu Hong Jun Wang Min Yao Litao Fan Teng Luo 《Intelligent Geoengineering》 2025年第3期109-125,共17页
Substantial advancements have been achieved in Tunnel Boring Machine(TBM)technology and monitoring systems,yet the presence of missing data impedes accurate analysis and interpretation of TBM monitoring results.This s... Substantial advancements have been achieved in Tunnel Boring Machine(TBM)technology and monitoring systems,yet the presence of missing data impedes accurate analysis and interpretation of TBM monitoring results.This study aims to investigate the issue of missing data in extensive TBM datasets.Through a comprehensive literature review,we analyze the mechanism of missing TBM data and compare different imputation methods,including statistical analysis and machine learning algorithms.We also examine the impact of various missing patterns and rates on the efficacy of these methods.Finally,we propose a dynamic interpolation strategy tailored for TBM engineering sites.The research results show that K-Nearest Neighbors(KNN)and Random Forest(RF)algorithms can achieve good interpolation results;As the missing rate increases,the interpolation effect of different methods will decrease;The interpolation effect of block missing is poor,followed by mixed missing,and the interpolation effect of sporadic missing is the best.On-site application results validate the proposed interpolation strategy's capability to achieve robust missing value interpolation effects,applicable in ML scenarios such as parameter optimization,attitude warning,and pressure prediction.These findings contribute to enhancing the efficiency of TBM missing data processing,offering more effective support for large-scale TBM monitoring datasets. 展开更多
关键词 Tunnel boring machine(TBM) Missing data imputation Machine learning(ML) Time series interpolation Data preprocessing Real-time data stream
暂未订购 下载PDF
A Diffusion Model for Traffic Data Imputation 认领 引用 被引量:1
15
作者 Bo Lu Qinghai Miao +5 位作者 Yahui Liu Tariku Sinshaw Tamir Hongxia Zhao Xiqiao Zhang Yisheng Lv Fei-Yue Wang 《IEEE/CAA Journal of Automatica Sinica》 SCIE EI CSCD 2025年第3期606-617,共12页
Imputation of missing data has long been an important topic and an essential application for intelligent transportation systems(ITS)in the real world.As a state-of-the-art generative model,the diffusion model has prov... Imputation of missing data has long been an important topic and an essential application for intelligent transportation systems(ITS)in the real world.As a state-of-the-art generative model,the diffusion model has proven highly successful in image generation,speech generation,time series modelling etc.and now opens a new avenue for traffic data imputation.In this paper,we propose a conditional diffusion model,called the implicit-explicit diffusion model,for traffic data imputation.This model exploits both the implicit and explicit feature of the data simultaneously.More specifically,we design two types of feature extraction modules,one to capture the implicit dependencies hidden in the raw data at multiple time scales and the other to obtain the long-term temporal dependencies of the time series.This approach not only inherits the advantages of the diffusion model for estimating missing data,but also takes into account the multiscale correlation inherent in traffic data.To illustrate the performance of the model,extensive experiments are conducted on three real-world time series datasets using different missing rates.The experimental results demonstrate that the model improves imputation accuracy and generalization capability. 展开更多
关键词 Data imputation diffusion model implicit feature time series traffic data
暂未订购 下载PDF
Prediction of radionuclide diffusion enabled by missing data imputation and ensemble machine learning 认领 引用 被引量:1
16
作者 Jun-Lei Tian Jia-Xing Feng +4 位作者 Jia-Cong Shen Lei Yao Jing-Yan Wang Tao Wu Yao-Lin Zhao 《Nuclear Science and Techniques》 SCIE EI CAS CSCD 2025年第10期47-61,共15页
Missing values in radionuclide diffusion datasets can undermine the predictive accuracy and robustness of the machine learning(ML)models.In this study,regression-based missing data imputation method using a light grad... Missing values in radionuclide diffusion datasets can undermine the predictive accuracy and robustness of the machine learning(ML)models.In this study,regression-based missing data imputation method using a light gradient boosting machine(LGBM)algorithm was employed to impute more than 60%of the missing data,establishing a radionuclide diffusion dataset containing 16 input features and 813 instances.The effective diffusion coefficient(De)was predicted using ten ML models.The predictive accuracy of the ensemble meta-models,namely LGBM-extreme gradient boosting(XGB)and LGBM-categorical boosting(CatB),surpassed that of the other ML models,with R2values of 0.94.The models were applied to predict the Devalues of EuEDTAand HCrO4in saturated compacted bentonites at compactions ranging from 1200 to 1800 kg/m3,which were measured using a through-diffusion method.The generalization ability of the LGBM-XGB model surpassed that of LGB-CatB in predicting the Deof HCrO4.Shapley additive explanations identified total porosity as the most significant influencing factor.Additionally,the partial dependence plot analysis technique yielded clearer results in the univariate correlation analysis.This study provides a regression imputation technique to refine radionuclide diffusion datasets,offering deeper insights into analyzing the diffusion mechanism of radionuclides and supporting the safety assessment of the geological disposal of high-level radioactive waste. 展开更多
关键词 Machine learning Radionuclide diffusion Bentonite Regression imputation Missing data Diffusion experiments
暂未订购 下载PDF
Longevity prediction and missing data treatment of landslide dams 认领 引用
17
作者 WANG Danyan YANG Xingguo +2 位作者 ZHOU Jiawen FENG Zhenyu LIAO Haimei 《Journal of Mountain Science》 SCIE CSCD 2025年第7期2640-2653,共14页
Landslide dam failures can cause significant damage to both society and ecosystems.Predicting the failure of these dams in advance enables early preventive measures,thereby minimizing potential harm.This paper aims to... Landslide dam failures can cause significant damage to both society and ecosystems.Predicting the failure of these dams in advance enables early preventive measures,thereby minimizing potential harm.This paper aims to propose a fast and accurate model for predicting the longevity of landslide dams while also addressing the issue of missing data.Given the wide variation in the survival times of landslide dams—from mere minutes to several thousand years—predicting their longevity presents a considerable challenge.The study develops predictive models by considering key factors such as dam geometry,hydrodynamic conditions,materials,and triggering parameters.A dataset of 1045 landslide dam cases is analyzed,categorizing their longevity into three distinct groups:C1(1 year).Multiple imputation and knearest neighbor algorithms are used to handle missing data on geometric size,hydrodynamic conditions,materials,and triggers.Based on the imputed data,two predictive models are developed:a classification model for dam longevity categories and a regression model for precise longevity predictions.The classification model achieves an accuracy of 88.38%while the regression model outperforms existing models with an R2 value of 0.966.Two real-life landslide dam cases are used to validate the models,which show correct classification and small prediction errors.The longevity of landslide dams is jointly influenced by factors such as geometric size,hydrodynamic conditions,materials,and triggering events.Among these,geometric size has the greatest impact,followed by hydrodynamic conditions,materials,and triggers,as confirmed by variable importance in the model development. 展开更多
关键词 Category Longevity range Imputation Prediction models Decision Tree
暂未订购 下载PDF
Effective and efficient handling of missing data in supervised machine learning 认领 引用
18
作者 Peter Ayokunle Popoola Jules-Raymond Tapamo Alain Guy Honoré Assounga 《Data Science and Management》 EI CSCD 2025年第3期361-373,共13页
The prevailing consensus in statistical literature is that multiple imputation is generally the most suitable method for addressing missing data in statistical analyses,whereas a complete case analysis is deemed appro... The prevailing consensus in statistical literature is that multiple imputation is generally the most suitable method for addressing missing data in statistical analyses,whereas a complete case analysis is deemed appropriate only when the rate of missingness is negligible or when the missingness mechanism is missing completely at random(MCAR).This study investigates the applicability of this consensus within the context of supervised machine learning,with particular emphasis on the interactions between the imputation method,missingness mechanism,and missingness rate.Furthermore,we examine the time efficiency of these“state-of-the-art”imputation methods considering the time-sensitive nature of certain machine learning applications.Utilizing ten real-world datasets,we introduced missingness at rates ranging from approximately 5%–75%under the MCAR,missing at random(MAR),and missing not at random(MNAR)mechanisms.We subsequently address missing data using five methods:complete case analysis(CCA),mean imputation,hot deck imputation,regression imputation,and multiple imputation(MI).Statistical tests are conducted on the machine learning outcomes,and the findings are presented and analyzed.Our investigation reveals that in nearly all scenarios,CCA performs comparably to MI,even with substantial levels of missingness under the MAR and MNAR conditions and with missingness in the output variable for regression problems.Under some conditions,CCA surpasses MI in terms of its performance.Thus,given the considerable computational demands associated with MI,the application of CCA is recommended within the broader context of supervised machine learning,particularly in big-data environments. 展开更多
关键词 Classification Imputation Learning Missing data Prediction
暂未订购 下载PDF
A Novel Reduced Error Pruning Tree Forest with Time-Based Missing Data Imputation(REPTF-TMDI)for Traffic Flow Prediction 认领 引用
19
作者 Yunus Dogan Goksu Tuysuzoglu +4 位作者 Elife Ozturk Kiyak Bita Ghasemkhani Kokten Ulas Birant Semih Utku Derya Birant 《Computer Modeling in Engineering & Sciences》 SCIE EI 2025年第8期1677-1715,共39页
Accurate traffic flow prediction(TFP)is vital for efficient and sustainable transportation management and the development of intelligent traffic systems.However,missing data in real-world traffic datasets poses a sign... Accurate traffic flow prediction(TFP)is vital for efficient and sustainable transportation management and the development of intelligent traffic systems.However,missing data in real-world traffic datasets poses a significant challenge to maintaining prediction precision.This study introduces REPTF-TMDI,a novel method that combines a Reduced Error Pruning Tree Forest(REPTree Forest)with a newly proposed Time-based Missing Data Imputation(TMDI)approach.The REP Tree Forest,an ensemble learning approach,is tailored for time-related traffic data to enhance predictive accuracy and support the evolution of sustainable urbanmobility solutions.Meanwhile,the TMDI approach exploits temporal patterns to estimate missing values reliably whenever empty fields are encountered.The proposed method was evaluated using hourly traffic flow data from a major U.S.roadway spanning 2012-2018,incorporating temporal features(e.g.,hour,day,month,year,weekday),holiday indicator,and weather conditions(temperature,rain,snow,and cloud coverage).Experimental results demonstrated that the REPTF-TMDI method outperformed conventional imputation techniques across various missing data ratios by achieving an average 11.76%improvement in terms of correlation coefficient(R).Furthermore,REPTree Forest achieved improvements of 68.62%in RMSE and 70.52%in MAE compared to existing state-of-the-art models.These findings highlight the method’s ability to significantly boost traffic flow prediction accuracy,even in the presence of missing data,thereby contributing to the broader objectives of sustainable urban transportation systems. 展开更多
关键词 Machine learning traffic flow prediction missing data imputation reduced error pruning tree(REPTree) sustainable transportation systems traffic management artificial intelligence
暂未订购 下载PDF
An Integrated Perception Model for Predicting and Analyzing Urban Rail Transit Emergencies Based on Unstructured Data 认领 引用
20
作者 Liang Mu Yurui Kang +1 位作者 Zixu Yan Guangyu Zhu 《Computers, Materials & Continua》 SCIE EI 2025年第8期2495-2512,共18页
The accurate prediction and analysis of emergencies in Urban Rail Transit Systems(URTS)are essential for the development of effective early warning and prevention mechanisms.This study presents an integrated perceptio... The accurate prediction and analysis of emergencies in Urban Rail Transit Systems(URTS)are essential for the development of effective early warning and prevention mechanisms.This study presents an integrated perception model designed to predict emergencies and analyze their causes based on historical unstructured emergency data.To address issues related to data structuredness and missing values,we employed label encoding and an Elastic Net Regularization-based Generative Adversarial Interpolation Network(ER-GAIN)for data structuring and imputation.Additionally,to mitigate the impact of imbalanced data on the predictive performance of emergencies,we introduced an Adaptive Boosting Ensemble Model(AdaBoost)to forecast the key features of emergencies,including event types and levels.We also utilized Information Gain(IG)to analyze and rank the causes of various significant emergencies.Experimental results indicate that,compared to baseline data imputation models,ER-GAIN improved the prediction accuracy of key emergency features by 3.67%and 3.78%,respectively.Furthermore,AdaBoost enhanced the accuracy by over 4.34%and 3.25%compared to baseline predictivemodels.Through causation analysis,we identified the critical causes of train operation and fire incidents.The findings of this research will contribute to the establishment of early warning and prevention mechanisms for emergencies in URTS,potentially leading to safer and more reliable URTS operations. 展开更多
关键词 Urban rail transit system emergency prediction generative adversarial imputation network ensemble learning cause analysis
暂未订购 下载PDF
上一页 1 2 8 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈