Handling high-dimensional data with missing values by modern machine learning techniques

被引:4
|
作者
Chen, Sixia [1 ]
Xu, Chao [1 ]
机构
[1] Univ Oklahoma, Dept Biostat & Epidemiol, Hlth Sci Ctr, Oklahoma City, OK 73126 USA
基金
美国国家卫生研究院;
关键词
Deep learning; high-dimensional data; imputation; machine learning; missing data; JACKKNIFE VARIANCE-ESTIMATION; MULTIPLE IMPUTATION; FRACTIONAL IMPUTATION; ITEM NONRESPONSE; INFERENCE; VARIABLES; SELECTION;
D O I
10.1080/02664763.2022.2068514
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
High-dimensional data have been regarded as one of the most important types of big data in practice. It happens frequently in practice including genetic study, financial study, and geographical study. Missing data in high dimensional data analysis should be handled properly to reduce nonresponse bias. We discuss some modern machine learning techniques including penalized regression approaches, tree-based approaches, and deep learning (DL) for handling missing data with high dimensionality. Specifically, our proposed methods can be used for estimating general parameters of interest including population means and percentiles with imputation-based estimators, propensity score estimators, and doubly robust estimators. We compare those methods through some limited simulation studies and a real application. Both simulation studies and real application show the benefits of DL and XGboost approaches compared with other methods in terms of balancing bias and variance.
引用
收藏
页码:786 / 804
页数:19
相关论文
共 50 条
  • [41] Handling missing values in support vector machine classifiers
    Pelckmans, K
    De Brabanter, J
    Suykens, JAK
    De Moor, B
    NEURAL NETWORKS, 2005, 18 (5-6) : 684 - 692
  • [42] An overview of modern machine learning methods for effect measure modification analyses in high-dimensional settings
    Cheung, Michael
    Dimitrova, Anna
    Benmarhnia, Tarik
    SSM-POPULATION HEALTH, 2025, 29
  • [43] Efficient Learning on High-dimensional Operational Data
    Samani, Forough Shahab
    Zhang, Hongyi
    Stadler, Rolf
    2019 15TH INTERNATIONAL CONFERENCE ON NETWORK AND SERVICE MANAGEMENT (CNSM), 2019,
  • [44] Machine learning for high-dimensional dynamic stochastic economies
    Scheidegger, Simon
    Bilionis, Ilias
    JOURNAL OF COMPUTATIONAL SCIENCE, 2019, 33 : 68 - 82
  • [45] PCA learning for sparse high-dimensional data
    Hoyle, DC
    Rattray, M
    EUROPHYSICS LETTERS, 2003, 62 (01): : 117 - 123
  • [46] Metric Learning for High-Dimensional Tensor Data
    Shi Jiarong
    Jiao Licheng
    Shang Fanhua
    CHINESE JOURNAL OF ELECTRONICS, 2011, 20 (03): : 495 - 498
  • [47] Similarity Learning for High-Dimensional Sparse Data
    Liu, Kuan
    Bellet, Aurelien
    Sha, Fei
    ARTIFICIAL INTELLIGENCE AND STATISTICS, VOL 38, 2015, 38 : 653 - 662
  • [48] Group Learning for High-Dimensional Sparse Data
    Cherkassky, Vladimir
    Chen, Hsiang-Han
    Shiao, Han-Tai
    2019 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2019,
  • [49] Water-Quality Data Imputation with a High Percentage of Missing Values: A Machine Learning Approach
    Rodriguez, Rafael
    Pastorini, Marcos
    Etcheverry, Lorena
    Chreties, Christian
    Fossati, Monica
    Castro, Alberto
    Gorgoglione, Angela
    SUSTAINABILITY, 2021, 13 (11)
  • [50] A Performance Analysis of Prediction Techniques in Handling High-Dimensional Uncertain Data for the Application of Skyline Query Over Data Stream
    Mohamud, Mudathir Ahmed
    Ibrahim, Hamidah
    Sidi, Fatimah
    Rum, Siti Nurulain Mohd
    Dzolkhifli, Zarina Binti
    Xiaowei, Zhang
    IEEE ACCESS, 2024, 12 : 120877 - 120898