Handling high-dimensional data with missing values by modern machine learning techniques

被引:4
|
作者
Chen, Sixia [1 ]
Xu, Chao [1 ]
机构
[1] Univ Oklahoma, Dept Biostat & Epidemiol, Hlth Sci Ctr, Oklahoma City, OK 73126 USA
基金
美国国家卫生研究院;
关键词
Deep learning; high-dimensional data; imputation; machine learning; missing data; JACKKNIFE VARIANCE-ESTIMATION; MULTIPLE IMPUTATION; FRACTIONAL IMPUTATION; ITEM NONRESPONSE; INFERENCE; VARIABLES; SELECTION;
D O I
10.1080/02664763.2022.2068514
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
High-dimensional data have been regarded as one of the most important types of big data in practice. It happens frequently in practice including genetic study, financial study, and geographical study. Missing data in high dimensional data analysis should be handled properly to reduce nonresponse bias. We discuss some modern machine learning techniques including penalized regression approaches, tree-based approaches, and deep learning (DL) for handling missing data with high dimensionality. Specifically, our proposed methods can be used for estimating general parameters of interest including population means and percentiles with imputation-based estimators, propensity score estimators, and doubly robust estimators. We compare those methods through some limited simulation studies and a real application. Both simulation studies and real application show the benefits of DL and XGboost approaches compared with other methods in terms of balancing bias and variance.
引用
收藏
页码:786 / 804
页数:19
相关论文
共 50 条
  • [1] Missing Data Imputation with High-Dimensional Data
    Brini, Alberto
    van den Heuvel, Edwin R.
    AMERICAN STATISTICIAN, 2024, 78 (02) : 240 - 252
  • [2] An ensemble learning method for variable selection: application to high-dimensional data and missing values
    Bar-Hen, Avner
    Audigier, Vincent
    JOURNAL OF STATISTICAL COMPUTATION AND SIMULATION, 2022, 92 (16) : 3488 - 3510
  • [3] Improving Penalized Logistic Regression Model with Missing Values in High-Dimensional Data
    Alharthi, Aiedh Mrisi
    Lee, Muhammad Hisyam
    Algamal, Zakariya Yahya
    INTERNATIONAL JOURNAL OF ONLINE AND BIOMEDICAL ENGINEERING, 2022, 18 (02) : 40 - 54
  • [4] A Deep Learning-Cuckoo Search Method for Missing Data Estimation in High-Dimensional Datasets
    Leke, Collins
    Ndjiongue, Alain Richard
    Twala, Bhekisipho
    Marwala, Tshilidzi
    ADVANCES IN SWARM INTELLIGENCE, ICSI 2017, PT I, 2017, 10385 : 561 - 572
  • [5] Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques
    Liu, Mingxuan
    Li, Siqi
    Yuan, Han
    Ong, Marcus Eng Hock
    Ning, Yilin
    Xie, Feng
    Saffari, Seyed Ehsan
    Shang, Yuqing
    Volovici, Victor
    Chakraborty, Bibhas
    Liu, Nan
    ARTIFICIAL INTELLIGENCE IN MEDICINE, 2023, 142
  • [6] Missing values handling for machine learning portfolios
    Chen, Andrew Y.
    McCoy, Jack
    JOURNAL OF FINANCIAL ECONOMICS, 2024, 155
  • [7] Flexible High-Dimensional Unsupervised Learning with Missing Data
    Wei, Yuhong
    Tang, Yang
    McNicholas, Paul D.
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2020, 42 (03) : 610 - 621
  • [8] Dimension reduction of high-dimensional dataset with missing values
    Zhang, Ran
    Ye, Bin
    Liu, Peng
    JOURNAL OF ALGORITHMS & COMPUTATIONAL TECHNOLOGY, 2019, 13
  • [9] Prediction of vancomycin dose on high-dimensional data using machine learning techniques
    Huang, Xiaohui
    Yu, Ze
    Wei, Xin
    Shi, Junfeng
    Wang, Yu
    Wang, Zeyuan
    Chen, Jihui
    Bu, Shuhong
    Li, Lixia
    Gao, Fei
    Zhang, Jian
    Xu, Ajing
    EXPERT REVIEW OF CLINICAL PHARMACOLOGY, 2021, 14 (06) : 761 - 771
  • [10] Variable selection techniques after multiple imputation in high-dimensional data
    Zahid, Faisal Maqbool
    Faisal, Shahla
    Heumann, Christian
    STATISTICAL METHODS AND APPLICATIONS, 2020, 29 (03) : 553 - 580