Multilevel Stochastic Optimization for Imputation in Massive Medical Data Records

被引:1
|
作者
Li, Wenrui [1 ]
Wang, Xiaoyu [1 ]
Sun, Yuetian [1 ]
Milanovic, Snezana [1 ,2 ]
Kon, Mark [1 ]
Castrillon-Candas, Julio Enrique [1 ]
机构
[1] Boston Univ, Dept Math & Stat, Boston, MA 02215 USA
[2] Sunov Pharmaceut, Marlborough, MA 01752 USA
基金
美国国家科学基金会;
关键词
Covariance matrices; Optimization; Stochastic processes; Deep learning; Iterative methods; Costs; Big Data; Best linear unbiased predictor; computational applied mathematics; machine learning; massive datasets; numerical stability; APPROXIMATION; EQUATIONS; PDES;
D O I
10.1109/TBDATA.2023.3328433
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
It has long been a recognized problem that many datasets contain significant levels of missing numerical data. A potentially critical predicate for application of machine learning methods to datasets involves addressing this problem. However, this is a challenging task. In this article, we apply a recently developed multi-level stochastic optimization approach to the problem of imputation in massive medical records. The approach is based on computational applied mathematics techniques and is highly accurate. In particular, for the Best Linear Unbiased Predictor (BLUP) this multi-level formulation is exact, and is significantly faster and more numerically stable. This permits practical application of Kriging methods to data imputation problems for massive datasets. We test this approach on data from the National Inpatient Sample (NIS) data records, Healthcare Cost and Utilization Project (HCUP), Agency for Healthcare Research and Quality. Numerical results show that the multi-level method significantly outperforms current approaches and is numerically robust. It has superior accuracy as compared with methods recommended in the recent report from HCUP. Benchmark tests show up to 75% reductions in error. Furthermore, the results are also superior to recent state of the art methods such as discriminative deep learning.
引用
收藏
页码:122 / 131
页数:10
相关论文
共 50 条
  • [31] A fully adaptive multilevel stochastic collocation strategy for solving elliptic PDEs with random data
    Lang, J.
    Scheichl, R.
    Silvester, D.
    JOURNAL OF COMPUTATIONAL PHYSICS, 2020, 419
  • [32] Implementation of a Big Data Accessing and Processing Platform for Medical Records in Cloud
    Chao-Tung Yang
    Jung-Chun Liu
    Shuo-Tsung Chen
    Hsin-Wen Lu
    Journal of Medical Systems, 2017, 41
  • [33] An Adaptive Multilevel Monte Carlo Method with Stochastic Bounds for Quantities of Interest with Uncertain Data
    Eigel, Martin
    Merdon, Christian
    Neumann, Johannes
    SIAM-ASA JOURNAL ON UNCERTAINTY QUANTIFICATION, 2016, 4 (01): : 1219 - 1245
  • [34] Co-Optimization of VaR and CVaR for Data-Driven Stochastic Demand Response Auction
    Roveto, Matt
    Mieth, Robert
    Dvorkin, Yury
    IEEE CONTROL SYSTEMS LETTERS, 2020, 4 (04): : 940 - 945
  • [35] Monitoring Model Based on Data-Driven Optimization Stochastic Configuration Network and Its Applications
    Yan, Aijun
    Hu, Kaicheng
    Wang, Dianhui
    IEEE SENSORS JOURNAL, 2025, 25 (06) : 10087 - 10096
  • [36] Online Stochastic Optimization for Unknown Linear Systems: Data-Driven Controller Synthesis and Analysis
    Bianchin, Gianluca
    Vaquero, Miguel
    Cortes, Jorge
    Dall'Anese, Emiliano
    IEEE TRANSACTIONS ON AUTOMATIC CONTROL, 2024, 69 (07) : 4411 - 4426
  • [37] Editorial: Multilevel medical security systems and big data in healthcare: trends and developments
    Fan, Fei
    Wang, Siqin
    FRONTIERS IN PUBLIC HEALTH, 2025, 12
  • [38] Data-Driven-Based Stochastic Robust Optimization for a Virtual Power Plant With Multiple Uncertainties
    Fang, Fang
    Yu, Songyuan
    Xin, Xiuli
    IEEE TRANSACTIONS ON POWER SYSTEMS, 2022, 37 (01) : 456 - 466
  • [39] ON PRECISION OF STOCHASTIC OPTIMIZATION BASED ON ESTIMATES FROM CENSORED DATA
    Volf, Petr
    KYBERNETIKA, 2014, 50 (03) : 297 - 309
  • [40] Dual strategy based missing completely at random type missing data imputation on the internet of medical things
    Punitha, P. Iris
    Sathiaseelan, J. G. R.
    INTERNATIONAL JOURNAL OF INTELLIGENT ENGINEERING INFORMATICS, 2023, 11 (04) : 317 - 336