Imputation of Missing Data in Industrial Databases

被引:0
作者
Kamakshi Lakshminarayan
Steven A. Harp
Tariq Samad
机构
[1] Honeywell Technology Center,
来源
Applied Intelligence | 1999年 / 11卷
关键词
missing data; industrial databases; multiple imputation; machine learning;
D O I
暂无
中图分类号
学科分类号
摘要
A limiting factor for the application of IDA methods in many domains is the incompleteness of data repositories. Many records have fields that are not filled in, especially, when data entry is manual. In addition, a significant fraction of the entries can be erroneous and there may be no alternative but to discard these records. But every cell in a database is not an independent datum. Statistical relationships will constrain and, often determine, missing values. Data imputation, the filling in of missing values for partially missing data, can thus be an invaluable first step in many IDA projects. New imputation methods that can handle the large-scale problems and large-scale sparsity of industrial databases are needed. To illustrate the incomplete database problem, we analyze one database with instrumentation maintenance and test records for an industrial process. Despite regulatory requirements for process data collection, this database is less than 50% complete. Next, we discuss possible solutions to the missing data problem. Several approaches to imputation are noted and classified into two categories: data-driven and model-based. We then describe two machine-learning-based approaches that we have worked with. These build upon well-known algorithms: AutoClass and C4.5. Several experiments are designed, all using the maintenance database as a common test-bed but with various data splits and algorithmic variations. Results are generally positive with up to 80% accuracies of imputation. We conclude the paper by outlining some considerations in selecting imputation methods, and by discussing applications of data imputation for intelligent data analysis.
引用
收藏
页码:259 / 275
页数:16
相关论文
共 8 条
  • [1] Rubin D.B.(1976)Inference and missing data Biometrika 63 581-592
  • [2] Dempster A.P.(1977)Maximum likelihood from incomplete data via the EM algorithm (with discussion) J. Roy. Statist. Soci. B39 1-38
  • [3] Laird N.M.(1987)The calculation of posterior distributions by data augmentation (with discussion) Journal of the American Statistical Association 82 528-550
  • [4] Rubin D.B.(1996)Multiple imputation after 18+ years Journal of the American Statistical Association 91 473-489
  • [5] Tanner M.A.(1960)A method of estimation of missing values in multivariate data suitable for use with an electronic computer J. Roy. Statist. Soci. B22 302-306
  • [6] Wong W.H.(undefined)undefined undefined undefined undefined-undefined
  • [7] Rubin D.B.(undefined)undefined undefined undefined undefined-undefined
  • [8] Buck S.F.(undefined)undefined undefined undefined undefined-undefined