Review: A gentle introduction to imputation of missing values

被引:1858
作者
Donders, A. Rogier T. [1 ]
van der Heijden, Geert J. M. G.
Stijnen, Theo
Moons, Karel G. M.
机构
[1] Univ Utrecht, Ctr Biostat, NL-3508 TC Utrecht, Netherlands
[2] Univ Utrecht, Copernicus Inst, Dept Innovat Studies, NL-3508 TC Utrecht, Netherlands
[3] Univ Utrecht, Med Ctr, Julius Ctr Hlth Sci & Primary Care, NL-3508 TC Utrecht, Netherlands
[4] Erasmus Univ, Sch Med, Dept Epidemiol & Biostat, NL-3000 DR Rotterdam, Netherlands
关键词
missing data; single imputation; multiple imputation; indicator method; bias; precision;
D O I
10.1016/j.jclinepi.2006.01.014
中图分类号
R19 [保健组织与事业(卫生事业管理)];
学科分类号
摘要
In most situations, simple techniques for handling missing data (such as complete case analysis, overall mean imputation, and the missing-indicator method) produce biased results, whereas imputation techniques yield valid results without complicating the analysis once the imputations are carried out. Imputation techniques are based on the idea that any subject in a study sample can be replaced by a new randomly chosen subject from the same source population. Imputation of missing data on a variable is replacing that missing by a value that is drawn from an estimate of the distribution of this variable. In single imputation, only one estimate is used. In multiple imputation, various estimates are used, reflecting the uncertainty in the estimation of this distribution. Under the general conditions of so-called missing at random and missing completely at random, both single and multiple imputations result in unbiased estimates of study associations. But single imputation results in too small estimated standard errors, whereas multiple imputation results in correctly estimated standard errors and confidence intervals. In this article we explain why all this is the case, and use a simple simulation study to demonstrate our explanations. We also explain and illustrate why two frequently used methods to handle missing data, i.e., overall mean imputation and the missing-indicator method, almost always result in biased estimates. (c) 2006 Elsevier Inc. All rights reserved.
引用
收藏
页码:1087 / 1091
页数:5
相关论文
共 16 条
  • [11] RUBIN DB, 1976, BIOMETRIKA, V63, P581, DOI 10.1093/biomet/63.3.581
  • [12] Rubin DonaldB., 1987, MULTIPLE IMPUTATIONS
  • [13] Schafer J.L, 1997, ANAL INCOMPLETE MULT
  • [14] Vach W., 1998, ENCY BIOSTATISTICS, P2641
  • [15] Van Buuren S, 1999, STAT MED, V18, P681, DOI 10.1002/(SICI)1097-0258(19990330)18:6<681::AID-SIM71>3.0.CO
  • [16] 2-R