Categorical linkage-data analysis

被引:0
作者
Zhang, Li-Chun [1 ]
Tuoto, Tiziana [2 ]
机构
[1] Univ Southampton, Dept Social Stat & Demog, Southampton, England
[2] Ist Nazl Stat, Dept Social Stat & Demog, Rome, Italy
关键词
analysis of contingency table; heterogeneous linkage error; incomplete match space; linkage data structure; logistic regression; secondary analysis; REGRESSION;
D O I
10.1002/sim.10134
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Analysis of integrated data often requires record linkage in order to join together the data residing in separate sources. In case linkage errors cannot be avoided, due to the lack a unique identity key that can be used to link the records unequivocally, standard statistical techniques may produce misleading inference if the linked data are treated as if they were true observations. In this paper, we propose methods for categorical data analysis based on linked data that are not prepared by the analyst, such that neither the match-key variables nor the unlinked records are available. The adjustment is based on the proportion of false links in the linked file and our approach allows the probabilities of correct linkage to vary across the records without requiring that one is able to estimate this probability for each individual record. It accommodates also the general situation where unmatched records that cannot possibly be correctly linked exist in all the sources. The proposed methods are studied by simulation and applied to real data.
引用
收藏
页码:3463 / 3483
页数:21
相关论文
共 22 条