Regression analysis after bipartite Bayesian record linkage
In many settings, a data curator or data analyst links records from two files to produce an integrated dataset. These linked data are then used to estimate regression models of interest. This two-stage approach does not necessarily account for the uncertainty in the model parameters that results from uncertainty in the linkages. Further, it does not leverage the relationships among the non-linking variables in the two files to help identify incorrect linkages. A multiple imputation framework is proposed to address these shortcomings. First, a bipartite Bayesian record linkage model is used to generate multiple plausible linked datasets. This model does not use the non-linking variables. Second, each linked file is presumed to comprise a mixture of true links and false links. The mixture model is estimated using an EM algorithm that leverages the information provided by the non-linking variables. Point and variance estimates of the regression parameters in each plausible linked file are combined via multiple imputation inferences. Using simulation studies of linear regressions, it is demonstrated that the mixture modeling approach can have desirable repeated sampling properties. The mixture modeling approach is illustrated using Bayesian record linkage of data from the Survey of Household Income and Wealth, examining a regression involving the persistence of income.
Duke Scholars
Altmetric Attention Stats
Dimensions Citation Stats
Published In
DOI
ISSN
Publication Date
Volume
Related Subject Headings
- Statistics & Probability
- 4905 Statistics
- 3802 Econometrics
Citation
Published In
DOI
ISSN
Publication Date
Volume
Related Subject Headings
- Statistics & Probability
- 4905 Statistics
- 3802 Econometrics