Skip to main content

Unique entity estimation with application to the syrian conflict

Publication ,  Journal Article
Chen, B; Shrivastava, A; Steorts, RC
Published in: Annals of Applied Statistics
June 1, 2018

Entity resolution identifies and removes duplicate entities in large, noisy databases and has grown in both usage and new developments as a result of increased data availability. Nevertheless, entity resolution has tradeoffs regarding assumptions of the data generation process, error rates, and computational scalability that make it a difficult task for real applications. In this paper, we focus on a related problem of unique entity estimation, which is the task of estimating the unique number of entities and associated standard errors in a data set with duplicate entities. Unique entity estimation shares many fundamental challenges of entity resolution, namely, that the computational cost of all-to-all entity comparisons is intractable for large databases. To circumvent this computational barrier, we propose an efficient (near-linear time) estimation algorithm based on locality sensitive hashing. Our estimator, under realistic assumptions, is unbiased and has provably low variance compared to existing random sampling based approaches. In addition, we empirically show its superiority over the state-of-the-art estimators on three real applications. The motivation for our work is to derive an accurate estimate of the documented, identifiable deaths in the ongoing Syrian conflict. Our methodology, when applied to the Syrian data set, provides an estimate of 191,874 ± 1,772 documented, identifiable deaths, which is very close to the Human Rights Data Analysis Group (HRDAG) estimate of 191,369. Our work provides an example of challenges and efforts involved in solving a real, noisy challenging problem where modeling assumptions may not hold.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Annals of Applied Statistics

DOI

EISSN

1941-7330

ISSN

1932-6157

Publication Date

June 1, 2018

Volume

12

Issue

2

Start / End Page

1039 / 1067

Related Subject Headings

  • Statistics & Probability
  • 4905 Statistics
  • 1403 Econometrics
  • 0104 Statistics
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Chen, B., Shrivastava, A., & Steorts, R. C. (2018). Unique entity estimation with application to the syrian conflict. Annals of Applied Statistics, 12(2), 1039–1067. https://doi.org/10.1214/18-AOAS1163
Chen, B., A. Shrivastava, and R. C. Steorts. “Unique entity estimation with application to the syrian conflict.” Annals of Applied Statistics 12, no. 2 (June 1, 2018): 1039–67. https://doi.org/10.1214/18-AOAS1163.
Chen B, Shrivastava A, Steorts RC. Unique entity estimation with application to the syrian conflict. Annals of Applied Statistics. 2018 Jun 1;12(2):1039–67.
Chen, B., et al. “Unique entity estimation with application to the syrian conflict.” Annals of Applied Statistics, vol. 12, no. 2, June 2018, pp. 1039–67. Scopus, doi:10.1214/18-AOAS1163.
Chen B, Shrivastava A, Steorts RC. Unique entity estimation with application to the syrian conflict. Annals of Applied Statistics. 2018 Jun 1;12(2):1039–1067.

Published In

Annals of Applied Statistics

DOI

EISSN

1941-7330

ISSN

1932-6157

Publication Date

June 1, 2018

Volume

12

Issue

2

Start / End Page

1039 / 1067

Related Subject Headings

  • Statistics & Probability
  • 4905 Statistics
  • 1403 Econometrics
  • 0104 Statistics