Skip to main content
Journal cover image

sureLDA: A multidisease automated phenotyping method for the electronic health record.

Publication ,  Journal Article
Ahuja, Y; Zhou, D; He, Z; Sun, J; Castro, VM; Gainer, V; Murphy, SN; Hong, C; Cai, T
Published in: J Am Med Inform Assoc
August 1, 2020

OBJECTIVE: A major bottleneck hindering utilization of electronic health record data for translational research is the lack of precise phenotype labels. Chart review as well as rule-based and supervised phenotyping approaches require laborious expert input, hampering applicability to studies that require many phenotypes to be defined and labeled de novo. Though International Classification of Diseases codes are often used as surrogates for true labels in this setting, these sometimes suffer from poor specificity. We propose a fully automated topic modeling algorithm to simultaneously annotate multiple phenotypes. MATERIALS AND METHODS: Surrogate-guided ensemble latent Dirichlet allocation (sureLDA) is a label-free multidimensional phenotyping method. It first uses the PheNorm algorithm to initialize probabilities based on 2 surrogate features for each target phenotype, and then leverages these probabilities to constrain the LDA topic model to generate phenotype-specific topics. Finally, it combines phenotype-feature counts with surrogates via clustering ensemble to yield final phenotype probabilities. RESULTS: sureLDA achieves reliably high accuracy and precision across a range of simulated and real-world phenotypes. Its performance is robust to phenotype prevalence and relative informativeness of surogate vs nonsurrogate features. It also exhibits powerful feature selection properties. DISCUSSION: sureLDA combines attractive properties of PheNorm and LDA to achieve high accuracy and precision robust to diverse phenotype characteristics. It offers particular improvement for phenotypes insufficiently captured by a few surrogate features. Moreover, sureLDA's feature selection ability enables it to handle high feature dimensions and produce interpretable computational phenotypes. CONCLUSIONS: sureLDA is well suited toward large-scale electronic health record phenotyping for highly multiphenotype applications such as phenome-wide association studies .

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

J Am Med Inform Assoc

DOI

EISSN

1527-974X

Publication Date

August 1, 2020

Volume

27

Issue

8

Start / End Page

1235 / 1243

Location

England

Related Subject Headings

  • Translational Research, Biomedical
  • ROC Curve
  • Precision Medicine
  • Natural Language Processing
  • Medical Informatics
  • Humans
  • Electronic Health Records
  • Algorithms
  • 46 Information and computing sciences
  • 42 Health sciences
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Ahuja, Y., Zhou, D., He, Z., Sun, J., Castro, V. M., Gainer, V., … Cai, T. (2020). sureLDA: A multidisease automated phenotyping method for the electronic health record. J Am Med Inform Assoc, 27(8), 1235–1243. https://doi.org/10.1093/jamia/ocaa079
Ahuja, Yuri, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian Gainer, Shawn N. Murphy, Chuan Hong, and Tianxi Cai. “sureLDA: A multidisease automated phenotyping method for the electronic health record.J Am Med Inform Assoc 27, no. 8 (August 1, 2020): 1235–43. https://doi.org/10.1093/jamia/ocaa079.
Ahuja Y, Zhou D, He Z, Sun J, Castro VM, Gainer V, et al. sureLDA: A multidisease automated phenotyping method for the electronic health record. J Am Med Inform Assoc. 2020 Aug 1;27(8):1235–43.
Ahuja, Yuri, et al. “sureLDA: A multidisease automated phenotyping method for the electronic health record.J Am Med Inform Assoc, vol. 27, no. 8, Aug. 2020, pp. 1235–43. Pubmed, doi:10.1093/jamia/ocaa079.
Ahuja Y, Zhou D, He Z, Sun J, Castro VM, Gainer V, Murphy SN, Hong C, Cai T. sureLDA: A multidisease automated phenotyping method for the electronic health record. J Am Med Inform Assoc. 2020 Aug 1;27(8):1235–1243.
Journal cover image

Published In

J Am Med Inform Assoc

DOI

EISSN

1527-974X

Publication Date

August 1, 2020

Volume

27

Issue

8

Start / End Page

1235 / 1243

Location

England

Related Subject Headings

  • Translational Research, Biomedical
  • ROC Curve
  • Precision Medicine
  • Natural Language Processing
  • Medical Informatics
  • Humans
  • Electronic Health Records
  • Algorithms
  • 46 Information and computing sciences
  • 42 Health sciences