Skip to main content

LLM-assisted systematic review of large language models in clinical medicine.

Journal articles  - Systematic Review, Journal Article
Chen, SF; Alyakin, A; Seas, A; Yang, E; Choi, JJ; Lee, JV; Chen, AL; Warman, PI; Bitolas, RT; Steele, RJ; Alber, DA; Oermann, EK
Published in: Nature medicine
March 2026

Clinical evaluations of large language models (LLMs) have rapidly expanded since 2022, yet their evidence base remains opaque. The overwhelming volume of studies creates challenges for manual curation and review. However, LLMs themselves offer the scalability and capability to evaluate the ever-growing evidence base. This LLM-assisted review identified 4,609 peer-reviewed studies in clinical medicine between January 2022 and September 2025, equating to roughly 3.2 papers per day. Only 1,048 studies used real-world patient data and of these only 19 were prospective randomized trials; most addressed simulated scenarios (n = 1,857) or exam-style tasks (n = 1,704). ChatGPT and related OpenAI models constitute 65.7% of evaluated models, with Gemini/Bard a distant second constituting 13.1% of evaluated models. Patient-facing communication and education comprised 17% of tasks, followed by knowledge retrieval, and education and assessment simulation. Across 1,046 head-to-head comparisons, LLMs outperformed humans in 33% of comparisons, with a strong dependency on task realism and level of training. At least 25% of studies had sample sizes less than 30. Despite the growth of LLMs in medicine, rigorous, patient-centered evidence remains scarce, underscoring the need for larger prospective trials before clinical adoption.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Nature medicine

DOI

EISSN

1546-170X

ISSN

1078-8956

Publication Date

March 2026

Volume

32

Issue

3

Start / End Page

1152 / 1159

Related Subject Headings

  • Large Language Models
  • Immunology
  • Humans
  • Generative Artificial Intelligence
  • 42 Health sciences
  • 32 Biomedical and clinical sciences
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Chen, S. F., Alyakin, A., Seas, A., Yang, E., Choi, J. J., Lee, J. V., … Oermann, E. K. (2026). LLM-assisted systematic review of large language models in clinical medicine. Nature Medicine, 32(3), 1152–1159. https://doi.org/10.1038/s41591-026-04229-5
Chen, Sully F., Anton Alyakin, Andreas Seas, Eunice Yang, Joanne J. Choi, Jin Vivian Lee, Amelia L. Chen, et al. “LLM-assisted systematic review of large language models in clinical medicine.Nature Medicine 32, no. 3 (March 2026): 1152–59. https://doi.org/10.1038/s41591-026-04229-5.
Chen SF, Alyakin A, Seas A, Yang E, Choi JJ, Lee JV, et al. LLM-assisted systematic review of large language models in clinical medicine. Nature medicine. 2026 Mar;32(3):1152–9.
Chen, Sully F., et al. “LLM-assisted systematic review of large language models in clinical medicine.Nature Medicine, vol. 32, no. 3, Mar. 2026, pp. 1152–59. Epmc, doi:10.1038/s41591-026-04229-5.
Chen SF, Alyakin A, Seas A, Yang E, Choi JJ, Lee JV, Chen AL, Warman PI, Bitolas RT, Steele RJ, Alber DA, Oermann EK. LLM-assisted systematic review of large language models in clinical medicine. Nature medicine. 2026 Mar;32(3):1152–1159.

Published In

Nature medicine

DOI

EISSN

1546-170X

ISSN

1078-8956

Publication Date

March 2026

Volume

32

Issue

3

Start / End Page

1152 / 1159

Related Subject Headings

  • Large Language Models
  • Immunology
  • Humans
  • Generative Artificial Intelligence
  • 42 Health sciences
  • 32 Biomedical and clinical sciences