Skip to main content

Assessing the capability of large language models in answering pediatric critical care board-style questions.

Journal articles  - Journal Article
Chanci, D; Moore, R; Foote, HP; Goldstein, MA; Kumar, KR; Rotta, AT; Hornik, CP; Burriss-West, M; Hamilton, M; Stanco, N; Alvarez, M; Brown, A ...
Published in: Sci Rep
June 13, 2026

The potential of Large Language Models (LLMs) in medicine is often linked to massive, resource-intensive models. However, their practical application in specialized fields like pediatric critical care requires exploring the capability of more efficient, locally-deployable open-source alternatives. In this study, we evaluated the accuracy and clinical reasoning of open-source LLMs of varying sizes, specifically assessing if smaller, efficient models can perform comparably to larger ones on pediatric critical care multiple-choice questions. A set of 100 pediatric critical care MCQs across six clinical domains, i.e. calculation, diagnosis, ethics, management, pharmacology, and physiology, was curated by two pediatric specialists to evaluate eight open-source LLMs, ranging from 2 to 70 billion parameters. The LLMs were assessed using the overall and category-specific accuracy and clinical reasoning quality score based on a 5-points Likert scale. Additionally, six pediatric critical care fellows completed the MCQs for comparison. Cochran's Q test, McNemar's test, the Friedman test, Cohen's kappa, and Fleiss's kappa were used for the statistical analysis. While the largest model (Llama-3.3-70B) achieved the highest accuracy (78%; 95% CI, 69%-86%), a key finding was the performance of the much smaller, 14.7-billion parameter Phi-4. This efficient model was strikingly comparable, with 75% accuracy (95% CI, 65%-83%) and a similar reasoning score (4.40 vs. 4.49/5). Both models' performance was comparable to that of the pediatric critical care fellows included in this study. The LLMs showed a strong performance in ethics but struggled with calculations. Inter-rater reliability was excellent for the clinical reasoning assessment (κ = 0.92). Our findings demonstrate that smaller, efficient LLMs can approach the performance of much larger models and pediatric critical care fellows for complex pediatric critical care reasoning. These results support further investigation for developing secure, locally-deployable decision support tools without relying on massive, proprietary systems. At the same time, these models hold potential as complementary resources for trainee education in pediatric critical care. However, their identified weaknesses, especially in calculations, pose a prohibitive barrier that underscores the need for rigorous, domain-specific validation followed by evaluation in more realistic clinical scenarios to ensure safe use in both clinical and educational contexts.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Sci Rep

DOI

EISSN

2045-2322

Publication Date

June 13, 2026

Location

England
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Chanci, D., Moore, R., Foote, H. P., Goldstein, M. A., Kumar, K. R., Rotta, A. T., … Kamaleswaran, R. (2026). Assessing the capability of large language models in answering pediatric critical care board-style questions. Sci Rep. https://doi.org/10.1038/s41598-026-57353-0
Chanci, Daniela, Ronald Moore, Henry P. Foote, Matthew A. Goldstein, Karan R. Kumar, Alexandre T. Rotta, Christoph P. Hornik, et al. “Assessing the capability of large language models in answering pediatric critical care board-style questions.Sci Rep, June 13, 2026. https://doi.org/10.1038/s41598-026-57353-0.
Chanci D, Moore R, Foote HP, Goldstein MA, Kumar KR, Rotta AT, et al. Assessing the capability of large language models in answering pediatric critical care board-style questions. Sci Rep. 2026 Jun 13;
Chanci, Daniela, et al. “Assessing the capability of large language models in answering pediatric critical care board-style questions.Sci Rep, June 2026. Pubmed, doi:10.1038/s41598-026-57353-0.
Chanci D, Moore R, Foote HP, Goldstein MA, Kumar KR, Rotta AT, Hornik CP, Burriss-West M, Hamilton M, Stanco N, Alvarez M, Brown A, Johnson M, Kamaleswaran R. Assessing the capability of large language models in answering pediatric critical care board-style questions. Sci Rep. 2026 Jun 13;

Published In

Sci Rep

DOI

EISSN

2045-2322

Publication Date

June 13, 2026

Location

England