Skip to main content

Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content.

Journal articles  - Journal Article
Ataídes, RJC; Campos, MAG; Souza, JVPD; Rocha, RC; Lacalle, AA; Vieira, CB; Artioli, T; Medeiros, TC; Souza Filho, EMD; Gismondi, R; Romeo, FJ ...
Published in: Arq Bras Cardiol
February 2026

BACKGROUND: Artificial intelligence (AI) language models are increasingly used to generate patient education materials. However, their accuracy, completeness, and adherence to clinical guidelines remain uncertain. OBJECTIVES: To compare ChatGPT-Mini and ChatGPT-4.0 in the generation of hypertension education content with respect to accuracy, completeness, structural quality using the Ensuring Quality Information for Patients (EQIP), response consistency, and alignment with established guidelines. METHODS: A standardized set of 31 hypertension-related questions was submitted to both models. Outputs were independently evaluated by 10 blinded clinicians using a modified EQIP score, a 5-point accuracy scale, and a 3-point completeness scale. Response consistency was assessed using BERTScore. Between-model comparisons were performed using the two-sided Wilcoxon rank-sum test (p < 0.05). Effect sizes were reported as Hodges-Lehmann (HL) median differences and Cliff's delta (δ), both with 95% CIs. Inter-rater reliability was estimated using the intraclass correlation coefficient (ICC; two-way random effects model, absolute agreement). RESULTS: Central tendency measures favored ChatGPT-4.0, although differences were small. Median scores were as follows: accuracy, 4.10 (3.70-4.20) versus 3.73 (3.60-4.05); completeness, 1.26 (1.17-1.41) versus 1.10 (0.96-1.23); and total EQIP score, 19.5 (18.0-25.0) versus 18.5 (16.0-23.0) for ChatGPT-4.0 and ChatGPT-Mini, respectively. HL median differences were small, with 95% CIs crossing zero (accuracy: +0.37, -0.25 to +0.50; completeness: +0.16, -0.06 to +0.36; EQIP: +1.0, -1.0 to +6.0). Cliff's δ values were consistently small and positive across primary outcomes, indicating only modest stochastic dominance of ChatGPT-4.0. Identification clarity tended to be higher with ChatGPT-4.0, whereas response consistency measured by BERTScore F1 was generally higher for ChatGPT-Mini (> 0.92 versus 0.885-0.932). Inter-rater reliability was good to excellent across all measures (ICC > 0.80). CONCLUSIONS: ChatGPT-4.0 demonstrated small, non-significant improvements in accuracy, completeness, and structural quality compared with ChatGPT-Mini. Effect sizes were modest, and all 95% CIs included zero. ChatGPT-Mini produced more consistent responses. These findings underscore the importance of routinely reporting effect sizes with 95% CIs and support the use of standardized evaluation methods and real-time validation frameworks for AI-generated medical education content.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Arq Bras Cardiol

DOI

EISSN

1678-4170

Publication Date

February 2026

Volume

123

Issue

2

Start / End Page

e20250498

Location

Brazil

Related Subject Headings

  • Statistics, Nonparametric
  • Reproducibility of Results
  • Practice Guidelines as Topic
  • Patient Education as Topic
  • Large Language Models
  • Hypertension
  • Humans
  • Generative Artificial Intelligence
  • Cardiovascular System & Hematology
  • 3201 Cardiovascular medicine and haematology
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Ataídes, R. J. C., Campos, M. A. G., Souza, J. V. P. D., Rocha, R. C., Lacalle, A. A., Vieira, C. B., … Lopes, R. D. (2026). Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content. Arq Bras Cardiol, 123(2), e20250498. https://doi.org/10.36660/abc.20250498
Ataídes, Rômullo José Costa, Marcos Adriano Garcia Campos, João Vítor Perez de Souza, Rafael Cardoso Rocha, Almir Alamino Lacalle, Ciro Bezerra Vieira, Thiago Artioli, et al. “Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content.Arq Bras Cardiol 123, no. 2 (February 2026): e20250498. https://doi.org/10.36660/abc.20250498.
Ataídes RJC, Campos MAG, Souza JVPD, Rocha RC, Lacalle AA, Vieira CB, et al. Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content. Arq Bras Cardiol. 2026 Feb;123(2):e20250498.
Ataídes, Rômullo José Costa, et al. “Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content.Arq Bras Cardiol, vol. 123, no. 2, Feb. 2026, p. e20250498. Pubmed, doi:10.36660/abc.20250498.
Ataídes RJC, Campos MAG, Souza JVPD, Rocha RC, Lacalle AA, Vieira CB, Artioli T, Medeiros TC, Souza Filho EMD, Gismondi R, Campana ÉMG, Romeo FJ, Razuk V, Vissoci JRN, Lopes RD. Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content. Arq Bras Cardiol. 2026 Feb;123(2):e20250498.

Published In

Arq Bras Cardiol

DOI

EISSN

1678-4170

Publication Date

February 2026

Volume

123

Issue

2

Start / End Page

e20250498

Location

Brazil

Related Subject Headings

  • Statistics, Nonparametric
  • Reproducibility of Results
  • Practice Guidelines as Topic
  • Patient Education as Topic
  • Large Language Models
  • Hypertension
  • Humans
  • Generative Artificial Intelligence
  • Cardiovascular System & Hematology
  • 3201 Cardiovascular medicine and haematology