The readability and quality paradox: Comparing ChatGPT, Gemini, and Perplexity outputs on pediatric chest pain queries

Pediatric chest pain represents a frequently encountered symptom in emergency departments and outpatient clinics throughout childhood, often generating substantial parental anxiety. This study aims to comparatively analyze the readability, information quality, and scientific reliability of textual contents generated by the artificial intelligence (AI) chatbots ChatGPT, Gemini, and Perplexity regarding pediatric chest pain, utilizing multidimensional analytical indices. Out of the top 25 queries with the highest global search volume on Google Trends, 17 unique keywords meeting the predefined inclusion criteria were filtered on June 1, 2026. These inquiries were directed to all three AI platforms in distinct, independent user sessions, yielding a total of 51 textual responses. Linguistic readability levels were calculated across six separate digital interfaces using the FKGL, FRES and other formulas, and the results were benchmarked against the sixth-grade reading level” threshold. Scientific reliability was audited via the Modified DISCERN and JAMA benchmarks, while content quality was examined using the GQS and EQIP instruments. The median readability scores computed across all three AI models were found to be statistically and significantly above the targeted sixth-grade comprehension threshold, indicating a high level of difficulty (p < 0.001). Post-hoc pairwise evaluations revealed that ChatGPT and Gemini offered linguistically more accessible and structurally less complex textual architectures; conversely, Perplexity exhibited a significantly more difficult and academic linguistic framework (p < 0.0167). On the other hand, regarding content quality and source credibility, Perplexity demonstrated remarkably superior and more optimized scores across all evaluated instruments, namely GQS (p = 0.002, p = 0.001), JAMA (p < 0.001, p < 0.001), mDISCERN (p = 0.005, p = 0.001), and EQIP (p = 0.003, p < 0.001), compared directly to both ChatGPT and Gemini, respectively. Although generative AI formulations harbor considerable potential to provide extensive data on pediatric chest pain, the sophisticated, university-level linguistic architecture of these texts poses a substantial access barrier for individuals who lack proficient digital health literacy skills. While Perplexity achieved the performance closest to the “gold standard” in terms of informational accuracy, no model is currently mature enough to substitute for a professional medical consultation due to observed citation biases and omissions. It is of paramount importance that clinicians actively guide families to scrutinize online health data through a critical lens.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-25
DOI
https://doi.org/10.1371/journal.pone.0359170
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The readability and quality paradox: Comparing ChatGPT, Gemini, and Perplexity outputs on pediatric chest pain queries

Hikmet Kıztanır, Ebru Çetin Özbek
PLoS ONE
Artificial Intelligence in Healthcare and Education
article

The readability and quality paradox: Comparing ChatGPT, Gemini, and Perplexity outputs on pediatric chest pain queries

Hikmet Kıztanır, Ebru Çetin Özbek
article en

Abstract

Pediatric chest pain represents a frequently encountered symptom in emergency departments and outpatient clinics throughout childhood, often generating substantial parental anxiety. This study aims to comparatively analyze the readability, information quality, and scientific reliability of textual contents generated by the artificial intelligence (AI) chatbots ChatGPT, Gemini, and Perplexity regarding pediatric chest pain, utilizing multidimensional analytical indices. Out of the top 25 queries with the highest global search volume on Google Trends, 17 unique keywords meeting the predefined inclusion criteria were filtered on June 1, 2026. These inquiries were directed to all three AI platforms in distinct, independent user sessions, yielding a total of 51 textual responses. Linguistic readability levels were calculated across six separate digital interfaces using the FKGL, FRES and other formulas, and the results were benchmarked against the sixth-grade reading level” threshold. Scientific reliability was audited via the Modified DISCERN and JAMA benchmarks, while content quality was examined using the GQS and EQIP instruments. The median readability scores computed across all three AI models were found to be statistically and significantly above the targeted sixth-grade comprehension threshold, indicating a high level of difficulty (p < 0.001). Post-hoc pairwise evaluations revealed that ChatGPT and Gemini offered linguistically more accessible and structurally less complex textual architectures; conversely, Perplexity exhibited a significantly more difficult and academic linguistic framework (p < 0.0167). On the other hand, regarding content quality and source credibility, Perplexity demonstrated remarkably superior and more optimized scores across all evaluated instruments, namely GQS (p = 0.002, p = 0.001), JAMA (p < 0.001, p < 0.001), mDISCERN (p = 0.005, p = 0.001), and EQIP (p = 0.003, p < 0.001), compared directly to both ChatGPT and Gemini, respectively. Although generative AI formulations harbor considerable potential to provide extensive data on pediatric chest pain, the sophisticated, university-level linguistic architecture of these texts poses a substantial access barrier for individuals who lack proficient digital health literacy skills. While Perplexity achieved the performance closest to the “gold standard” in terms of informational accuracy, no model is currently mature enough to substitute for a professional medical consultation due to observed citation biases and omissions. It is of paramount importance that clinicians actively guide families to scrutinize online health data through a critical lens.

PLoS ONEVol. 21(9)
Recep Tayyip Erdoğan University (TR), University of Health Sciences Antigua (AG)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.