Comparative Evaluation of a New Autonomous AI Agent Versus Frontier LLMs for AI-Generated Patient Information Sheets on Pediatric Pathologies

Background: At present, there are very limited studies evaluating autonomous artificial intelligence (AI) agents in oral oncology or healthcare education, and no studies have directly compared traditional frontier large language models (LLMs) with autonomous AI agents in the field of pediatric oral oncologic pathology. This study evaluated the performance of AI LLMs and an autonomous AI scientific agent in generating engaging and accessible patient information sheets for common pediatric oral pathologic conditions and rare head and neck tumors. The development of high-quality patient education materials is particularly important in pediatric pathology because parents must often navigate complex and emotionally sensitive diagnoses, including rare tumors and developmental lesions, for which accessible, patient-friendly educational resources are frequently unavailable. AI-generated patient information sheets therefore represent a potential strategy to improve communication, understanding, and shared decision-making for families facing these uncommon conditions. Methods: AI-generated patient information sheets from popular frontier chatbots on various pediatric pathological conditions and rare tumors were evaluated by five platforms (Eli5a v2.0, ChatGPT-5.4, Claude Sonnet 4.6, Perplexity and Doximity), in a blinded fashion, by three expert evaluators using the Global Quality Score (GQS), DISCERN score, understandability score and actionability score using PEMAT, and the Flesch-Kincaid Grade Level. Because the same 20 conditions were assessed on every platform, observations were paired; platforms were compared using Friedman tests with paired Wilcoxon signed-rank post hoc tests and Holm correction, and interrater reliability was quantified using an absolute-agreement intraclass correlation coefficient (ICC). Results: Platform differences were significant for all five outcomes (all p < 0.001). Claude Sonnet 4.6 obtained the best overall information quality (GQS 4.25 ± 0.39), scoring significantly higher than all other platforms, while Perplexity, Eli5a v2.0 and ChatGPT-5.4 performed similarly to one another and better than Doximity. For information reliability, Doximity (DISCERN 70.50 ± 3.20) and Claude Sonnet 4.6 (70.25 ± 3.02) performed best, while Eli5a v2.0 recorded the lowest DISCERN score (54.50 ± 4.26). Eli5a v2.0 obtained the best results for patient-centered communication, with the highest understandability (PEMAT-U 95.05 ± 1.23) and readability (FKGL 5.45 ± 0.51) and an actionability score among the highest of the five platforms (PEMAT-A 89.00 ± 2.62, not significantly different from Perplexity or Doximity). Interrater absolute agreement for GQS was poor to moderate (ICC(2,1) = 0.236; ICC(2,3) = 0.481). Conclusion: No single platform was superior across all evaluated domains. Claude Sonnet 4.6 led in overall quality whereas Eli5a v2.0 achieved the highest understandability and the most appropriatereading grade level. Eli5a’s actionability was high but did not differ significantly from Perplexity or Doximity. These findings indicate a trade-off between information quality and reliability on one hand and lay accessibility on the other. Platform selection should therefore be matched to the communication task, and all AI-generated patient materials require clinician review before use.

Authors

Institutions

Publication Details

Journal
Cancers
Published
2026-09-20
DOI
https://doi.org/10.3390/cancers18183047
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative Evaluation of a New Autonomous AI Agent Versus Frontier LLMs for AI-Generated Patient Information Sheets on Pediatric Pathologies

Ahmed S. Sultan, Zaid H. Khoury, Jeffery B. Price, Kimia Sadat Kazemi et al.
Cancers
Artificial Intelligence in Healthcare and Education
article

Comparative Evaluation of a New Autonomous AI Agent Versus Frontier LLMs for AI-Generated Patient Information Sheets on Pediatric Pathologies

Ahmed S. Sultan, Zaid H. Khoury, Jeffery B. Price, Kimia Sadat Kazemi, Tiffany Tavares, Rata Rokhshad, Neda Najafimakhsoos, Mohamed S. Sultan
article en

Abstract

Background: At present, there are very limited studies evaluating autonomous artificial intelligence (AI) agents in oral oncology or healthcare education, and no studies have directly compared traditional frontier large language models (LLMs) with autonomous AI agents in the field of pediatric oral oncologic pathology. This study evaluated the performance of AI LLMs and an autonomous AI scientific agent in generating engaging and accessible patient information sheets for common pediatric oral pathologic conditions and rare head and neck tumors. The development of high-quality patient education materials is particularly important in pediatric pathology because parents must often navigate complex and emotionally sensitive diagnoses, including rare tumors and developmental lesions, for which accessible, patient-friendly educational resources are frequently unavailable. AI-generated patient information sheets therefore represent a potential strategy to improve communication, understanding, and shared decision-making for families facing these uncommon conditions. Methods: AI-generated patient information sheets from popular frontier chatbots on various pediatric pathological conditions and rare tumors were evaluated by five platforms (Eli5a v2.0, ChatGPT-5.4, Claude Sonnet 4.6, Perplexity and Doximity), in a blinded fashion, by three expert evaluators using the Global Quality Score (GQS), DISCERN score, understandability score and actionability score using PEMAT, and the Flesch-Kincaid Grade Level. Because the same 20 conditions were assessed on every platform, observations were paired; platforms were compared using Friedman tests with paired Wilcoxon signed-rank post hoc tests and Holm correction, and interrater reliability was quantified using an absolute-agreement intraclass correlation coefficient (ICC). Results: Platform differences were significant for all five outcomes (all p < 0.001). Claude Sonnet 4.6 obtained the best overall information quality (GQS 4.25 ± 0.39), scoring significantly higher than all other platforms, while Perplexity, Eli5a v2.0 and ChatGPT-5.4 performed similarly to one another and better than Doximity. For information reliability, Doximity (DISCERN 70.50 ± 3.20) and Claude Sonnet 4.6 (70.25 ± 3.02) performed best, while Eli5a v2.0 recorded the lowest DISCERN score (54.50 ± 4.26). Eli5a v2.0 obtained the best results for patient-centered communication, with the highest understandability (PEMAT-U 95.05 ± 1.23) and readability (FKGL 5.45 ± 0.51) and an actionability score among the highest of the five platforms (PEMAT-A 89.00 ± 2.62, not significantly different from Perplexity or Doximity). Interrater absolute agreement for GQS was poor to moderate (ICC(2,1) = 0.236; ICC(2,3) = 0.481). Conclusion: No single platform was superior across all evaluated domains. Claude Sonnet 4.6 led in overall quality whereas Eli5a v2.0 achieved the highest understandability and the most appropriatereading grade level. Eli5a’s actionability was high but did not differ significantly from Perplexity or Doximity. These findings indicate a trade-off between information quality and reliability on one hand and lay accessibility on the other. Platform selection should therefore be matched to the communication task, and all AI-generated patient materials require clinician review before use.

CancersVol. 18(18)
Meharry Medical College (US), University of Maryland, Baltimore (US), The University of Texas at San Antonio Health Science Center (US), Loma Linda University (US), University of Saskatchewan (CA), Eli Lilly (Switzerland) (CH), University of Maryland Marlene and Stewart Greenebaum Comprehensive Cancer Center, University of Bologna (IT)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.