Comparative Evaluation of a New Autonomous AI Agent Versus Frontier LLMs for AI-Generated Patient Information Sheets on Pediatric Pathologies
Background: At present, there are very limited studies evaluating autonomous artificial intelligence (AI) agents in oral oncology or healthcare education, and no studies have directly compared traditional frontier large language models (LLMs) with autonomous AI agents in the field of pediatric oral oncologic pathology. This study evaluated the performance of AI LLMs and an autonomous AI scientific agent in generating engaging and accessible patient information sheets for common pediatric oral pathologic conditions and rare head and neck tumors. The development of high-quality patient education materials is particularly important in pediatric pathology because parents must often navigate complex and emotionally sensitive diagnoses, including rare tumors and developmental lesions, for which accessible, patient-friendly educational resources are frequently unavailable. AI-generated patient information sheets therefore represent a potential strategy to improve communication, understanding, and shared decision-making for families facing these uncommon conditions. Methods: AI-generated patient information sheets from popular frontier chatbots on various pediatric pathological conditions and rare tumors were evaluated by five platforms (Eli5a v2.0, ChatGPT-5.4, Claude Sonnet 4.6, Perplexity and Doximity), in a blinded fashion, by three expert evaluators using the Global Quality Score (GQS), DISCERN score, understandability score and actionability score using PEMAT, and the Flesch-Kincaid Grade Level. Because the same 20 conditions were assessed on every platform, observations were paired; platforms were compared using Friedman tests with paired Wilcoxon signed-rank post hoc tests and Holm correction, and interrater reliability was quantified using an absolute-agreement intraclass correlation coefficient (ICC). Results: Platform differences were significant for all five outcomes (all p < 0.001). Claude Sonnet 4.6 obtained the best overall information quality (GQS 4.25 ± 0.39), scoring significantly higher than all other platforms, while Perplexity, Eli5a v2.0 and ChatGPT-5.4 performed similarly to one another and better than Doximity. For information reliability, Doximity (DISCERN 70.50 ± 3.20) and Claude Sonnet 4.6 (70.25 ± 3.02) performed best, while Eli5a v2.0 recorded the lowest DISCERN score (54.50 ± 4.26). Eli5a v2.0 obtained the best results for patient-centered communication, with the highest understandability (PEMAT-U 95.05 ± 1.23) and readability (FKGL 5.45 ± 0.51) and an actionability score among the highest of the five platforms (PEMAT-A 89.00 ± 2.62, not significantly different from Perplexity or Doximity). Interrater absolute agreement for GQS was poor to moderate (ICC(2,1) = 0.236; ICC(2,3) = 0.481). Conclusion: No single platform was superior across all evaluated domains. Claude Sonnet 4.6 led in overall quality whereas Eli5a v2.0 achieved the highest understandability and the most appropriatereading grade level. Eli5a’s actionability was high but did not differ significantly from Perplexity or Doximity. These findings indicate a trade-off between information quality and reliability on one hand and lay accessibility on the other. Platform selection should therefore be matched to the communication task, and all AI-generated patient materials require clinician review before use.
Authors
- Ahmed S. Sultan (ORCID: https://orcid.org/0000-0001-5286-4562)
- Zaid H. Khoury (ORCID: https://orcid.org/0000-0001-9596-3560)
- Jeffery B. Price (ORCID: https://orcid.org/0000-0002-9699-7591)
- Kimia Sadat Kazemi
- Tiffany Tavares
- Rata Rokhshad (ORCID: https://orcid.org/0000-0001-9668-7684)
- Neda Najafimakhsoos (ORCID: https://orcid.org/0009-0007-2376-7289)
- Mohamed S. Sultan
Institutions
- Meharry Medical College (US)
- University of Maryland, Baltimore (US)
- The University of Texas at San Antonio Health Science Center (US)
- Loma Linda University (US)
- University of Saskatchewan (CA)
- Eli Lilly (Switzerland) (CH)
- University of Maryland Marlene and Stewart Greenebaum Comprehensive Cancer Center
- University of Bologna (IT)
Publication Details
- Journal
- Cancers
- Published
- 2026-09-20
- DOI
- https://doi.org/10.3390/cancers18183047
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00