Structural quality, readability, and educational usability of large language model chatbot responses to parental questions on teething
Large language model (LLM)–based chatbots are increasingly used by parents to obtain information about pediatric oral health, including concerns related to tooth eruption and teething. However, the structural quality, readability, and educational usability of AI-generated responses addressing these parental inquiries remain insufficiently explored. In this cross-sectional analytical study, ten frequently asked parental questions retrieved from an online forum were submitted to four widely accessible chatbots ChatGPT-4o, Gemini Advanced, Microsoft Copilot, and Claude Sonnet 4 across three consecutive days using standardized interaction protocols. A total of 120 responses were analyzed. Structural quality was assessed using the DISCERN instrument and the Global Quality Score (GQS). Readability was evaluated using the Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), and Gunning Fog Index (GFI). Educational usability was assessed using the Patient Education Materials Assessment Tool (PEMAT). Inter-model comparisons were performed using the Friedman test to account for the repeated-measures structure, with multiplicity adjustment using the Holm method. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and associations among quality, readability, and usability metrics were examined using Spearman's rank correlation, with false discovery rate correction for multiple comparisons. After adjustment for multiple testing, DISCERN scores differed significantly among chatbots on Day 3 (adjusted p = 0.046), whereas no significant adjusted inter-model differences were detected on Days 1 or 2. No significant adjusted differences in GQS scores were observed across the three evaluation sessions. Inter-rater agreement was moderate for DISCERN (ICC = 0.620), GQS (ICC = 0.508), PEMAT-Understandability (ICC = 0.740), and PEMAT-Actionability (ICC = 0.710). Significant inter-model differences were observed for all three readability indices (FRE, FKGL, and GFI; adjusted p < 0.05), whereas no significant differences were detected for PEMAT understandability or actionability. Correlation analyses demonstrated a strong positive association between DISCERN and GQS scores (ρ = 0.76, FDR-adjusted q < 0.001), while associations between structural quality and readability measures did not remain significant after FDR correction. Differences among the evaluated chatbots depended on the dimension assessed. Structural quality differences were session-dependent, while significant inter-model differences in readability were consistently identified, and no significant differences were detected in educational usability. These findings underscore the need to evaluate AI-generated pediatric dental information across multiple dimensions and to optimize its accessibility for caregivers.
Authors
- Yelda Kasımoğlu (ORCID: https://orcid.org/0000-0003-1022-2486)
- Selin Saygılı (ORCID: https://orcid.org/0000-0001-9772-7828)
- Eda Sır
- Çağan Taş
- Hülya Çerçi Akçay
Institutions
- Kocaeli Üniversitesi (TR)
- Istanbul University (TR)
Publication Details
- Journal
- BMC Oral Health
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1186/s12903-026-09805-2
- Primary Topic
- Mobile Health and mHealth Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00