Structural quality, readability, and educational usability of large language model chatbot responses to parental questions on teething

Large language model (LLM)–based chatbots are increasingly used by parents to obtain information about pediatric oral health, including concerns related to tooth eruption and teething. However, the structural quality, readability, and educational usability of AI-generated responses addressing these parental inquiries remain insufficiently explored. In this cross-sectional analytical study, ten frequently asked parental questions retrieved from an online forum were submitted to four widely accessible chatbots ChatGPT-4o, Gemini Advanced, Microsoft Copilot, and Claude Sonnet 4 across three consecutive days using standardized interaction protocols. A total of 120 responses were analyzed. Structural quality was assessed using the DISCERN instrument and the Global Quality Score (GQS). Readability was evaluated using the Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), and Gunning Fog Index (GFI). Educational usability was assessed using the Patient Education Materials Assessment Tool (PEMAT). Inter-model comparisons were performed using the Friedman test to account for the repeated-measures structure, with multiplicity adjustment using the Holm method. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and associations among quality, readability, and usability metrics were examined using Spearman's rank correlation, with false discovery rate correction for multiple comparisons. After adjustment for multiple testing, DISCERN scores differed significantly among chatbots on Day 3 (adjusted p = 0.046), whereas no significant adjusted inter-model differences were detected on Days 1 or 2. No significant adjusted differences in GQS scores were observed across the three evaluation sessions. Inter-rater agreement was moderate for DISCERN (ICC = 0.620), GQS (ICC = 0.508), PEMAT-Understandability (ICC = 0.740), and PEMAT-Actionability (ICC = 0.710). Significant inter-model differences were observed for all three readability indices (FRE, FKGL, and GFI; adjusted p < 0.05), whereas no significant differences were detected for PEMAT understandability or actionability. Correlation analyses demonstrated a strong positive association between DISCERN and GQS scores (ρ = 0.76, FDR-adjusted q < 0.001), while associations between structural quality and readability measures did not remain significant after FDR correction. Differences among the evaluated chatbots depended on the dimension assessed. Structural quality differences were session-dependent, while significant inter-model differences in readability were consistently identified, and no significant differences were detected in educational usability. These findings underscore the need to evaluate AI-generated pediatric dental information across multiple dimensions and to optimize its accessibility for caregivers.

Authors

Institutions

Publication Details

Journal
BMC Oral Health
Published
2026-09-15
DOI
https://doi.org/10.1186/s12903-026-09805-2
Primary Topic
Mobile Health and mHealth Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Structural quality, readability, and educational usability of large language model chatbot responses to parental questions on teething

Yelda Kasımoğlu, Selin Saygılı, Eda Sır, Çağan Taş et al.
BMC Oral Health
Mobile Health and mHealth Applications
article

Structural quality, readability, and educational usability of large language model chatbot responses to parental questions on teething

Yelda Kasımoğlu, Selin Saygılı, Eda Sır, Çağan Taş, Hülya Çerçi Akçay
article en

Abstract

Large language model (LLM)–based chatbots are increasingly used by parents to obtain information about pediatric oral health, including concerns related to tooth eruption and teething. However, the structural quality, readability, and educational usability of AI-generated responses addressing these parental inquiries remain insufficiently explored. In this cross-sectional analytical study, ten frequently asked parental questions retrieved from an online forum were submitted to four widely accessible chatbots ChatGPT-4o, Gemini Advanced, Microsoft Copilot, and Claude Sonnet 4 across three consecutive days using standardized interaction protocols. A total of 120 responses were analyzed. Structural quality was assessed using the DISCERN instrument and the Global Quality Score (GQS). Readability was evaluated using the Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), and Gunning Fog Index (GFI). Educational usability was assessed using the Patient Education Materials Assessment Tool (PEMAT). Inter-model comparisons were performed using the Friedman test to account for the repeated-measures structure, with multiplicity adjustment using the Holm method. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and associations among quality, readability, and usability metrics were examined using Spearman's rank correlation, with false discovery rate correction for multiple comparisons. After adjustment for multiple testing, DISCERN scores differed significantly among chatbots on Day 3 (adjusted p = 0.046), whereas no significant adjusted inter-model differences were detected on Days 1 or 2. No significant adjusted differences in GQS scores were observed across the three evaluation sessions. Inter-rater agreement was moderate for DISCERN (ICC = 0.620), GQS (ICC = 0.508), PEMAT-Understandability (ICC = 0.740), and PEMAT-Actionability (ICC = 0.710). Significant inter-model differences were observed for all three readability indices (FRE, FKGL, and GFI; adjusted p < 0.05), whereas no significant differences were detected for PEMAT understandability or actionability. Correlation analyses demonstrated a strong positive association between DISCERN and GQS scores (ρ = 0.76, FDR-adjusted q < 0.001), while associations between structural quality and readability measures did not remain significant after FDR correction. Differences among the evaluated chatbots depended on the dimension assessed. Structural quality differences were session-dependent, while significant inter-model differences in readability were consistently identified, and no significant differences were detected in educational usability. These findings underscore the need to evaluate AI-generated pediatric dental information across multiple dimensions and to optimize its accessibility for caregivers.

BMC Oral Health
Kocaeli Üniversitesi (TR), Istanbul University (TR)
Quality Education
Openalex Percentile: Top 6%
Mobile Health and mHealth Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.