Comparative Evaluation of Large Language Models in Responding to Common Misconceptions Regarding Periodontal Disease and Oral Hygiene
Background/Objective: As large language models (LLMs) are increasingly used to obtain oral health information, their ability to provide accurate and comprehensive responses to patient-focused periodontal questions derived from commonly reported misconceptions is increasingly relevant to patient education. This study aimed to evaluate the performance of contemporary LLMs in answering such questions with respect to accuracy, comprehensiveness, readability, and temporal consistency. Methods: Following a comprehensive review of the existing literature, structured patient-focused periodontal questions derived from commonly reported misconceptions were presented to seven different LLMs using a standardized prompt, and their performance was compared. Responses were evaluated for accuracy and comprehensiveness using a Likert scale, while readability was assessed using the Simple Measure of Gobbledygook (SMOG), the Flesch Reading Ease (FRE) score, and the Flesch–Kincaid Grade Level (FKGL). Results: The evaluated LLMs demonstrated uniformly high performance across the benchmark, with no individual response scoring below the scale midpoint (score < 3) for accuracy or comprehensiveness, confirming the absence of clinically unsafe or misleading information. Although statistically significant inter-model differences were observed for accuracy (p < 0.001, Kendall’s W = 0.197) and comprehensiveness (p < 0.001, Kendall’s W = 0.257), the rank-based effect sizes were small, indicating that relative rankings reflect minor numerical variations rather than clinically substantial performance gaps. Regarding readability, absolute metrics across all models revealed a high level of structural and linguistic difficulty that exceeds layperson health literacy requirements. Mean FRE scores across models ranged from 39.5 ± 9.1 to 52.1 ± 12.0 (indicating “difficult” to “fairly difficult” reading ease), mean FKGL scores ranged from 9.1 ± 2.2 to 10.8 ± 1.3 (requiring a 9th to 11th-grade education level), and mean SMOG scores ranged from 11.8 ± 1.6 to 14.0 ± 2.2. Statistically significant differences with large effect sizes were observed among models for FRE (p < 0.001, partial ηp2 = 0.308, ηG2 = 0.178), FKGL (p < 0.001, partial ηp2 = 0.162, ηG2 = 0.104) and SMOG (p < 0.001, partial ηp2 = 0.217, ηG2 = 0.140), with DeepSeek V4 and Gemini 3.1 exhibiting relatively more favorable readability profiles. In the temporal consistency analysis, significant differences in change scores were observed only for SMOG (p = 0.004), whereas no significant temporal changes occurred for accuracy, comprehensiveness, FRE, or FKGL. Conclusions: Contemporary LLMs provide uniformly accurate and clinically safe responses to patient-focused periodontal misconceptions, with scores consistently concentrated at the ceiling. However, absolute readability metrics demonstrate that AI-generated outputs consistently exceed standard public health literacy comprehension levels, presenting a critical barrier for general patient communication that necessitates targeted optimization.
Authors
- Resül Çolak (ORCID: https://orcid.org/0000-0001-5210-1119)
- Orhan Çiçek (ORCID: https://orcid.org/0000-0002-8172-6043)
- Merve Küçükoğlu Çolak (ORCID: https://orcid.org/0000-0002-7652-0016)
- İsmail Gül (ORCID: https://orcid.org/0009-0001-8550-3250)
Institutions
- Zonguldak Bülent Ecevit University (TR)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/diagnostics16193158
- Primary Topic
- Health Literacy and Information Accessibility
- Type
- article
- Field-Weighted Citation Impact
- 0.00