Assessing the accuracy and usability of artificial intelligence–based language models in responding to common periodontal patient questions

{"BACKGROUND:":[0],"Artificial":[1],"intelligence-powered":[2],"large":[3],"language":[4,288],"models":[5,450],"(LLMs)":[6],"are":[7,319,459,476],"increasingly":[8,382],"used":[9,48],"by":[10,89,96,259],"patients":[11,294,363,381],"seeking":[12],"quick":[13],"information":[14,361,401],"regarding":[15,26],"dental":[16,245,428],"and":[17,30,44,73,81,87,126,142,234,265,279,290,316,325,335,351,368,406,479],"medical":[18],"problems.":[19],"Despite":[20],"their":[21,427],"growing":[22],"popularity,":[23],"concerns":[24],"remain":[25,284],"the":[27,42,114,165,201,215,343,394,409,447,482],"accuracy,":[28],"clarity,":[29],"clinical":[31,344],"usefulness":[32],"of":[33,46],"LLM-generated":[34],"responses.":[35],"This":[36,430],"study":[37,431],"aimed":[38],"to":[39,52,75,321,364,397,421,485],"comparatively":[40],"evaluate":[41],"accuracy":[43],"usability":[45,119,207,292],"widely":[47],"LLMs":[49,77,219,248,299],"in":[50,78,252,286,291,437,468,471],"responding":[51],"frequently":[53],"asked":[54,64],"periodontal":[55,65,255,338],"questions.":[56,229,429],"METHODS:":[57],"In":[58],"this":[59],"analytical-comparative":[60],"study,":[61],"15":[62],"commonly":[63],"questions":[66,225,442],"were":[67,84,94],"selected":[68],"from":[69],"real":[70],"patient":[71,306],"encounters":[72],"administered":[74],"ten":[76],"both":[79],"English":[80,179,469],"Persian.":[82],"Responses":[83],"evaluated":[85,197],"independently":[86],"blindly":[88],"two":[90,460],"board-certified":[91],"periodontists;":[92],"disagreements":[93],"resolved":[95],"discussion":[97],"or":[98],"third-reviewer.":[99],"Each":[100],"response":[101],"was":[102,147],"rated":[103],"on":[104,223,227,384],"a":[105,375,492,499,503],"5-point":[106],"Likert":[107],"scale":[108],"across":[109,168],"six":[110],"criteria:":[111],"correlation":[112,199],"with":[113,122,139,160,176,210,240,295,362,408,426,502],"question,":[115],"adequacy,":[116],"comprehensiveness,":[117],"clarity/readability,":[118],"for":[120,208,293,305,347,386,424,481],"individuals":[121,209],"limited":[123,211,296],"scientific":[124,127,212,297],"literacy,":[125],"accuracy.":[128],"Statistical":[129],"analyses":[130],"included":[131],"descriptive":[132],"statistics,":[133],"Mann-Whitney":[134],"U":[135,190],"test,":[136],"Kruskal-Wallis":[137],"test":[138],"Bonferroni":[140],"correction,":[141],"univariate":[143],"ANOVA.":[144],"Significance":[145],"level":[146],"set":[148],"at":[149],"p":[150,156,193],"<":[151,157,194],"0.05.":[152],"RESULTS:":[153],"=":[154,191],"294.78,":[155],"0.001).":[158,195],"ChatGPT-4.5":[159,231,241,274],"deep":[161],"search":[162],"enabled":[163],"achieved":[164],"highest":[166,202],"scores":[167],"most":[169,448],"criteria,":[170,198],"while":[171,488],"ChatGPT-4o":[172],"generally":[173,336],"underperformed":[174],"compared":[175],"other":[177],"LLMs.":[178],"responses":[180],"scored":[181],"significantly":[182,221,466],"higher":[183],"than":[184,226,470],"Persian":[185,287],"(mean":[186],"4.25":[187],"vs.":[188],"3.96;":[189],"333,058,":[192],"Among":[196],"had":[200,214],"mean":[203],"score":[204,217],"(4.81),":[205],"whereas":[206],"literacy":[213],"lowest":[216],"(3.80).":[218],"performed":[220],"better":[222,467],"treatment/prevention":[224],"diagnosis/pathogenesis":[228],"Only":[230],"(deep":[232,236],"search)":[233,237],"Grok-3":[235],"provided":[238],"citations,":[239],"referencing":[242],"more":[243],"specialized":[244],"sources.":[246],"CONCLUSIONS:":[247],"show":[249],"variable":[250],"performance":[251,464],"answering":[253],"frequent":[254],"patients'":[256],"questions,":[257],"influenced":[258],"model":[260],"architecture,":[261],"language,":[262],"question":[263],"type,":[264],"evaluative":[266],"domain.":[267],"Although":[268],"advanced":[269,449],"configurations":[270],"such":[271],"as":[272,302,393,452],"deep-search":[273],"offer":[275],"highly":[276,333],"accurate,":[277,405,455],"comprehensive,":[278],"well-referenced":[280],"responses,":[281],"significant":[282],"limitations":[283],"particularly":[285],"output":[289],"literacy.":[298],"may":[300],"serve":[301],"complementary":[303],"tools":[304,331,435],"communication,":[307],"but":[308],"should":[309,356,496],"not":[310],"replace":[311,498],"professional":[312,500],"advice.":[313],"Further":[314],"development":[315],"language-specific":[317],"optimization":[318],"needed":[320],"improve":[322],"global":[323],"accessibility":[324],"reliability.":[326],"KEY":[327],"POINTS:":[328],"While":[329,446],"AI":[330,385,489],"provide":[332,454],"readable":[334],"accurate":[337],"information,":[339,457],"they":[340,440],"often":[341],"lack":[342],"nuance":[345],"required":[346],"personalized":[348],"risk":[349],"assessment":[350],"complex":[352],"treatment":[353,412],"planning.":[354],"Practitioners":[355],"proactively":[357],"discuss":[358],"AI-generated":[359],"health":[360,388],"address":[365],"potential":[366],"oversimplifications":[367],"ensure":[369],"online":[370,400],"advice":[371],"is":[372,465],"integrated":[373],"into":[374],"professional,":[376],"evidence-based":[377],"care":[378],"plan.":[379],"As":[380],"rely":[383],"initial":[387],"guidance,":[389],"practitioners":[390],"must":[391],"act":[392],"authoritative":[395],"\\"filter\\"":[396],"verify":[398],"that":[399,433],"remains":[402],"clinically":[403],"safe,":[404],"aligned":[407],"patient's":[410],"specific":[411],"goals.":[413],"PLAIN":[414],"LANGUAGE":[415],"SUMMARY:":[416],"Many":[417],"people":[418],"now":[419],"turn":[420],"artificial":[422],"intelligence":[423],"help":[425],"found":[432],"these":[434],"vary":[436],"how":[438],"well":[439],"answer":[441],"about":[443],"gum":[444],"disease.":[445],"(such":[451],"ChatGPT-4.5)":[453],"high-quality":[456],"there":[458],"main":[461],"limitations:":[462],"first,":[463],"Persian;":[472],"second,":[473],"many":[474],"explanations":[475],"too":[477],"technical":[478],"hard":[480],"average":[483],"person":[484],"understand.":[486],"Overall,":[487],"can":[490],"be":[491],"helpful":[493],"guide,":[494],"it":[495],"never":[497],"consultation":[501],"dentist.":[504]}

Authors

Institutions

Publication Details

Journal
Clinical Advances in Periodontics
Published
2026-09-16
DOI
https://doi.org/10.1002/cap.70105
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Assessing the accuracy and usability of artificial intelligence–based language models in responding to common periodontal patient questions

Anahita Moscowchi, Reza Amid, Fazele Atarbashi‐Moghadam, Alireza Akbarzadeh Baghban et al.
Clinical Advances in Periodontics
Artificial Intelligence in Healthcare and Education
article

Assessing the accuracy and usability of artificial intelligence–based language models in responding to common periodontal patient questions

Anahita Moscowchi, Reza Amid, Fazele Atarbashi‐Moghadam, Alireza Akbarzadeh Baghban, Sajad Jahantigh
article en

Abstract

BACKGROUND: Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aimed to comparatively evaluate the accuracy and usability of widely used LLMs in responding to frequently asked periodontal questions. METHODS: In this analytical-comparative study, 15 commonly asked periodontal questions were selected from real patient encounters and administered to ten LLMs in both English and Persian. Responses were evaluated independently and blindly by two board-certified periodontists; disagreements were resolved by discussion or third-reviewer. Each response was rated on a 5-point Likert scale across six criteria: correlation with the question, adequacy, comprehensiveness, clarity/readability, usability for individuals with limited scientific literacy, and scientific accuracy. Statistical analyses included descriptive statistics, Mann-Whitney U test, Kruskal-Wallis test with Bonferroni correction, and univariate ANOVA. Significance level was set at p < 0.05. RESULTS: = 294.78, p < 0.001). ChatGPT-4.5 with deep search enabled achieved the highest scores across most criteria, while ChatGPT-4o generally underperformed compared with other LLMs. English responses scored significantly higher than Persian (mean 4.25 vs. 3.96; U = 333,058, p < 0.001). Among evaluated criteria, correlation had the highest mean score (4.81), whereas usability for individuals with limited scientific literacy had the lowest score (3.80). LLMs performed significantly better on treatment/prevention questions than on diagnosis/pathogenesis questions. Only ChatGPT-4.5 (deep search) and Grok-3 (deep search) provided citations, with ChatGPT-4.5 referencing more specialized dental sources. CONCLUSIONS: LLMs show variable performance in answering frequent periodontal patients' questions, influenced by model architecture, language, question type, and evaluative domain. Although advanced configurations such as deep-search ChatGPT-4.5 offer highly accurate, comprehensive, and well-referenced responses, significant limitations remain particularly in Persian language output and in usability for patients with limited scientific literacy. LLMs may serve as complementary tools for patient communication, but should not replace professional advice. Further development and language-specific optimization are needed to improve global accessibility and reliability. KEY POINTS: While AI tools provide highly readable and generally accurate periodontal information, they often lack the clinical nuance required for personalized risk assessment and complex treatment planning. Practitioners should proactively discuss AI-generated health information with patients to address potential oversimplifications and ensure online advice is integrated into a professional, evidence-based care plan. As patients increasingly rely on AI for initial health guidance, practitioners must act as the authoritative "filter" to verify that online information remains clinically safe, accurate, and aligned with the patient's specific treatment goals. PLAIN LANGUAGE SUMMARY: Many people now turn to artificial intelligence for help with their dental questions. This study found that these tools vary in how well they answer questions about gum disease. While the most advanced models (such as ChatGPT-4.5) provide accurate, high-quality information, there are two main limitations: first, performance is significantly better in English than in Persian; second, many explanations are too technical and hard for the average person to understand. Overall, while AI can be a helpful guide, it should never replace a professional consultation with a dentist.

Clinical Advances in Periodontics
Research Institute for Endocrine Sciences (IR), Shahid Beheshti University of Medical Sciences (IR)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.