Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

{"Background:":[0],"Although":[1,116],"artificial":[2],"intelligence":[3],"(AI)":[4],"has":[5],"been":[6],"used":[7],"in":[8,157,243],"patient":[9,190],"education":[10,191],"for":[11],"some":[12],"time,":[13],"the":[14,37,49,61,129,134,140,147,153,168,171,205,210,214,220,233],"accuracy,":[15],"reliability,":[16,38],"and":[17,27,40,58,82,100,110,202,219,239],"clinical":[18,69,131,230,237],"appropriateness":[19],"of":[20,42,68,96,207,209,229,236],"AI-generated":[21],"medical":[22],"content":[23],"remain":[24],"inadequately":[25],"defined":[26],"continue":[28],"to":[29,35,48,55,59,79,164],"be":[30],"debated.":[31],"This":[32],"study":[33],"aimed":[34],"evaluate":[36],"readability,":[39],"comprehensibility":[41,101],"responses":[43,62,83,156,169,194,211],"generated":[44],"by":[45,63,89,192],"ChatGPT":[46,80,181],"5.2":[47,182],"most":[50],"frequently":[51],"asked":[52,78],"questions":[53,73],"related":[54],"sleeve":[56,75,196],"gastrectomy,":[57],"assess":[60],"surgeons":[64,92],"with":[65,93,128,226],"varying":[66,94],"levels":[67,95,228],"expertise.":[70,97],"Methods:":[71],"Twenty-four":[72],"regarding":[74,195],"gastrectomy":[76,197],"were":[77,84,160,175],"5.2,":[81],"evaluated":[85],"using":[86,106],"a":[87,184],"Likert-Scale":[88],"three":[90],"general":[91],"The":[98,126],"readability":[99,154,217],"analysis":[102],"was":[103,120,124],"also":[104],"conducted,":[105],"Flesch–Kincaid":[107],"Grade":[108],"Level":[109],"Flesch":[111],"Reading":[112],"Ease":[113],"Scores.":[114],"Results:":[115],"excellent":[117],"intra-rater":[118],"reliability":[119,123],"observed,":[121],"inter-rater":[122],"poor.":[125],"evaluator":[127,141],"greatest":[130],"expertise":[132,238],"assigned":[133,146],"highest":[135],"mean":[136,149],"score":[137,150],"(2.92±0.776),":[138],"whereas":[139,166],"possessing":[142],"predominantly":[143],"theoretical":[144],"knowledge":[145],"lowest":[148],"(2.58±0.881).":[151],"In":[152],"analysis,":[155],"all":[158],"subcategories":[159],"classified":[161],"as":[162,177,212],"“difficult":[163],"read,”":[165],"only":[167],"within":[170],"postoperative":[172],"course":[173],"category":[174],"categorized":[176],"“fairly":[178],"difficult”.":[179],"Conclusions:":[180],"is":[183],"valuable":[185],"AI–assisted":[186],"chatbot":[187],"that":[188,198],"facilitates":[189],"providing":[193],"are":[199],"generally":[200],"accurate":[201],"acceptable.":[203],"Nevertheless,":[204],"categorization":[206],"4–16%":[208],"\\"Incorrect,\\"":[213],"overall":[215],"difficult":[216],"levels,":[218],"significant":[221],"variability":[222],"observed":[223],"among":[224],"evaluators":[225],"different":[227],"experience":[231],"underscore":[232],"indispensable":[234],"role":[235],"specialized":[240],"professional":[241],"guidance":[242],"surgical":[244],"practice.":[245]}

Authors

Institutions

Publication Details

Journal
Archives of Current Medical Research
Published
2026-09-16
DOI
https://doi.org/10.47482/acmr.1958370
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

Furkan Turkoglu, Elif Nur GENCER, Emre Erdogan
Archives of Current Medical Research
Artificial Intelligence in Healthcare and Education
article

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

Furkan Turkoglu, Elif Nur GENCER, Emre Erdogan
article en

Abstract

Background: Although artificial intelligence (AI) has been used in patient education for some time, the accuracy, reliability, and clinical appropriateness of AI-generated medical content remain inadequately defined and continue to be debated. This study aimed to evaluate the reliability, readability, and comprehensibility of responses generated by ChatGPT 5.2 to the most frequently asked questions related to sleeve gastrectomy, and to assess the responses by surgeons with varying levels of clinical expertise. Methods: Twenty-four questions regarding sleeve gastrectomy were asked to ChatGPT 5.2, and responses were evaluated using a Likert-Scale by three general surgeons with varying levels of expertise. The readability and comprehensibility analysis was also conducted, using Flesch–Kincaid Grade Level and Flesch Reading Ease Scores. Results: Although excellent intra-rater reliability was observed, inter-rater reliability was poor. The evaluator with the greatest clinical expertise assigned the highest mean score (2.92±0.776), whereas the evaluator possessing predominantly theoretical knowledge assigned the lowest mean score (2.58±0.881). In the readability analysis, responses in all subcategories were classified as “difficult to read,” whereas only the responses within the postoperative course category were categorized as “fairly difficult”. Conclusions: ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable. Nevertheless, the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability observed among evaluators with different levels of clinical experience underscore the indispensable role of clinical expertise and specialized professional guidance in surgical practice.

Archives of Current Medical ResearchVol. 7(3)
State Hospital (GB), Akita International University (JP)
Quality Education
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.