Consistency of ChatGPT responses with the American Psychiatric Association guidelines in adolescent eating disorders: evaluation of quality, reliability, and readability

The increasing use of artificial intelligence systems in healthcare has raised concerns regarding the accuracy and reliability of information generated by large language models. This study aimed to evaluate the concordance of ChatGPT responses related to adolescent eating disorders with the American Psychiatric Association (APA) eating disorders guideline and to assess the quality, reliability, and readability of these responses. A cross-sectional content analysis was conducted using 34 clinical questions derived from the APA 2023 eating disorders guideline. ChatGPT responses were evaluated by nine independent experts, including adolescent medicine physicians, child and adolescent psychiatrists, and dietitians. Response quality and reliability were assessed using the Global Quality Score (GQS) and modified DISCERN (mDISCERN) scales. Inter-rater agreement was analyzed using Fleiss’ kappa coefficients. Readability was evaluated using multiple standard readability indices. Mean GQS scores indicated high quality in diagnosis and assessment (4.27 ± 0.39), nutritional management and medical monitoring (4.25 ± 0.50), and treatment and clinical interventions (4.37 ± 0.33). Corresponding mDISCERN scores indicated reasonable reliability (29.24 ± 4.25, 30.99 ± 3.12, and 29.75 ± 4.66, respectively). No significant differences in GQS or mDISCERN scores were observed between professional groups (all p > 0.05). Inter-rater agreement varied across disciplines (Fleiss’ κ = 0.28–0.56). Readability ranged from high school to college level. ChatGPT responses demonstrated substantial alignment with guideline-based information in adolescent eating disorders. However, variability in reliability and high readability levels suggest that these systems should be used cautiously. Artificial intelligence tools may serve as complementary sources for accessing guideline-based information but should not replace clinical expertise or evidence-based guidelines in clinical decision-making.

Authors

Institutions

Publication Details

Journal
Journal of Eating Disorders
Published
2026-09-18
DOI
https://doi.org/10.1186/s40337-026-01749-w
Primary Topic
Eating Disorders and Behaviors
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Consistency of ChatGPT responses with the American Psychiatric Association guidelines in adolescent eating disorders: evaluation of quality, reliability, and readability

Zehra Aycan, Rahime Duygu Temeltürk, Ayşe Gül Güven, Simay Mirioğlu et al.
Journal of Eating Disorders
Eating Disorders and Behaviors
article

Consistency of ChatGPT responses with the American Psychiatric Association guidelines in adolescent eating disorders: evaluation of quality, reliability, and readability

Zehra Aycan, Rahime Duygu Temeltürk, Ayşe Gül Güven, Simay Mirioğlu, Burak Kurt
article en

Abstract

The increasing use of artificial intelligence systems in healthcare has raised concerns regarding the accuracy and reliability of information generated by large language models. This study aimed to evaluate the concordance of ChatGPT responses related to adolescent eating disorders with the American Psychiatric Association (APA) eating disorders guideline and to assess the quality, reliability, and readability of these responses. A cross-sectional content analysis was conducted using 34 clinical questions derived from the APA 2023 eating disorders guideline. ChatGPT responses were evaluated by nine independent experts, including adolescent medicine physicians, child and adolescent psychiatrists, and dietitians. Response quality and reliability were assessed using the Global Quality Score (GQS) and modified DISCERN (mDISCERN) scales. Inter-rater agreement was analyzed using Fleiss’ kappa coefficients. Readability was evaluated using multiple standard readability indices. Mean GQS scores indicated high quality in diagnosis and assessment (4.27 ± 0.39), nutritional management and medical monitoring (4.25 ± 0.50), and treatment and clinical interventions (4.37 ± 0.33). Corresponding mDISCERN scores indicated reasonable reliability (29.24 ± 4.25, 30.99 ± 3.12, and 29.75 ± 4.66, respectively). No significant differences in GQS or mDISCERN scores were observed between professional groups (all p > 0.05). Inter-rater agreement varied across disciplines (Fleiss’ κ = 0.28–0.56). Readability ranged from high school to college level. ChatGPT responses demonstrated substantial alignment with guideline-based information in adolescent eating disorders. However, variability in reliability and high readability levels suggest that these systems should be used cautiously. Artificial intelligence tools may serve as complementary sources for accessing guideline-based information but should not replace clinical expertise or evidence-based guidelines in clinical decision-making.

Journal of Eating Disorders
Ankara University (TR)
Openalex Percentile: Top 7%
Eating Disorders and Behaviors
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.