Multidimensional Evaluation of Guideline-Based Generative Artificial Intelligence Responses in Traditional Chinese: A Ménière’s Disease Study

Background: The reliability of generative artificial intelligence for Chinese medical information remains uncertain. This study evaluates the concordance and linguistic performance of three generative artificial intelligence systems in generating Chinese information on Ménière’s disease against clinical practice guidelines. Methods: Seventeen questions, adapted from the Key Action Statements of the American Academy of Otolaryngology–Head and Neck Surgery guidelines, were posed to ChatGPT o4-mini-high, Gemini 2.5 Pro, and Grok 3 (51 total responses). Responses were assessed for guideline concordance, communication features, and readability with matched analyses (Cochran’s Q and Friedman tests), with Holm–Bonferroni correction across nine communication characteristics. Results: Correctness rates did not differ significantly among the three models (ChatGPT o4-mini-high: 100%, Gemini 2.5 Pro: 100%, Grok 3: 94.1%; Q = 2.00, p = 0.368). Six of the nine communication characteristics differed significantly, with moderate to large effect sizes (Kendall’s W = 0.26–0.76), including guideline quotation, citation quality, key point emphasis and recommendations beyond the guideline. The proportion of difficult words also differed significantly (p = 0.0033). Gemini 2.5 Pro had a lower proportion of difficult words than ChatGPT o4-mini-high (adjusted p = 0.040) and Grok 3 (adjusted p = 0.0002). Conclusion: High guideline concordance does not necessarily indicate reliable citations, effective communication, or accessible language. These findings reflect responses generated under single-query conditions rather than consistent model performance, highlighting the need for expert oversight in clinical use.

Authors

Institutions

Publication Details

Journal
Life
Published
2026-09-30
DOI
https://doi.org/10.3390/life16101649
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multidimensional Evaluation of Guideline-Based Generative Artificial Intelligence Responses in Traditional Chinese: A Ménière’s Disease Study

Chin‐Kuo Chen, Li‐Chun Hsieh, Yun-Chiao Wen, Mien-Jen Lin
Life
Artificial Intelligence in Healthcare and Education
article

Multidimensional Evaluation of Guideline-Based Generative Artificial Intelligence Responses in Traditional Chinese: A Ménière’s Disease Study

Chin‐Kuo Chen, Li‐Chun Hsieh, Yun-Chiao Wen, Mien-Jen Lin
article en

Abstract

Background: The reliability of generative artificial intelligence for Chinese medical information remains uncertain. This study evaluates the concordance and linguistic performance of three generative artificial intelligence systems in generating Chinese information on Ménière’s disease against clinical practice guidelines. Methods: Seventeen questions, adapted from the Key Action Statements of the American Academy of Otolaryngology–Head and Neck Surgery guidelines, were posed to ChatGPT o4-mini-high, Gemini 2.5 Pro, and Grok 3 (51 total responses). Responses were assessed for guideline concordance, communication features, and readability with matched analyses (Cochran’s Q and Friedman tests), with Holm–Bonferroni correction across nine communication characteristics. Results: Correctness rates did not differ significantly among the three models (ChatGPT o4-mini-high: 100%, Gemini 2.5 Pro: 100%, Grok 3: 94.1%; Q = 2.00, p = 0.368). Six of the nine communication characteristics differed significantly, with moderate to large effect sizes (Kendall’s W = 0.26–0.76), including guideline quotation, citation quality, key point emphasis and recommendations beyond the guideline. The proportion of difficult words also differed significantly (p = 0.0033). Gemini 2.5 Pro had a lower proportion of difficult words than ChatGPT o4-mini-high (adjusted p = 0.040) and Grok 3 (adjusted p = 0.0002). Conclusion: High guideline concordance does not necessarily indicate reliable citations, effective communication, or accessible language. These findings reflect responses generated under single-query conditions rather than consistent model performance, highlighting the need for expert oversight in clinical use.

LifeVol. 16(10)
Chang Gung University (TW), Mackay Memorial Hospital (TW), Chang Gung Memorial Hospital (TW), Mackay Medical University (TW), Keelung Chang Gung Memorial Hospital (TW), Linkou Chang Gung Memorial Hospital (TW)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.