Quality and Readability of Four AI Chatbots Answering Search-Derived Patient Questions About Primary Aldosteronism: A Cross-Sectional Comparative Study

Background/Objectives: Patients use artificial intelligence (AI) chatbots for medical information, but response quality, source transparency, and readability for primary aldosteronism (PA) remain uncertain. This study compared four chatbot configurations and assessed whether their responses met patient-education reading levels. Methods: During a single-time-point snapshot on 27 July 2026, 17 questions derived from Medical Subject Headings and five-year worldwide Google Trends data were submitted once to GPT-5.5 through ChatGPT Plus, Copilot in Smart mode, Gemini 3.5 Flash, and the default Perplexity model, yielding 68 responses. DISCERN, the Ensuring Quality Information for Patients (EQIP) instrument, the Journal of the American Medical Association (JAMA) benchmarks, and the Global Quality Score (GQS) assessed structural information quality; six formulas assessed readability. These instruments did not assess statement-level factual accuracy. Differences were examined using Friedman tests and paired Wilcoxon signed-rank tests with Holm adjustment. Results: Scores differed for DISCERN (χ2 = 16.717, p < 0.001), EQIP (χ2 = 31.125, p < 0.001), JAMA (χ2 = 48.851, p < 0.001), and GQS (χ2 = 9.874, p = 0.020). Copilot had the highest median DISCERN and EQIP scores (43.00 and 75.00), followed by Perplexity (42.00 and 70.00). Median JAMA scores were 0 for ChatGPT and Gemini and 1 for Copilot and Perplexity. These instrument-based differences do not establish greater factual accuracy, guideline concordance, or clinical superiority; moreover, the full-set DISCERN comparison has limited interpretability because only two questions were explicitly treatment-related. Item-level kappa values ranged from 0.844 to 0.924, and total-score ICC(2,1) values ranged from 0.846 to 0.883. No readability measure met the sixth-grade benchmark. Conclusions: In these configurations, chatbots produced PA information. However, source transparency remained low, and language was overly complex. Relative score differences should not be interpreted as factual superiority. Patient-facing use requires verified sources, scope limits, human oversight, and routes to care.

Authors

Institutions

Publication Details

Journal
Healthcare
Published
2026-09-21
DOI
https://doi.org/10.3390/healthcare14183119
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Quality and Readability of Four AI Chatbots Answering Search-Derived Patient Questions About Primary Aldosteronism: A Cross-Sectional Comparative Study

Jianyu Chen, Chenxia Wang, Jie Qin, Yuxia Ma et al.
Healthcare
Artificial Intelligence in Healthcare and Education
article

Quality and Readability of Four AI Chatbots Answering Search-Derived Patient Questions About Primary Aldosteronism: A Cross-Sectional Comparative Study

Jianyu Chen, Chenxia Wang, Jie Qin, Yuxia Ma, Shangyu Han, Jianxun Cao, Zaihang Zhao
article en

Abstract

Background/Objectives: Patients use artificial intelligence (AI) chatbots for medical information, but response quality, source transparency, and readability for primary aldosteronism (PA) remain uncertain. This study compared four chatbot configurations and assessed whether their responses met patient-education reading levels. Methods: During a single-time-point snapshot on 27 July 2026, 17 questions derived from Medical Subject Headings and five-year worldwide Google Trends data were submitted once to GPT-5.5 through ChatGPT Plus, Copilot in Smart mode, Gemini 3.5 Flash, and the default Perplexity model, yielding 68 responses. DISCERN, the Ensuring Quality Information for Patients (EQIP) instrument, the Journal of the American Medical Association (JAMA) benchmarks, and the Global Quality Score (GQS) assessed structural information quality; six formulas assessed readability. These instruments did not assess statement-level factual accuracy. Differences were examined using Friedman tests and paired Wilcoxon signed-rank tests with Holm adjustment. Results: Scores differed for DISCERN (χ2 = 16.717, p < 0.001), EQIP (χ2 = 31.125, p < 0.001), JAMA (χ2 = 48.851, p < 0.001), and GQS (χ2 = 9.874, p = 0.020). Copilot had the highest median DISCERN and EQIP scores (43.00 and 75.00), followed by Perplexity (42.00 and 70.00). Median JAMA scores were 0 for ChatGPT and Gemini and 1 for Copilot and Perplexity. These instrument-based differences do not establish greater factual accuracy, guideline concordance, or clinical superiority; moreover, the full-set DISCERN comparison has limited interpretability because only two questions were explicitly treatment-related. Item-level kappa values ranged from 0.844 to 0.924, and total-score ICC(2,1) values ranged from 0.846 to 0.883. No readability measure met the sixth-grade benchmark. Conclusions: In these configurations, chatbots produced PA information. However, source transparency remained low, and language was overly complex. Relative score differences should not be interpreted as factual superiority. Patient-facing use requires verified sources, scope limits, human oversight, and routes to care.

HealthcareVol. 14(18)
Gansu Provincial Hospital (CN), Lanzhou University Second Hospital (CN), First Hospital of Lanzhou University (CN), Lanzhou University (CN)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.