Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-oriented questions about breast cancer and its surgical treatment regarding scientific accuracy, clarity, and unnecessary detail. Methods Forty frequently asked questions were collected from Turkish online sources. Both models were queried under identical conditions on August 13, 2025, and the responses were anonymized for blinded evaluation. Four general surgeons experienced in breast surgery independently assessed each response using a five-point Likert scale across three domains: scientific accuracy, clarity, and unnecessary detail. Results Response lengths were recorded and compared. ChatGPT achieved significantly higher median scores than Gemini in scientific accuracy [4.75 (3.50–5.00) vs. 4.25 (3.50–5.00); p < 0.001], clarity [4.75 (3.50–5.00) vs. 4.25 (3.25–5.00); p = 0.005], and unnecessary detail [5.00 (5.00–5.00) vs. 4.50 (3.50–5.00); p < 0.001]. The overall median score was 4.83 (4.08–5.00) for ChatGPT and 4.33 (3.50–4.92) for Gemini (p < 0.001). Gemini’s responses were significantly longer (244 ± 84 vs. 170 ± 47 words; p < 0.001). Conclusion Both models demonstrated acceptable overall performance. However, ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions. In their current form, these models should not be used for diagnostic or clinical decision-making purposes. AI-generated health information should be used only under expert supervision and for educational or supportive purposes.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-16
DOI
https://doi.org/10.1371/journal.pone.0358325
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Kemal Atahan, Cem Karaalı, Yeliz Yılmaz, Nihan Acar
PLoS ONE
Artificial Intelligence in Healthcare and Education
article

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Kemal Atahan, Cem Karaalı, Yeliz Yılmaz, Nihan Acar
article en

Abstract

Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-oriented questions about breast cancer and its surgical treatment regarding scientific accuracy, clarity, and unnecessary detail. Methods Forty frequently asked questions were collected from Turkish online sources. Both models were queried under identical conditions on August 13, 2025, and the responses were anonymized for blinded evaluation. Four general surgeons experienced in breast surgery independently assessed each response using a five-point Likert scale across three domains: scientific accuracy, clarity, and unnecessary detail. Results Response lengths were recorded and compared. ChatGPT achieved significantly higher median scores than Gemini in scientific accuracy [4.75 (3.50–5.00) vs. 4.25 (3.50–5.00); p < 0.001], clarity [4.75 (3.50–5.00) vs. 4.25 (3.25–5.00); p = 0.005], and unnecessary detail [5.00 (5.00–5.00) vs. 4.50 (3.50–5.00); p < 0.001]. The overall median score was 4.83 (4.08–5.00) for ChatGPT and 4.33 (3.50–4.92) for Gemini (p < 0.001). Gemini’s responses were significantly longer (244 ± 84 vs. 170 ± 47 words; p < 0.001). Conclusion Both models demonstrated acceptable overall performance. However, ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions. In their current form, these models should not be used for diagnostic or clinical decision-making purposes. AI-generated health information should be used only under expert supervision and for educational or supportive purposes.

PLoS ONEVol. 21(9)
Izmir Kâtip Çelebi University (TR)
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.