When Treatment Gets More Complicated: AI Readability Declines in Breast Cancer Patient Education

Background The proliferation of Large Language Models necessitates evaluating their performance in communication accessibility. This study compares the linguistic architecture, readability, and automated psycholinguistic text properties of AI-generated breast cancer materials across varying clinical complexities. Methods Twenty standardized questions across diagnosis and treatment categories were queried through three independent iterations. Readability was analyzed using Flesch-Kincaid Grade Level, SMOG, and Gunning Fog indices. Linguistic dimensions were evaluated via LIWC-22 software as computational style metrics. Statistical analyses utilized Wilcoxon, Mann-Whitney U, Spearman’s correlation, and with False Discovery Rate (FDR) corrections. Cross-iteration stability was assessed with the intraclass correlation coefficient (ICC(2,1), two-way random-effects, and absolute agreement). Results Gemini had better readability than ChatGPT (FKGL 9.23 vs 10.07, q = 0.022); both exceeded the 6th-grade threshold. Treatment queries increased linguistic complexity for both (ChatGPT FKGL q = 0.012; Gemini FKGL q = 0.047). Gemini scored higher on Analytic (q = 0.001) and Authentic (q = 0.021) dimensions; ChatGPT scored higher on I-words (q = 0.001) and Cognitive Processes (q < 0.001). Treatment-specific shifts in Gemini Positive Tone (raw P = .041, q = 0.367) and ChatGPT Authentic (raw P = .025, q = 0.114) lost significance after FDR adjustment. ChatGPT showed a negative correlation between readability and Authentic score (ρ = −0.63, q = 0.026). Cross-iteration ICC values were poor-to-moderate for both models (mean ICC: ChatGPT = 0.58, Gemini = 0.59). Conclusions Both LLMs create practical information gaps for patients with limited health literacy, particularly when addressing complex treatment pathways. Finer-grained psycholinguistic claims did not withstand statistical correction, and cross-iteration stability was modest. These automated analyses highlight structural limitations of current LLMs, reinforcing the necessity of clinician-guided curation before deploying them for patient education.

Authors

Publication Details

Journal
The American Surgeon
Published
2026-09-29
DOI
https://doi.org/10.1177/00031348261494138
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

When Treatment Gets More Complicated: AI Readability Declines in Breast Cancer Patient Education

Burak Altunpak
The American Surgeon
Artificial Intelligence in Healthcare and Education
article

When Treatment Gets More Complicated: AI Readability Declines in Breast Cancer Patient Education

Burak Altunpak
article en

Abstract

Background The proliferation of Large Language Models necessitates evaluating their performance in communication accessibility. This study compares the linguistic architecture, readability, and automated psycholinguistic text properties of AI-generated breast cancer materials across varying clinical complexities. Methods Twenty standardized questions across diagnosis and treatment categories were queried through three independent iterations. Readability was analyzed using Flesch-Kincaid Grade Level, SMOG, and Gunning Fog indices. Linguistic dimensions were evaluated via LIWC-22 software as computational style metrics. Statistical analyses utilized Wilcoxon, Mann-Whitney U, Spearman’s correlation, and with False Discovery Rate (FDR) corrections. Cross-iteration stability was assessed with the intraclass correlation coefficient (ICC(2,1), two-way random-effects, and absolute agreement). Results Gemini had better readability than ChatGPT (FKGL 9.23 vs 10.07, q = 0.022); both exceeded the 6th-grade threshold. Treatment queries increased linguistic complexity for both (ChatGPT FKGL q = 0.012; Gemini FKGL q = 0.047). Gemini scored higher on Analytic (q = 0.001) and Authentic (q = 0.021) dimensions; ChatGPT scored higher on I-words (q = 0.001) and Cognitive Processes (q < 0.001). Treatment-specific shifts in Gemini Positive Tone (raw P = .041, q = 0.367) and ChatGPT Authentic (raw P = .025, q = 0.114) lost significance after FDR adjustment. ChatGPT showed a negative correlation between readability and Authentic score (ρ = −0.63, q = 0.026). Cross-iteration ICC values were poor-to-moderate for both models (mean ICC: ChatGPT = 0.58, Gemini = 0.59). Conclusions Both LLMs create practical information gaps for patients with limited health literacy, particularly when addressing complex treatment pathways. Finer-grained psycholinguistic claims did not withstand statistical correction, and cross-iteration stability was modest. These automated analyses highlight structural limitations of current LLMs, reinforcing the necessity of clinician-guided curation before deploying them for patient education.

The American Surgeon
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

When Treatment Gets More Complicated: AI Readability Declines in Breast Cancer Patient Education — Burak Altunpak · The American Surgeon (2026) | TGRS Research Map | TGRS