Quality, safety, and readability of consumer-facing generative AI responses to end-of-life questions relevant to surrogate decision-makers for older adults: a benchmark evaluation

As the population ages and the prevalence of multiple coexisting conditions increases, end-of-life decision-making for older adults is becoming increasingly complex in clinical practice. limited access to relevant medical information and differences in medical knowledge between clinicians and surrogate decision-makers. Large language models (LLMs) are being widely used for accessing health information and facilitating doctor-patient communication, but their value and risks in end-of-life decision-making support for older adults have yet to be systematically evaluated. This study aimed to benchmark the safety, accuracy, observable empathetic communication, information quality, transparency indicators, and readability of responses generated by five publicly accessible consumer-facing generative AI systems to standardized end-of-life questions relevant to surrogate decision-makers for older adults. This study is a cross-sectional benchmark evaluation using a 30-question standardized question set covering six prespecified end-of-life domains. The question set was developed using an evidence-informed approach incorporating clinical guidelines, consensus statements, relevant literature, published qualitative studies involving surrogate decision-makers and family caregivers, and public search terminology. Each question was submitted once, in a new session, to five consumer-facing generative AI systems under their default web-interface conditions on June 14, 2026, yielding 150 paired system–question outputs. Two blinded clinical experts independently evaluated the outputs using predefined criteria for safety, accuracy, textual empathy, DISCERN, EQIP, JAMA transparency benchmarks, GQS, and six readability indices. Inter-rater reliability, paired overall comparisons, effect sizes, and post hoc comparisons with Benjamini–Hochberg correction were assessed. Safe-response rates ranged from 76.7% to 86.7% across systems, with no evidence of an overall between-system difference in safety (Cochran’s Q = 1.368, df = 4, P = 0.850; effect size = 0.011, 95% CI 0.005–0.117). Accuracy differed significantly across systems, although the overall effect was small ( P = 0.003; Kendall’s W = 0.132, 95% CI 0.036–0.332). Differences in expert-rated textual empathy were more pronounced ( P < 0.001; Kendall’s W = 0.373, 95% CI 0.194–0.598). Significant between-system differences were also observed for DISCERN ( P < 0.001; W = 0.332), EQIP ( P = 0.016; W = 0.102), JAMA transparency benchmarks ( P < 0.001; W = 0.751), and GQS ( P = 0.001; W = 0.148). Readability differed significantly across systems across all six indices (all P < 0.001), with substantial variability in grade-level estimates. Inter-rater agreement was good to excellent across the evaluated dimensions. In this time-stamped benchmark, consumer-facing generative AI systems generally produced accurate responses with observable empathetic communication, but clinically relevant limitations remained in safety, transparency, and readability. The sampled systems did not differ significantly in safety, whereas differences in accuracy were small and differences in textual empathy were more pronounced. Because the study evaluated generated text rather than surrogate understanding, trust, preferences, decisions, decisional conflict, behavior, or clinical outcomes, the findings do not establish the effectiveness of these systems as end-of-life decision-support interventions. Consumer-facing generative AI should therefore not be used as a standalone information source for high-stakes end-of-life decisions.

Authors

Institutions

Publication Details

Journal
BMC Palliative Care
Published
2026-10-05
DOI
https://doi.org/10.1186/s12904-026-02360-1
Primary Topic
Palliative Care and End-of-Life Issues
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Quality, safety, and readability of consumer-facing generative AI responses to end-of-life questions relevant to surrogate decision-makers for older adults: a benchmark evaluation

Feng Cao, Jiayan Ling, Zhenliang Zhu, Xiaoli Zhou et al.
BMC Palliative Care
Palliative Care and End-of-Life Issues
article

Quality, safety, and readability of consumer-facing generative AI responses to end-of-life questions relevant to surrogate decision-makers for older adults: a benchmark evaluation

Feng Cao, Jiayan Ling, Zhenliang Zhu, Xiaoli Zhou, Zichen Liu
article en

Abstract

As the population ages and the prevalence of multiple coexisting conditions increases, end-of-life decision-making for older adults is becoming increasingly complex in clinical practice. limited access to relevant medical information and differences in medical knowledge between clinicians and surrogate decision-makers. Large language models (LLMs) are being widely used for accessing health information and facilitating doctor-patient communication, but their value and risks in end-of-life decision-making support for older adults have yet to be systematically evaluated. This study aimed to benchmark the safety, accuracy, observable empathetic communication, information quality, transparency indicators, and readability of responses generated by five publicly accessible consumer-facing generative AI systems to standardized end-of-life questions relevant to surrogate decision-makers for older adults. This study is a cross-sectional benchmark evaluation using a 30-question standardized question set covering six prespecified end-of-life domains. The question set was developed using an evidence-informed approach incorporating clinical guidelines, consensus statements, relevant literature, published qualitative studies involving surrogate decision-makers and family caregivers, and public search terminology. Each question was submitted once, in a new session, to five consumer-facing generative AI systems under their default web-interface conditions on June 14, 2026, yielding 150 paired system–question outputs. Two blinded clinical experts independently evaluated the outputs using predefined criteria for safety, accuracy, textual empathy, DISCERN, EQIP, JAMA transparency benchmarks, GQS, and six readability indices. Inter-rater reliability, paired overall comparisons, effect sizes, and post hoc comparisons with Benjamini–Hochberg correction were assessed. Safe-response rates ranged from 76.7% to 86.7% across systems, with no evidence of an overall between-system difference in safety (Cochran’s Q = 1.368, df = 4, P = 0.850; effect size = 0.011, 95% CI 0.005–0.117). Accuracy differed significantly across systems, although the overall effect was small ( P = 0.003; Kendall’s W = 0.132, 95% CI 0.036–0.332). Differences in expert-rated textual empathy were more pronounced ( P < 0.001; Kendall’s W = 0.373, 95% CI 0.194–0.598). Significant between-system differences were also observed for DISCERN ( P < 0.001; W = 0.332), EQIP ( P = 0.016; W = 0.102), JAMA transparency benchmarks ( P < 0.001; W = 0.751), and GQS ( P = 0.001; W = 0.148). Readability differed significantly across systems across all six indices (all P < 0.001), with substantial variability in grade-level estimates. Inter-rater agreement was good to excellent across the evaluated dimensions. In this time-stamped benchmark, consumer-facing generative AI systems generally produced accurate responses with observable empathetic communication, but clinically relevant limitations remained in safety, transparency, and readability. The sampled systems did not differ significantly in safety, whereas differences in accuracy were small and differences in textual empathy were more pronounced. Because the study evaluated generated text rather than surrogate understanding, trust, preferences, decisions, decisional conflict, behavior, or clinical outcomes, the findings do not establish the effectiveness of these systems as end-of-life decision-support interventions. Consumer-facing generative AI should therefore not be used as a standalone information source for high-stakes end-of-life decisions.

BMC Palliative Care
Zhejiang Hospital (CN)
Good health and well-being
Openalex Percentile: Top 10%
Palliative Care and End-of-Life Issues
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.