Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation

Abstract Background Polypharmacy and multimorbidity make medication review in older adults a high-stakes clinical task. Large language models (LLMs) may support medication review, but whether item-level concordance with explicit prescribing criteria reflects expert-rated reasoning quality and safety prioritisation is uncertain. We compared the quality and safety of outputs generated by three LLMs using a two-stage evaluation framework. Methods In this two-stage comparative vignette-based benchmark evaluation, GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro responded to 20 standardised geriatric pharmacotherapy vignettes representing fictional older adults aged 72–88 years across four clinical domains. Each vignette contained three potentially inappropriate medications and one clinically relevant START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3. Models received an identical master prompt under default end-user settings; memory features were disabled where available and separate sessions were used for each vignette. Two geriatricians, blinded to model identity, independently rated anonymised outputs using a 100-point rubric comprising medication review output quality (0–80) and critical safety-risk prioritisation (0–20). Stage 1 assessed answer-key concordance; exploratory Stage 2 characterised clinically contextualised error patterns. Comparisons used Friedman and Wilcoxon signed-rank tests, Cochran’s Q, and exact McNemar tests with Bonferroni correction. Results Stage 1 concordance was uniformly high, with limited between-model discrimination. Expert-rated total scores differed across models ( p < 0.001; Kendall’s W = 0.700): Gemini 3 Pro scored highest (97.93 ± 2.33; 95% CI 96.83–99.02), followed by Claude Sonnet 4.5 (93.90 ± 3.37) and GPT-5.2 (89.85 ± 5.23); all pairwise comparisons were significant, with large standardised effect sizes ( r = 0.55–0.61). Safety scores showed the same ranking ( p < 0.001; Kendall’s W = 0.861; r = 0.52–0.62). Reliability of averaged total ratings was high (ICC(A,2) = 0.935; 95% CI 0.897–0.957). In the exploratory Stage 2 error analysis, Gemini was less frequently flagged than GPT-5.2 for superficial reasoning, weak emphasis on life-threatening risk, and any flagged error, and less frequently than Claude Sonnet 4.5 for weak emphasis on life-threatening risk and any flagged error. Conclusions Within this vignette-based benchmark, high item-level concordance did not ensure high expert-rated output quality or safety prioritisation. Models also differed in safety prioritisation and in the pattern of flagged errors. These findings support supervised use of LLMs and evaluation approaches that assess reasoning and prioritisation as well as target detection. They should not be interpreted as evidence that one model is clinically superior in real-world practice. Trial registration Clinical trial not applicable.

Authors

Publication Details

Journal
BMC Geriatrics
Published
2026-09-16
DOI
https://doi.org/10.1186/s12877-026-08253-5
Primary Topic
Pharmaceutical Practices and Patient Outcomes
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation

Tugce Emiroglu Gedik, Duygu Ozata, Alper Döventaş, Suna Avcı et al.
BMC Geriatrics
Pharmaceutical Practices and Patient Outcomes
article

Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation

Tugce Emiroglu Gedik, Duygu Ozata, Alper Döventaş, Suna Avcı, Ulev Deniz Erdinçler, Kübra Çıngar Alpay
article en

Abstract

No abstract available for this paper.

BMC Geriatrics
Openalex Percentile: Top 38%
Pharmaceutical Practices and Patient Outcomes
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.