Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation
Abstract Background Polypharmacy and multimorbidity make medication review in older adults a high-stakes clinical task. Large language models (LLMs) may support medication review, but whether item-level concordance with explicit prescribing criteria reflects expert-rated reasoning quality and safety prioritisation is uncertain. We compared the quality and safety of outputs generated by three LLMs using a two-stage evaluation framework. Methods In this two-stage comparative vignette-based benchmark evaluation, GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro responded to 20 standardised geriatric pharmacotherapy vignettes representing fictional older adults aged 72–88 years across four clinical domains. Each vignette contained three potentially inappropriate medications and one clinically relevant START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3. Models received an identical master prompt under default end-user settings; memory features were disabled where available and separate sessions were used for each vignette. Two geriatricians, blinded to model identity, independently rated anonymised outputs using a 100-point rubric comprising medication review output quality (0–80) and critical safety-risk prioritisation (0–20). Stage 1 assessed answer-key concordance; exploratory Stage 2 characterised clinically contextualised error patterns. Comparisons used Friedman and Wilcoxon signed-rank tests, Cochran’s Q, and exact McNemar tests with Bonferroni correction. Results Stage 1 concordance was uniformly high, with limited between-model discrimination. Expert-rated total scores differed across models ( p < 0.001; Kendall’s W = 0.700): Gemini 3 Pro scored highest (97.93 ± 2.33; 95% CI 96.83–99.02), followed by Claude Sonnet 4.5 (93.90 ± 3.37) and GPT-5.2 (89.85 ± 5.23); all pairwise comparisons were significant, with large standardised effect sizes ( r = 0.55–0.61). Safety scores showed the same ranking ( p < 0.001; Kendall’s W = 0.861; r = 0.52–0.62). Reliability of averaged total ratings was high (ICC(A,2) = 0.935; 95% CI 0.897–0.957). In the exploratory Stage 2 error analysis, Gemini was less frequently flagged than GPT-5.2 for superficial reasoning, weak emphasis on life-threatening risk, and any flagged error, and less frequently than Claude Sonnet 4.5 for weak emphasis on life-threatening risk and any flagged error. Conclusions Within this vignette-based benchmark, high item-level concordance did not ensure high expert-rated output quality or safety prioritisation. Models also differed in safety prioritisation and in the pattern of flagged errors. These findings support supervised use of LLMs and evaluation approaches that assess reasoning and prioritisation as well as target detection. They should not be interpreted as evidence that one model is clinically superior in real-world practice. Trial registration Clinical trial not applicable.
Authors
- Tugce Emiroglu Gedik (ORCID: https://orcid.org/0000-0002-5550-6477)
- Duygu Ozata (ORCID: https://orcid.org/0009-0001-3927-9411)
- Alper Döventaş (ORCID: https://orcid.org/0000-0001-5509-2625)
- Suna Avcı (ORCID: https://orcid.org/0000-0003-4322-5157)
- Ulev Deniz Erdinçler
- Kübra Çıngar Alpay
Publication Details
- Journal
- BMC Geriatrics
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1186/s12877-026-08253-5
- Primary Topic
- Pharmaceutical Practices and Patient Outcomes
- Type
- article
- Field-Weighted Citation Impact
- 0.00