Quality and Safety of Large Language Model–Generated Medication Review Outputs in Geriatric Pharmacotherapy: A Two-Stage Comparative Vignette-Based Benchmark Evaluation
Synopsis
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Interpretation
The work proposes and implements a two-stage evaluation framework: Stage 1 assesses item-level concordance of model outputs with explicit prescribing-criteria answer keys, and Stage 2 exploratorily characterises clinically contextualised error patterns. Prior evaluation of LLM medication review largely rested on item matching against explicit prescribing criteria; this work separates 'getting the items right' from expert-rated reasoning quality and safety prioritisation, making it possible to observe that the two do not coincide. Based on 20 standardised vignettes, an identical master prompt, default end-user settings, memory features disabled where available, and separate sessions per vignette, with two blinded geriatricians independently rating on a 100-point rubric and inter-rater reliability ICC(A,2) = 0.935 (95% CI 0.897–0.957).
Item-level concordance was uniformly high with limited between-model discrimination, yet expert-rated total output-quality scores differed significantly across the three models. This indicates that concordance with answer keys alone cannot discriminate models, whereas expert ratings can reveal differences, suggesting the evaluation dimensions need to be broadened. Total-score differences p < 0.001, Kendall's W = 0.700, all pairwise comparisons significant with large standardised effect sizes (r = 0.55–0.61); Gemini 3 Pro 97.93 ± 2.33 (95% CI 96.83–99.02), Claude Sonnet 4.5 93.90 ± 3.37, GPT-5.2 89.85 ± 5.23.
Safety-prioritisation scores showed the same model ranking as total scores, and models differed in safety prioritisation. Treating 'critical safety-risk prioritisation' as a scoring dimension independent of general output quality, and showing it discriminates models. Safety-score differences p < 0.001, Kendall's W = 0.861, r = 0.52–0.62, with ranking matching total scores.
Exploratory error analysis showed that the pattern of flagged errors differed across models. Suggests that, beyond score levels, the distribution of error types is an observable dimension of between-model difference. Gemini was less frequently flagged than GPT-5.2 for superficial reasoning, weak emphasis on life-threatening risk, and any flagged error, and less frequently than Claude Sonnet 4.5 for weak emphasis on life-threatening risk and any flagged error; the authors position this analysis as exploratory.
Perspective
The results apply to the vignette-based benchmark setting: fictional older adults aged 72–88, four clinical domains, each vignette containing three potentially inappropriate medications and one START-type omission, anchored to the AGS Beers Criteria and STOPP/START version 3, with an identical master prompt and default end-user settings. They support supervised use of LLMs to assist medication review and support including reasoning quality and safety prioritisation in evaluation design; they are directly useful to geriatrics, clinical pharmacy, and clinical informatics teams building clinical LLM evaluation pipelines.
The authors position the Stage 2 error analysis as exploratory, so its error-type distribution findings are best treated as observations awaiting further verification; the vignettes are fictional and the set comprises 20 cases, so whether the model ranking is stable across different vignette sets, prompts, or clinical domains remains an open question; in addition, the loaded content is abstract-level text lacking figures and full methodological detail, so the weighting of rubric items, the specific rules for flagging errors, and the vignette construction process cannot be verified from the available text, which are open points for readers to keep in mind.
