Public articles linked to the same research event.
BMC Geriatrics Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.