Do Large Language Models Use the Clinical Vignette? A Question Ablation Study on the Orthopaedic In-Training Examination
Synopsis
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
Interpretation
The clinical vignette was the primary contributor to LLM performance on the OITE, whereas images had minimal effect. Prior work often measured model ability by overall examination accuracy; this study removed question components one at a time, decomposing the total score into the contributions of the vignette, images, and answer options. Based on 792 OITE questions, eight models, five image-containing and three non-image-containing ablation conditions, with accuracy summarized by median and IQR and compared using paired analyses with Holm adjustment; accuracy on image-containing questions was essentially unchanged when images were removed (59.13% vs 59.13%) but dropped when the vignette was removed (to 49.05%).
When given only the answer options, without the question stem, clinical context, or images, every model still scored above the 25% random baseline. This condition compresses the question to options alone, directly testing whether models can answer without clinical content, rather than reporting only full-information scores. The highest accuracy was 45.62% (44.24-47.24) for Claude Opus 4.8 on image-containing questions and 46.09% (44.41-47.77) for GPT-5.6 Terra on non-image questions, with every model exceeding the 25% random expectation; the authors suggest models may be learning spurious correlations.
BioMedBERT semantic similarity remained high across models but was not significantly associated with accuracy and did not consistently distinguish correct from incorrect responses. The study examined explanation-text semantic similarity alongside answer accuracy, indicating that explanation fluency and answer correctness may not move together. BioMedBERT was used to assess semantic similarity between model-generated explanations and reference discussions; the Spearman correlation was ρ=-0.17, P=0.29, not significant, while similarity stayed high despite wide accuracy differences across models.
Proprietary models achieved substantially higher accuracy than open-source models, yet both performed above chance when given answer options alone. The study placed three open-source Ministral-3 models and five proprietary models within the same ablation framework, grounding cross-model comparison in identical questions and conditions. GPT-5.6 Terra achieved the highest full-context accuracy on both image-containing (81.80%, 80.41-82.95) and non-image-containing (94.13%, 93.30-94.97) questions, while all models remained above chance with options alone.
Perspective
The study defines the scope of its conclusions: the evaluation covers 792 OITE questions from 2020 through 2024 (434 with images, 358 without) and three open-source Ministral-3 models plus five proprietary models, under five image-containing and three non-image-containing ablation conditions. Within this setting, the results help clarify the relative contributions of question components in orthopaedic training examination items and provide a starting point for repeating similar ablations across broader specialties, more question formats, or more model versions.
Readers may still watch: the source of above-chance performance with options alone, whether it comes from cues in the options themselves, memorization of questions in training data, or another mechanism, is not distinguished in the text; BioMedBERT similarity was high but not significantly correlated with accuracy, so the relationship between explanation quality and correctness still needs further characterization; in addition, this reading is at the abstract level and does not include figures or supplementary materials, so if those contain model-level or year-level breakdowns, they could affect a detailed understanding of each component's contribution.
