Public articles linked to the same research event.
medRxiv This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.