OpenTumorBoard benchmarks 14 models on 611 real tumor board cases, with the best reaching only 2.78/5 alignment with actual conclusions
Synopsis
Researchers curated OpenTumorBoard from 219 recorded tumor board meetings totaling 12,534 minutes of YouTube video, yielding 611 patient cases, 19,157 discussion turns and ten specialist roles, and used it for two tasks—single Specialist Turn responses and full Board Simulation—finding that across 14 general-purpose and medical LLMs the best model reached only 3.43/5 in clinical equivalence and 2.78/5 in alignment with true board conclusions, while supervised finetuning plus reinforcement learning on Qwen2.5-VL-3B raised conclusion alignment from 1.58 to 1.86.
Interpretation
The paper builds OpenTumorBoard, a benchmark grounded in real tumor board recordings, containing 611 patient cases, 19,157 discussion turns and ten specialist roles drawn from 219 meetings totaling 12,534 minutes, spanning 17 cancer sites, 9 specialist turn types and 11 input modalities. Table 1 contrasts existing benchmarks (MedQA, CRAFT-MD, AgentClinic, MTBBench and others) as lacking real tumor board cases, a multidisciplinary team, ground-truth discussion or multimodal data, and the paper states OpenTumorBoard is the first benchmark providing complete tumor board discussion trajectories paired with observed clinical consensus; it is substantially larger than MTBBench's 66 patient cases. Scale and structure come from the main text and appendix; three M.D. experts with oncology expertise reviewed a random subset (15 cases, each seen by two reviewers, with a third expert resolving disagreement), rating case coverage, case factuality and consensus fidelity 100% "High", response correctness 92.6% "High" and discussion support 94.4% "High".
The paper derives two evaluation tasks: Specialist Turn, where a model plays one specialist answering a question actually posed in the meeting (4,844 test items across 9 categories), and Board Simulation, where only the case summary and slides are given and the model must generate a back-and-forth discussion and reach consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. The paper stresses that references are not written retrospectively from case descriptions but grounded in observed specialist responses and board-level consensus from the recordings, and describes this as the first capture of the real-world questions a tumor board actually poses. Task definitions, test-set scale (66 recordings, 184 cases, 4,844 questions) and the evaluation protocol appear in Section 2.4; judging uses Qwen3.8-27B as an LLM judge with a clinical-equivalence rubric and a conclusion-alignment rubric, alongside ROUGE-L and BERTScore.
Evaluation exposes clear gaps: on Specialist Turn the best model, DeepSeek-V4-Pro in reasoning mode, reaches only 3.43/5 clinical equivalence, with critical-error rates from 5.6% to 43.5% and unsupported-claim rates from 10.0% to 74.5%; on Board Simulation the best model, Gemini 3.7 Flash, reaches only 2.78/5 conclusion alignment. The paper reports that medical-specific models did not outperform general-purpose ones, with three of four medical models ranking at the bottom and HuatuoGPT-3-32B and MedReason-8B showing the highest unsupported-claim rates (74.5% and 51.9%), suggesting specialist questions require case-specific reasoning rather than generic medical knowledge. Results come from evaluations of 9 models (Specialist Turn) and 14 models (Board Simulation), with per-item scores in Appendix Tables A2 to A5; the paper also reports that longer discussion trajectories are strongly associated with higher conclusion alignment.
The paper also treats the benchmark as a training ground: supervised finetuning of Qwen2.5-VL-3B on 366 training cases raised conclusion alignment from 1.58 to 1.67, and reinforcement learning with Dr.GRPO raised it further to 1.86, which the paper describes as roughly an 18% relative improvement and as closing half the gap to MedGemma 27B. The paper presents this as evidence that real discussion trajectories can serve directly as supervision, teaching a small model to generate complementary specialist perspectives and orchestrate them toward a board-level decision, with RL requiring no additional labeled examples. Training settings (50 finetuning epochs, learning rate, batch size, 128,000-token context, fixed seed; RL with 4 cases and 16 samples per case per step, checkpoint at step 430) are given in Appendix E, the improvement over the finetuned model is confirmed by a paired test, and per-dimension scores appear in Table A4.
Perspective
This work targets researchers and developers who want to evaluate or train models for multidisciplinary cancer decision-making: the benchmark supplies case summaries, slides, discussion trajectories and board consensus, supports both single-turn specialist response and end-to-end meeting simulation, and can serve as supervision for finetuning and reinforcement learning. The paper states it will release OpenTumorBoard and its automated curation pipeline so the benchmark can expand with newly posted recordings, and supports local extension to institutional data subject to applicable privacy and governance requirements. The results apply to the tumor board settings reflected in public recordings, where discussions frequently concern advanced and metastatic disease.
The paper explicitly lists agreement between the LLM judge and clinician assessments as important future work, so current scores reflect rubric-based automated judging. Data come only from public YouTube recordings, which the paper notes may limit representativeness, and the paper states that no independent patient-identifier screening was performed beyond the privacy protections in the publicly released source recordings. The curation pipeline relies on GPT-5.4 for role inference, segmentation, evidence and conclusion extraction, with quality supported by manual inspection and expert review at limited scale (for example, 15 recordings and 15 cases). In addition, although this evidence bundle is full text, some appendices (such as the verbatim judge prompts and rubrics and the full reward formula) are referenced rather than reproduced, so reproducing exact scoring details would require the original appendices.
