SEER uses skill-evolving image-grounded reasoning to lift worst-case Dice from 79.34 to 95.47 and cut standard deviation to 0.98 in free-text-prompted 3D medical segmentation
Synopsis
The work proposes SEER, a framework that curates the skill-tagged, image-grounded reasoning-trace dataset SEER-Trace (22,330 multimodal instruction instances from 1,811 cases), extracts anatomical evidence and synthesizes an executable task specification at inference, and uses SEER-Loop to distill high-reward reasoning episodes into reusable skills stored in SEER-Bank, thereby improving accuracy and stability in free-text-promptable 3D medical image segmentation, reporting an 81.94% reduction in performance variance and an 18.60% improvement in worst-case Dice under linguistic perturbations.
Fig. 1. Illustration of the proposed SEER framework. SEER makes free-text prompt- able 3D medical segmentation robust by grounding clinical language in image evidence and evolving reusable reasoning skills.
· Page 3Interpretation
SEER converts a free-text clinical request into a structured evidence-rationale-executable-answer reasoning chain before handing execution to a frozen segmentation backbone, achieving semantic alignment prior to voxel-level decoding. Prior methods mainly relied on stronger vision-language fusion or larger vocabularies for robustness, whereas this work explicitly inserts an intermediate reasoning step that aligns ambiguous phrasing with image-derived anatomical evidence. The paper formalizes the triplet and the reasoning equation, and the PENGWIN ablation shows that a fine-tuned VLM with grounded reasoning raises mean Dice from the baseline 92.26 to 95.92 and lowers standard deviation from 7.49 to 3.84.
SEER-Loop distills high-reward reasoning episodes into reusable skills in SEER-Bank, with deduplication and pruning, so reasoning capability self-refines across rounds. Unlike a static reasoning module, skills here are auditable, retrievable, updatable memory units whose utility is estimated from marginal reward gain. The paper provides the bank-update and utility-estimation equations, and the PENGWIN ablation shows that adding SEER-Loop yields mean Dice 97.39, worst-case Dice 95.47, and standard deviation 0.98.
In free-text prompting mode, SEER outperforms compared baselines on both a partial domain-shift dataset and a strictly out-of-distribution dataset, and markedly reduces sensitivity to wording changes. The paper reports that several prior methods collapse to near-zero Dice under free-text prompting, while SEER simultaneously improves mean, worst-case, and variance on BrainMetShare and PENGWIN. Table 1 shows SEER on PENGWIN with free-text Dice 97.39, worst-case Dice 95.47, and standard deviation 0.98, versus VoxTell at 92.26, 79.34, and 7.49; on BrainMetShare SEER reports 53.83, 51.44, and 1.67.
SEER's front-end reasoning transfers to other segmentation backbones, improving zero-shot generalization and worst-case behavior when the backbone is replaced with MedSAM3. This indicates the gains are not tied to a single segmentation network but come from the front-end grounded reasoning and skill evolution. The paper reports that on the strictly out-of-distribution PENGWIN, integrating SEER raises MedSAM3 mean Dice from 5.75 (std 6.40) to 19.97 (std 13.30) and the worst-case minimum from 3.67 to 3.94.
Perspective
The work targets 3D medical imaging scenarios where segmentation targets are described in free text, and is aimed at researchers and system developers who want clinical natural language rather than fixed labels to drive segmentation; by design it produces an evidence-aligned, executable task specification before a frozen segmentation backbone, so it can be layered onto existing segmentation networks. SEER-Trace and the skill bank provide a supervision format and memory mechanism for reproducible research, facilitating later expansion of the skill taxonomy and evaluation on more clinical tasks.
Readers should still watch how the skill bank behaves over longer evolution rounds and larger skill taxonomies, how it performs on more anatomical targets and institutional sources, and the failure-mode audit the paper mentions; in addition, SEER-Trace requests are generated by a large model under constrained prompting and checked by 5% expert sampling, so the gap between its covered clinical styles and real-world distributions remains an open question. The paper also notes future work on volumetric VLM encoders and broader clinical tasks.
