SWiM distills working memory so multimodal LLMs improve reasoning segmentation even without memory at inference, reaching state-of-the-art on benchmarks
Synopsis
The work finds that multimodal large language models benefit from using their own self-generated reasoning traces and localization proposals as working memory when revisiting the same image and query, and proposes SWiM: rollouts are selected by segmentation quality to build working memory, the memory-conditioned model serves as a teacher giving token-level distributional supervision along student-generated trajectories while the student receives only the original image and query, and on-policy self-distillation is jointly optimized with outcome-based reinforcement learning, achieving state-of-the-art performance on reasoning segmentation benchmarks.
Figure 1: Working-memory-guided reasoning segmentation. Given an image and a query, the MLLM generates reasoning traces and localization proposals that form its working memory. It then revisits the same image and query with this memory as additional input, predicting boxes and points that prompt a frozen SAM2 model to produce the final mask.
arXivInterpretation
The authors observe that MLLMs can use their own previously generated reasoning traces and localization proposals as working memory, yielding improved reasoning segmentation when revisiting the same image and query. Prior reasoning segmentation typically treats the MLLM's reasoning and localization output as a one-shot result, whereas this work feeds the same model's earlier output back as context, explicitly framing the revisit as a usable self-generated memory signal. The finding comes from the authors' exploration, stated as 'Our exploration reveals'; the abstract gives no dataset, sample size, or control setting.
It proposes SWiM (Reasoning Segmenter with Working Memory), a working-memory distillation framework that selects rollouts by segmentation quality to construct working memory, uses the memory-conditioned model as a teacher providing token-level distributional supervision along student-generated trajectories, while the student receives only the original image and query. Unlike approaches relying only on outcome rewards or only on imitating a teacher's trajectory, this uses the model's own memory-conditioned self as teacher so the student inherits the benefit of revisiting without needing working memory at inference. The method description comes from the abstract and is at the framework level; the abstract does not report teacher or student scale, number of rollouts, or selection thresholds.
SWiM jointly optimizes on-policy self-distillation and outcome-based reinforcement learning, combining working-memory guidance with direct feedback on segmentation quality. Distribution-level distillation signals and outcome-level reward signals are combined in one training procedure, so working-memory guidance is tied to final segmentation quality rather than remaining purely imitative. The abstract names the joint optimization components but provides no loss weights, training steps, or ablation results.
Extensive experiments on reasoning segmentation benchmarks show SWiM achieves state-of-the-art performance, validating the effectiveness of working-memory distillation. It turns the benefit of revisiting self-generated memory from an inference-time trick into a trainable capability that can be deployed without memory. The abstract summarizes this as 'Extensive experiments on reasoning segmentation benchmarks' and 'state-of-the-art performance' without naming benchmarks, metric values, or comparison baselines.
Perspective
The work targets reasoning segmentation specifically and applies to training settings where a multimodal large language model can generate reasoning traces and localization proposals and where segmentation quality can be scored to select rollouts. Its goal is for the backbone model to benefit whether or not working memory is provided at inference, making it directly relevant to readers who want to improve segmentation quality without adding inference-time context overhead. The abstract-level information is suitable for judging direction and framework; reproduction or comparison would require the paper's benchmark setup, teacher and student configuration, and training details.
The abstract does not give specific benchmark names, metric values, comparison baselines, or ablations, so the magnitude of 'state-of-the-art' and the contribution of each component remain unclear. Constructing working memory depends on selecting rollouts by segmentation quality, and the selection criteria and quality assessment are not detailed in the abstract, which may affect stability across data distributions. The teacher is a memory-conditioned model while the student sees only the original image and query, so the relationship between the capability gap and distillation signal strength is an open question worth watching. The abstract also does not state how performance differs between inference with no working memory and inference with working memory.
