Skip to main content
Back to timeline
arXivSource publication:

SpatialOPSD distills verified spatial coding-agent traces into a 9B multimodal model, reaching 51.9 average accuracy with tool-free inference

Synopsis

The authors introduce SpatialOPSD, an on-policy self-distillation framework that treats verified execution traces from a spatial coding agent as privileged information, letting a frozen teacher initialized from the same checkpoint see the traces while the student sees only raw images, with Repetition-Aware Distillation (repetition masking plus unlikelihood regularization) suppressing verbatim copying; on MindCube-Tiny and ViewSpatial, the 9B student raises average accuracy from 45.9 to 51.9, above SFT at 47.6 and GRPO at 48.2, and retains about 91% of its tool-using teacher SpatialClaw at 57.3 in a single tool-free forward pass.

Source-provided article image: SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
Figure 1 ·

Figure 1 : MindCube-Tiny test accuracy . OPSD from spatial agents effectively boosts spatial reasoning.

arXiv

Interpretation

Conditioning a multimodal model on execution traces of a spatial coding agent alone lets its chain-of-thought match the spatial reasoning of the full tool-using agent, with markedly more references to 3D perception and counterfactual motion concepts. Prior tool-augmentation routes outsource spatial reasoning to code execution and geometry tools; this work identifies that inference-time benefit as a transferable training signal rather than a mandatory external dependency. On shared spatial benchmarks with Qwen3.5-9B, four settings are compared: direct QA, prompted CoT, spatial coding agent, and agent-trace-conditioned CoT, with counts of 3D perception and counterfactual motion concept references in trace-conditioned rationales.

SpatialOPSD uses a frozen teacher initialized from the same checkpoint (which sees verified traces) to distill into a student that sees only raw input, internalizing tool-augmented spatial reasoning into a standalone model. Unlike answer-level SFT or GRPO with sparse outcome rewards, the method supplies dense token-level supervision on the model's own rollouts and needs no reward, advantage, or value model. Of the 10,000 MindCube training questions, SpatialClaw-Qwen3.5-9B answers 5,008 correctly, and SpatialOPSD and GRPO share those 5,008 instances; SFT uses Qwen3.7-plus to generate tool-free rationales for all 10,000 and retains 4,459, with all three methods using byte-identical student inputs and the same evaluation protocol.

Repetition-Aware Distillation detects self-repetition and privileged-information transcription via n-gram matching, masks distillation supervision at flagged tokens, and applies an unlikelihood penalty, preserving knowledge transfer while sharply reducing privileged-information leakage. Prior work on privilege leakage typically redesigns the teacher's supervision; here the privilege is left intact and only student tokens that reproduce trace-exclusive content are penalized. Ablations show mask alone gives the highest average accuracy (53.0) and unlikelihood alone the lowest leakage (6.7%); combining both cuts leakage from 82.7% under vanilla OPSD to 9.1% at 51.9 average accuracy, and sweeps over penalty weight and minimum repeated-span length show a moderate penalty both suppresses leakage and improves accuracy.

The framework improves accuracy across multiple student scales and different coding agents, and on out-of-domain benchmarks it matches or slightly exceeds the CoT baseline. This indicates the route from tool-assisted data generation to standalone spatial reasoning does not depend on a particular model scale or a particular agent implementation. Average accuracy rises from 40.6% to 42.8% at 2B, 46.4% to 50.2% at 4B, and 50.6% to 54.2% at 27B; across four out-of-domain benchmarks it reaches 64.7% versus 38.4% for SFT and 63.4% for GRPO, slightly above the 64.4% CoT baseline; traces from both SpatialClaw and pySpatial improve the base model.

Perspective

The result targets multimodal spatial reasoning that infers latent 3D properties from 2D observations, and applies when verified tool traces are available at training time while inference must be a single tool-free forward pass. The method depends on a coding agent that can correctly solve the training queries as the privilege source; here SpatialClaw-Qwen3.5-9B answers 5,008 MindCube training questions correctly to build that source, and the text does not report how queries the agent cannot solve are handled. Evaluation covers two spatial benchmarks, MindCube-Tiny and ViewSpatial, plus four out-of-domain benchmarks, BLINK, OmniSpatial, MMStar, and Video-MME, which readers can use to judge applicability.

The leakage rate uses a deliberately loose keyword rule that only detects standalone occurrences of agent, trace, or tool and does not infer semantic intent, so it reflects surface wording rather than true dependence; readers should note the gap to semantic leakage. The quantitative analysis rests on lexical marker frequency and rollout coverage over 1,040 MindCube-tiny samples, and the authors state these measure articulation style rather than intrinsic reasoning ability, so they should not be read as causal evidence of improved reasoning. Among privileged-information types, summary performs best, attributed to a moderate teacher-student entropy gap, but how the entropy-gap trajectories of full-trace and fact specifically affect optimization remains an open question. In addition, the SFT training set is not matched sample-for-sample with the SpatialOPSD and GRPO training sets, a data-construction difference to keep in mind when comparing across methods.

Sources