Soft Spatial Reasoning uses AdaptSoft to tune softness per step, lifting OmniSpatial weighted-average accuracy to 49.68
Synopsis
The work proposes Soft Spatial Reasoning, a post-training framework in which a large vision-language model forms a continuous soft state at each chain-of-thought step by mixing token embeddings instead of committing to a single token, with an AdaptSoft controller that adapts the degree of softness from the current hidden state and predictive uncertainty and is trained by a gradient-alignment objective; across OmniSpatial, SpatiaLab, and MindCube it reaches higher weighted-average accuracy than same-backbone hard-thinking and fixed-soft chain-of-thought baselines and the compared existing models.
Interpretation
It introduces a post-training framework for soft thinking in spatial reasoning by large vision-language models, representing intermediate reasoning steps as continuous soft states built from mixtures of token embeddings while the final answer remains discrete language. The authors describe it as the first framework dedicated to soft thinking for spatial reasoning in large vision-language models; prior soft thinking was mostly explored in language models, or in vision-language models kept continuous visual representations while the chain of thought still advanced through discrete token selections. The paper provides method derivations, including a Gumbel-reparameterized likelihood and a soft-step GRPO ratio, and controlled comparisons against same-backbone hard-thinking and fixed-soft chain-of-thought baselines on three benchmarks.
It designs AdaptSoft, a controller that sets softness per reasoning step from the current hidden state and the predictive uncertainty of the next-token distribution, so temperature varies around a base value across steps. Existing soft-thinking approaches typically use a fixed degree of softness, whereas this work makes softness adaptive per step to trade off preserving multiple candidate interpretations against interference from conflicting spatial relations. Ablations show average accuracy drops 3.53 points when the hidden-state input is removed and 2.11 points without predictive uncertainty; the appendix analysis reports temperature falling as normalized entropy rises, with a correlation of -0.71.
It proposes a gradient-alignment learning objective that supplies a step-specific learning signal for softness control by measuring how each step's contribution to the policy gradient aligns with a reference gradient, without intermediate reasoning annotations. The shared rollout-level advantage in GRPO does not distinguish how softness should vary across steps, and this objective supplies that signal without external evaluators or intermediate reasoning supervision. In ablations, replacing step-specific gradient alignment with the rollout-level task advantage produces the largest average decline of 4.07 points, including a 6.52-point drop in Hypothetical reasoning; omitting gradient centering reduces accuracy by 1.57 points on average.
Across OmniSpatial, SpatiaLab, and MindCube, the method attains the highest weighted-average accuracy among the compared non-proprietary models and outperforms same-backbone hard-thinking and fixed-soft chain-of-thought. With the same backbone and post-training setup, basic soft thinking improves on hard thinking by 1.01 points and adaptive softness adds another 2.75 points; under zero-shot transfer to SpatiaLab, adaptive softness adds 1.85 points over basic soft thinking, compared with 0.65 points from hard to basic soft thinking. OmniSpatial weighted average is 49.68, surpassing InternVL3-14B, SoFar, and LaCoT by 3.74, 4.54, and 4.02 percentage points; SpatiaLab is 48.71, exceeding the strongest prior open-weights model LVR by 2.93 percentage points; MindCube is 38.13, 4.51 points above the backbone.
Perspective
The framework targets settings where a large vision-language model answers spatial reasoning questions; training uses the 6,902-sample OmniSpatial training split, and evaluation covers the 1,533-sample held-out OmniSpatial test split plus zero-shot SpatiaLab and MindCube. The method is instantiated with a Qwen3-VL-8B-Thinking backbone, and soft thinking applies only to reasoning tokens, switching to discrete answer generation after </think>. For researchers and engineering teams wanting to bring soft thinking into vision-language model reasoning, this offers a reproducible controller and training objective, with source code released.
AdaptSoft currently lacks an explicit measure of uncertainty in visual evidence, and the authors list incorporating visual uncertainty at individual chain-of-thought steps as future work. The analysis relating temperature to uncertainty is based on statistics over inference rollouts, and the magnitude of adaptive-softness gains varies considerably across spatial tasks, so the conditions under which it helps most remain to be tested in more settings. This summary is based on the full paper text and the homepage evidence bundle without the original figures, so the specific per-step temperature values in Figure 4 cannot be restated here.
