Scaffolding Minds swaps in a learnable scaffolding encoder and an adaptive Gaussian sampler for latent visual reasoning, gaining 9.5 points on FrozenLake and 5.6 points across nine visual reasoning benchmarks
Synopsis
The work proposes Scaffolding Minds, a two-stage framework in which Stage 1 replaces the frozen off-the-shelf vision encoder with a scaffolding encoder trained end-to-end through the downstream task loss to produce the latent target, and Stage 2 replaces deterministic latent regularization with a Gaussian sampler whose mean and variance are learned so that residual latent actions can be sampled for RL exploration, improving over the strongest latent-reasoning baseline by 9.5 points on average on FrozenLake (19 points at Level 32) and by 5.6 points on average across nine visual-centric reasoning benchmarks.
Interpretation
Stage 1 changes the latent supervision target from frozen general-purpose vision-encoder features to the output of a scaffolding encoder optimized end-to-end through the downstream task loss, raising FrozenLake average accuracy from 62.0 to 72.0 and the nine-benchmark average from 66.3 to 70.9. Prior latent visual reasoning methods (LVR, Mirage, CoVT, Monet, VaLR and others) derive their latent target from a frozen off-the-shelf vision encoder that is not aligned with the reasoning task; here the target itself becomes a trainable component, implemented as a trainable copy of the VLM vision encoder plus a cross-attention pooling module, used only during training and discarded at inference. Ablations show that training only the attention-pooling module already lifts accuracy, and letting the task loss also shape the vision encoder adds further gains, together accounting for the full gap over the frozen target; a parameter-matched control that fine-tunes the same vision encoder in plain SFT gains only a small amount, indicating the gain comes from optimizing the target rather than from added capacity.
Stage 2 samples residual latent actions from an input-adaptive Gaussian and treats the latent block as one more policy action inside the GRPO objective, updated from the same reward as the text tokens, raising FrozenLake average accuracy from 72.0 to 75.0. Existing recipes either optimize only the text trajectory (GRPO) or apply a fixed-variance Gaussian as deterministic regularization on the latent block (VLPO), neither of which samples alternative latent actions or explores latent trajectories; this work learns both mean and variance with two lightweight MLP heads, with the mean head zero-initialized and the variance head starting from small positive values. From the same Stage 1 checkpoint, text-only GRPO and VLPO give smaller gains and VLPO training fluctuates; learning only the mean or only the variance yields partial gains, while learning both gives the largest gain; on the weaker frozen-target checkpoint Scaffolding RL gains even more, consistent with more headroom for latent exploration.
The two stages are complementary: Stage 2 refines but cannot replace the latent target Stage 1 provides, since on the weaker checkpoint Scaffolding RL still lands below even the weakest recipe on top of the scaffolding encoder. This decomposes the contribution of target quality versus exploration mechanism within the two-stage paradigm, whereas prior work typically changes only one of the two. The same two-stage ablation is repeated on FrozenLake and on the nine visual-centric benchmarks, where both components improve results and the improvement covers every benchmark.
The method outperforms image-generation and tool-calling approaches on both spatial planning and real-world visual reasoning while keeping inference cost close to the base VLM. Prior thinking-with-images methods must generate or manipulate explicit images at inference time at substantial cost; this work performs multi-step visual reasoning entirely in latent space. On FrozenLake it improves by 10.7 points on average over image-generation methods; across the nine benchmarks it improves by 6.2 points on average over the strongest thinking-with-images method; single-query latency measurements show latent reasoning methods add modest overhead while tool-calling and image-generation methods add substantially more.
Perspective
The work targets latent visual reasoning systems trained with a two-stage paradigm, in settings where intermediate helper images are available during training, such as value-function heatmaps for FrozenLake and the crops, bounding boxes, diagrams, and candidate placements provided by Zebra-CoT. At inference the scaffolding encoder and helper images are discarded and only the VLM is used, so deployment cost stays close to the base model. For teams seeking to improve multimodal reasoning without generating explicit intermediate images, it offers a reusable two-stage recipe: optimize the latent target first, then explore in latent space under reward feedback. The authors also note that extending the framework to settings where helper images are unavailable or expensive, automatically proposing helper images for new domains, and adapting latent block length to the difficulty of each reasoning step are directions for future work.
The availability of helper images during training is a precondition of the framework, and how to extend it when helpers are missing or expensive remains an open question; the authors do not exhaustively search alternative helper-image designs, nor do they automatically propose helper images for new domains. Each latent block uses a fixed number of tokens, and adapting block length to the difficulty of each reasoning step is not yet implemented. The teacher-forcing versus self-predicted latent comparison shows only small accuracy differences, but it is run in the visual-centric setting, so behavior under other task distributions still needs observation. In addition, although this is a full-text read, some table values appear as placeholders in the text, so exact per-item numbers should be checked against the original tables.
