Skip to main content
Back to timeline
arXivSource publication:

ReaLVR supervises latent visual reasoning with visual evidence, reaching a 63.7% five-task average on Qwen2.5-VL-7B and scaling latent visual reasoning to 235B

Synopsis

The work first analyzes latent-token behavior in latent visual reasoning (LVR) and identifies a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer, which the authors attribute to the lack of explicit supervision during GRPO training; it then proposes ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories, using correct-versus-wrong answer contrast to decide where stronger supervision is needed and relevant-versus-mismatched visual evidence to specify what to preserve, consistently outperforming evaluated LVR baselines across three model families, reaching a 63.7% five-task average on Qwen2.5-VL-7B, and being the first to scale visual reasoning in latent space up to 235B parameters.

AI-generated editorial illustration: Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Interpretation

The paper identifies a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that change the correct answer, indicating that a final-answer reward alone cannot guide which visual evidence latent tokens should preserve or how credit should be assigned across them. Prior LVR work focused mainly on the form and flexibility of latent trajectories or evaluated them by final-answer accuracy; this work pushes diagnosis to whether latent tokens respond to answer-relevant visual evidence. Based on controlled behavioral and representation tests: four edit types each contain 512 original-edited pairs with a fixed question, the correct answer changes in most pairs, yet LVR changes its prediction in only a minority; free-running latent states also show weak cosine alignment with visual targets.

It proposes ReaLVR, which applies visual-evidence supervision to a regenerated, gradient-retaining latent trajectory of the current model, using correct-versus-wrong answer attention readout to locate positions needing stronger supervision and relevant-versus-mismatched visual prototypes to define what to preserve. Unlike GRPO, which optimizes only generated text, ReaLVR lets the evidence loss backpropagate through latent generation; allocating supervision by answer contrast is markedly more effective than spreading it uniformly. Component ablations show uniform routing costs 1.3 points on average, raw attention alone gives 62.9, removing negatives gives 62.1, supervising saved off-policy latents gives 61.9, and removing the stop-gradient gives 62.6, versus 63.7 for the full method; the method changes neither architecture nor inference procedure.

It consistently outperforms evaluated LVR baselines across three model families and six backbones, and is the first to train latent visual reasoning on a 235B-parameter multimodal backbone. Prior validation of latent visual reasoning was limited in scale; this work extends evidence supervision to frontier scale and reports improvements over LVR-SFT on all three evaluated benchmarks at Qwen3-VL-235B. Qwen2.5-VL-7B reaches the highest five-task average of 63.7% among evaluated latent-reasoning methods; Qwen3-VL-30B reaches 65.2%, 1.1 points above LVR; InternVL3-8B and Gemma-3-12B exceed LVR-RL by several points in five-task mean; the 235B evaluation covers only MMVP, BLINK, and MME-RealWorld, with no five-task mean.

Mechanism analyses show that ReaLVR's latent-token positions are more question-sensitive, align more strongly with relevant visual regions, and that the most answer-attended latent tokens are more load-bearing for answer likelihood. These analyses move evaluation beyond benchmark accuracy toward three separately measurable dimensions of latent states: variation, grounding, and use. Target-region attention enrichment reaches about 2x in middle layers versus near 1 for Monet and at or below 1 for LVR; replacing the eight most answer-attended latent tokens lowers ReaLVR's correct-answer probability from about 0.9 to about 0.3; a linear probe recovers the BLINK task label from the mean latent, while vocabulary readout does not expose a language rationale.

Perspective

The result targets research and engineering settings that use continuous latent tokens for multimodal reasoning: when region annotations are available, the visual evidence target is more spatially specific; when annotations are missing or the mask is empty, the target falls back to a whole-image prototype and spatial specificity decreases. The method fits backbones that already follow a two-stage LVR recipe, requiring the training side to regenerate latent trajectories with retained gradients, while inference still generates latent tokens and decodes answers with the original LVR procedure, so it can replace the existing LVR training stage directly. For teams seeking stronger visual grounding and greater answer dependence on latent states without changing the inference interface, this supervision offers a transferable starting point; for readers interested in latent-reasoning interpretability, the variation, grounding, and fixed-context replacement diagnostics serve as an evaluation template.

Vocabulary readout of latent tokens does not reveal a language rationale, with the top-1 token tied to a closing-tag prior, so whether latent states carry readable reasoning remains open; fixed-context replacement measures only local answer dependence, and regenerating later states could introduce further effects. Target-region attention enrichment curves are similar for correct and incorrect responses, indicating that looking at the right region is not sufficient for answering correctly. The latent-length ablation shows the best budget is task dependent, while inference currently uses a prescribed latent-token budget, leaving adaptive budgets unresolved. In addition, the 235B scale is evaluated on only three benchmarks with no five-task mean; the loaded text is the full paper, but some table values appear as placeholders in the prose, so exact numbers should be checked against the original tables.

Sources