Skip to main content
Back to timeline
arXivSource publication:

Same Reward, Different Skills: When Multimodal RL Learns to Look

Related research and updates

Synopsis

The work proposes a "visual resolvability" design rule requiring that correct answers depend on visual evidence while the task stays learnable, and tests it on counterfactual coordinate scenes with standard GRPO and correctness-and-format rewards on Qwen2.5-VL-7B: on held-out 20-point scenes denser than any seen in training, discovery accuracy rises from 0.425 to 0.875; replacing test images with gray canvases drops accuracy to zero, while training on gray canvases at matched step 30 across four seeds yields essentially none of the gain, indicating the gain comes from training images; adding a caption that answers the training question, with the same images, reward and budget, cuts the gain by nearly two thirds.

Source-provided article image: Same Reward, Different Skills: When Multimodal RL Learns to Look
Figure 1 ·

Figure 1: Same learning rule, different lessons. Left: shortcut-accessible tasks allow reward without visual evidence; gains in the reported 3B comparisons mainly favor cued readout. Right: constructed tasks demand visual evidence and relation-based target identification; discovery, cued readout, and named-target grounding improve. The hierarchy distinguishes supplied locations, supplied identities, and relation-defined targets. GRPO is shared across the study; rewards stay fixed within matched comparisons.

arXiv

Interpretation

Under ordinary RLVR, benchmark gains on vision-language tasks do not equal visual acquisition: on Geometry3K, replacing every training image with a gray canvas still lets the 3B model recover roughly half of the real-image gain and the 7B model nearly four fifths (7B recovery share 0.78, 95% CI [0.64, 0.92]). Prior readings of visual-reasoning benchmark gains assumed that rewarding correct answers develops the needed visual operations; this work separates "depends on the image at test" from "learned from images" via crossed training-time and test-time visual access. Three seeds at 3B across four visual-access conditions (Real/Caption/None/Gray), reporting gains and recovery shares; the 7B share carries a bootstrap interval.

Counterfactual coordinate tasks built under visual resolvability (fixed question, unnamed target, relation only) raise 7B discovery accuracy from 0.425 to 0.875 on a held-out 20-point confirmatory set (gain 0.450, 95% CI [0.365, 0.540]), and transfer to never-trained cued readout and named-target grounding, and to independently constructed coordinate-grounding tasks (pair accuracy 0.768 to 0.948). Unlike approaches that act through the objective or data selection, this work writes the required visual operation into task construction, using a cue hierarchy (L1 marks location, L2 names target, L3 gives only the relation) to separate the operations measured. Four seeds and five runs; Real exceeds Gray by +0.354 (seed SD 0.019) on the confirmatory set and +0.339 (SD 0.021) on an untouched set of 1,332 members; transfer tasks predate the training corpus and were never trained on.

Two controls locate the source of the gain: replacing test images with gray canvases reduces discovery and probe accuracy to exactly zero; training on gray canvases (same prompts, starting checkpoint, reward and budget) yields essentially none of the discovery gain at matched step 30 across four seeds, showing the gain comes from training-time visual information. This matched blind-training control separates "the model uses the image at test" from "the model learned the skill from images," a measurement prior vision-centric task constructions did not make. Paired comparison of Real and Gray under identical evaluation, with a seed-level interval of [0.323, 0.384]; under gray canvases the models give fixed default answers rather than refusals.

With the same images, reward and budget, adding a caption that answers the training question still earns reward (accuracy component 0.97-0.99 by step 30) but yields only +0.120 confirmatory discovery gain against Real's +0.323, a gap of +0.203 (seed-level 95% CI [+0.163, +0.243]). This shortcut-accessible control shows that changing what reward requires changes what RL learns, even when the reward value itself is unchanged. Three seeds, with every shortcut seed below every Real seed; even with its caption at test the arm reaches at most 0.655, still below Real's image-only level.

Perspective

The results target researchers and engineers training vision-language models with verifiable rewards, in a controlled testbed where visual necessity can be measured and controlled directly (counterfactual coordinate scenes). In this setting the vision encoder and projector stay frozen and only the language model is trained, so "learning to see" appears as learning to look: a given image supplies the same visual tokens before and after. The next test of the principle is applying it to other visual formats and to natural images, where visual necessity must first be measured; the authors also propose extending the same evidence-centered idea to audio, video and tool-use tasks.

Readers should still watch: coordinate scenes are a deliberate testbed, and generalization to natural images and other visual formats is untested; the long-horizon corrosion runs change horizon, encoder trainability, corpus filter and reward implementation together, so they show the loss under this recipe rather than isolating one cause; the answer-matcher revision affects interpretation of the shared-error sets (the current revision lowers image-removed accuracies and raises gains); the requirement that the artifact screen pass at each of three seeds was added after generation began and before any criterion was evaluated; and training reward exceeds 0.97 by step 13 while discovery keeps improving through step 75, so reward saturation is not a signal that learning has stopped.

Sources