Reinforcement-learning updates move into vision-language models: keeping only the leading direction beats full transfer by 8.55 points on MathVision
Lead
Isolating the reinforcement-learning-stage parameter update, keeping only each matrix's leading direction rescaled to its original magnitude, and transferring it into the language modules of a vision-language model beats full-update interpolation in 12 of 15 backbone-benchmark comparisons across three model families and five visual-reasoning benchmarks, including an 8.55 percentage-point MathVision gain on the Qwen recipient.
Story
When reasoning capability moves across models, what is worth transferring is the parameter change acquired during the reinforcement-learning stage, not the whole difference between donor and recipient. Model merging previously built a donor-recipient displacement from endpoints, and that displacement mixes differences that predate reinforcement learning with changes acquired during post-training. Once the reinforcement-learning update is isolated, it beats full-update interpolation sharing the same parameter support in 12 of 15 comparisons across the Qwen3-VL-8B-Instruct, LLaVA-NeXT-LLaMA3-8B, and Idefics2-8B recipients on five visual-reasoning benchmarks.
The construction takes the difference between the donor's post-trained checkpoint and the checkpoint its reinforcement learning started from, decomposes each eligible language-layer matrix, keeps only its largest singular direction, and rescales that direction to the original matrix's Frobenius norm. The complete update was optimized for the source language model, so its directions need not be equally useful to a multimodal recipient; retaining more directions progressively lowers MathVista accuracy. On Qwen3-VL, the norm-matched leading direction reaches 81.00 on MathVista, against about 74.1 for ten matched random directions, 73.10 for the trailing direction, and 76.70 for the full update.
What to watch
The next step is to test whether leading-direction selection holds on more language-vision pairings and more reasoning tasks, especially the MMStar and MathVision regressions seen on H2 and H3. Engineers working on merging or transfer can reuse the offline pipeline directly: decomposition, normalization, coefficient selection, and checkpoint construction all happen offline, and the checkpoint stays a standard dense model with no added inference module. Gains should be checked per task, since 3 of 15 comparisons did not exceed the full update.
Why the leading direction transfers better than the complete update is shown only through empirical controls, with no mechanism attribution, so that remains open. Gains are not uniform: H2 loses on MMStar, and H3 declines on MathVision and MMStar, so a careful reader should watch whether these reversals recur on more backbones. Fixing rank at one is an empirical choice, and accuracy falls as rank grows from one to sixteen, but the text does not state where that rule stops applying. Source-model reconstruction fidelity and recipient gains are distinct measurements and cannot substitute for each other.
