Skip to main content
Back to timeline
arXivSource publication:

Frozen vision-language models gain candidate diversity without training by reusing decoder-layer blocks, lifting pass@9 by 6.58 points on average

Synopsis

The work introduces architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers in a frozen model; across five Qwen checkpoints and twelve multimodal benchmarks it improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget, with the strongest gains from reusing early layers, candidate-coverage improvement persisting under greedy decoding, lower lexical overlap among candidates, and higher accuracy when the candidates serve as rollouts for label-free test-time reinforcement learning.

Source-provided article image: Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
Figure 1 ·

Figure 1: Two Sources of Candidate Diversity. Temperature sampling (left) draws several outputs from the same standard computation path, starting from the initial multimodal representation h ( 0 ) h^{(0)} . Architectural sampling (right) changes the computation path from the same h ( 0 ) h^{(0)} before producing one candidate from each path. The background is a conceptual illustration of the answer space.

arXiv

Interpretation

It proposes architectural sampling, which produces candidate answers through distinct forward computations rather than temperature perturbations along one fixed path. Conventional temperature sampling generates every candidate along the same fixed computation path, whereas this method varies the location and repetition count of reused decoder-layer blocks to introduce computational diversity. The abstract reports the method is training-free, updates no model weights, and adds no auxiliary parameters, evaluated across five Qwen checkpoints and twelve multimodal benchmarks.

At the same nine-candidate budget, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average. The gain comes from broader candidate coverage rather than more samples or changed model parameters. The abstract gives an average improvement across five checkpoints and twelve benchmarks and notes that reusing early layers yields the strongest gains.

The improvement in candidate coverage persists even under greedy decoding, and the resulting candidates show lower lexical overlap. This indicates the diversity is not solely a product of sampling temperature but is tied to differences in computation paths. The abstract explicitly reports that coverage improvement persists under greedy decoding and that candidates have lower lexical overlap.

These candidates improve accuracy when used as rollouts for label-free test-time reinforcement learning. It extends the value of architectural sampling from candidate coverage to more effective learning from a model's own outputs. The abstract reports accuracy improvement when the candidates are used in a label-free test-time reinforcement learning setting.

Perspective

The result targets settings that use frozen vision-language models and multi-candidate sampling for test-time scaling, especially for practitioners who want broader candidate coverage without updating weights or adding auxiliary parameters; the method generates candidates by reusing decoder-layer blocks and varying their location and repetition count, and the abstract notes that reusing early layers yields the strongest gains, so the benefit is clearest in that setting. Its value also extends to pipelines that use the candidates as rollouts for label-free test-time reinforcement learning.

The visible text is only the abstract and does not provide per-benchmark results, the specific configurations of block location and repetition count, computation-cost comparisons, or statistical significance information; the 6.58 percentage points is an average across checkpoints and benchmarks, and its distribution across tasks and model scales still needs confirmation in the full text. The mechanistic explanation for coverage improvement under greedy decoding, and the stability of the test-time reinforcement learning gains, are also directions a reader may continue to watch.

Sources