Skip to main content
Back to timeline
arXivSource publication:

Freezing a 7B backbone and training only 100K bias parameters: label-free TTRL reaches 76.67% on MATH-500 and transfers to 4,500 problems unseen during optimization

Synopsis

The work introduces label-free bias-only test-time reinforcement learning: the pretrained backbone stays frozen while only about 100K bias parameters (down_proj.bias across 28 MLP layers) are optimized against majority-vote pseudo-label rewards, reaching 76.67% on MATH-500 with Qwen2.5-7B and 79.50% with Qwen2.5-Math-7B, improving vision-language and audio reasoning under the same procedure, and transferring frozen bias vectors to 4,500 verified-disjoint MATH problems to lift Qwen2.5-7B from 46.3% to 70.9% and Qwen2.5-Math-7B from 52.5% to 75.4%, with pseudo-label reliability and accessible gradient energy explaining why a restricted subspace still adapts.

AI-generated editorial illustration: Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Interpretation

Substantial test-time adaptation survives an extreme optimization bottleneck: with a frozen backbone, only about 100K trainable bias parameters, and rewards derived entirely from the model's own rollouts via majority vote, all six model-benchmark settings improve over the untrained baseline, reaching 79.50% on MATH-500 with Qwen2.5-Math-7B and 76.67% with Qwen2.5-7B. Prior TTRL-style methods typically update most or all parameters; this work keeps TTRL's majority-vote reward and GRPO unchanged and restricts optimization to a bias subspace, turning parameter efficiency itself into the variable that probes what makes restricted test-time learning possible. Results are reported as mean plus standard deviation over three independent training seeds under a fixed 200-step budget, and the same 100K-parameter procedure is used across text, vision-language, and audio with only task-specific prompts and reward/grading functions changed.

The learned bias vectors act beyond the problems used for test-time optimization: freezing the step-200 checkpoints and evaluating without further training on 4,500 verified-disjoint MATH problems raises Qwen2.5-7B from 46.3% to 70.9% and Qwen2.5-Math-7B from 52.5% to 75.4%. This tests whether the adaptation is merely overfitting to the 500 MATH-500 problems, and the result indicates the intervention transfers to a problem set not used during optimization. Transfer evaluation is performed on frozen checkpoints with no further training, on problems described as verified-disjoint.

Under matched parameter budgets, the choice of trainable subspace matters more than parameter count: across nine single-layer bias subspaces of exactly 3,584 parameters each, accessible gradient energy correlates with post-training accuracy at Spearman rho = 0.83 (exact two-tailed permutation test, p < 0.01), while a parameter-matched LoRA (exact-match, 100K parameters) trails bias-only TTRL on all six benchmarks and falls below the untrained baseline on LogicVista. The paper characterizes restricted-subspace trainability through accessible gradient energy, the squared norm of the full gradient projected into the trainable subspace, and gives a restricted-subspace improvement bound showing guaranteed local improvement is governed by that quantity rather than by subspace dimensionality. Theory is supported by proof sketches of a proposition and a theorem; empirically, the evidence is a nine-layer sweep with a rank-correlation test plus controlled comparisons of different parameterizations (bias, LoRA, a ReFT-style variant) at the same budget.

Pseudo-label reward reliability improves with the number of rollouts: under a positive-margin assumption, the probability that the majority-vote pseudo-label is incorrect decays exponentially in the rollout count, and experiments show pseudo-label accuracy ranking tracks both rollout count and consensus strength. This places the reliability of the self-generated signal alongside subspace alignment as the two requirements for restricted test-time learning, whereas prior work focused mainly on constructing the reward itself. Theory is a proof sketch based on Hoeffding's inequality and a union bound; empirically, rollouts are generated once per problem on a batch of MATH-500 problems, majority vote is computed as a prefix for each rollout count, and a consensus proxy is recorded alongside accuracy.

Perspective

The result targets deployment-time label-free adaptation: a pretrained checkpoint is available, multiple rollouts can be sampled for the same unlabeled problems, and a fixed additive bias vector can be applied at inference. It applies to 7B-scale models with 28 MLP decoder layers (Qwen2.5-7B, Qwen2.5-Math-7B, Qwen2.5-VL-7B-Instruct, Qwen2.5-Omni-7B), with a fixed 200-step training budget and greedy pass@1 evaluation. Because the method needs no ground-truth labels, it suits label-scarce mathematical, vision-language, and audio reasoning tasks; the paper also reports a length-aware reward variant (ShorterBetter) that slightly exceeds the binary reward on Qwen2.5-7B (75.2% vs. 74.2%) while cutting average response length by 30%, and large absolute gains at 1.5B scale (Qwen2.5-1.5B +54pp, Qwen2.5-Math-1.5B +37pp), suggesting the route can start from smaller models as well.

The composition of the multimodal gains still needs task-by-task reading: on AI2D part of the improvement comes from a lower official-parser extraction-failure rate, that is, better output-format compliance; on LogicVista an exhaustive audit of all 448 problems finds 77 flipped from wrong to correct and 47 from correct to wrong with no single dominant confound; on MathVista the steered model answers exactly 0 on all 49 age-gap questions, a template contributing one net problem out of 62 net-improved problems; on MMAU the template asking to count words containing a stressed or unstressed phoneme covers 15.3% (153/1000) of the set and improves by 15.1pp while the remaining 847 problems improve by only 1.4pp, so roughly two-thirds of the MMAU gain traces to that template, and the paper states that excluding it leaves a gain within seed variation, making MMAU the weakest of the six benchmark results. In addition, the comparison with a labeled bias vector in Appendix D is not a controlled one: the labeled vector was trained with RLOO on 40K labeled DeepScaleR problems and evaluated out-of-distribution on MATH-500, reaching 75.6% versus 72.0% for the label-free vector, which points in a different direction from the main-text Table 2 finding that the label source is not the binding constraint at a matched recipe; the paper explains that the two answer different questions. The accessible-gradient-energy analysis rests on a single-layer sweep with a shared backward pass on a fixed 16-problem batch, so whether accessible gradient energy predicts trainability across layers, tasks, and model scales, and how the tendency of majority-vote pseudo-labels to converge on a fixed answer for narrow templates can be detected by the training process itself, remain open questions.

Sources