Skip to main content
Back to timeline
arXivSource publication:

TaH2 allocates loop iterations per token and raises the AIME accuracy-compute slope by 53%

Synopsis

This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.

AI-generated editorial illustration: Improving Test-Time Scaling with Adaptive Looped Transformers

Interpretation

The paper characterizes the test-time scaling behavior of looped transformers: with the same pretrained checkpoint and post-training data, fixed-depth Ouro, adaptive Ouro and Huginn reach accuracy-compute slopes of 2.44, 2.44 and 2.36 points per compute doubling, all steeper than the non-looped Standard baseline at 1.79, yet they remain less accurate than Standard over the overlapping compute range. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs; this work instead measures the accuracy gain per doubling of decoding FLOPs while sweeping output-token budgets from 4K to 16K, moving the comparison axis from model size to test-time compute. Based on a unified post-training setup from Qwen3-1.7B-Base, with 32 samples per AIME24-26 problem, output cutoffs swept from 4K to 16K in 2K increments, and linear fits of accuracy against the logarithm of decoding FLOPs.

Token-level analysis shows fixed-depth looping spends extra iterations on many tokens that do not benefit: the next-token loss change distributions for Ouro and Huginn concentrate near zero, with 68% and 66% of tokens changing by at most a threshold and 21% and 24% becoming worse by more than that threshold. This offers a token-granular explanation for why looping trails at matched compute and points directly to allocating iteration depth per token rather than further deepening or widening the loop. Measured on 1,000 validation samples randomly drawn from the training mixture and excluded from training, comparing next-token loss after the first and final iterations.

TaH2 jointly post-trains the backbone and an iteration decider, using lookahead depth supervision to generate online labels from the current backbone indicating whether further iteration improves the prediction, with a cost-sensitive loss supervising each depth decision; on AIME it improves the accuracy-compute slope by 53% (2.74 vs. 1.79) and exceeds Standard's peak accuracy by about 3.4 points at matched test-time compute. Unlike TaH's staged training with offline mismatch labels, and unlike Ouro's second stage that freezes the backbone and trains only the gate, TaH2 derives depth labels online from iteration gains measured on the current backbone while the backbone and decider are updated together. 1.7B scale, avg@32 evaluation on AIME24-26 with evaluation extended to 32K tokens; ablations show replacing gain labels with top-1 mismatch labels costs 1.1 points on average, uniform instead of cost-sensitive weights costs 3.6 points, and retaining all positive gains costs 1.9 points.

As the maximum iteration depth ceiling increases, existing looped models largely plateau while TaH2's gain over the non-looped baseline grows from +2.8 points at depth 2 to +3.9 points at depth 8; the advantage persists at 4B and 8B scales and extends to code, QA and tool-use tasks. Depth ceilings tend to bring saturation or degradation in existing looped models, whereas TaH2 converts the depth ceiling into sustained accuracy gains, indicating that adaptive depth allocation changes the payoff structure of depth scaling. At 1.7B the average gain across ten benchmarks grows from +2.9 to +4.8 points as the depth ceiling rises; 4B improves by 3.2 points and 8B by 2.4 points on average, with AIME gains up to 7.4 and 4.4 points respectively.

Perspective

The results target post-training with supervised fine-tuning on Qwen3-{1.7B,4B,8B}-Base backbones over math, code, QA and tool-use data, evaluated primarily by avg@32 on benchmarks such as AIME24-26 against decoding FLOPs, with throughput verified on a Mini-SGLang engine that batches requests at different iteration depths. It makes it plausible for the post-training route of adding recurrence to existing pretrained models to surpass the non-looped baseline at matched test-time compute, and offers a starting point for testing adaptive depth at larger scales, longer output budgets and more task families.

The authors state the method is studied only under supervised fine-tuning, leaving extension to on-policy distillation and reinforcement learning for future work; training FLOPs exceed standard SFT, though the authors note post-training costs far less than pretraining. When inference depth exceeds the training ceiling, accuracy is essentially unchanged while mean depth rises, so extra depth neither breaks the model nor improves it further. In addition, this reading is of the full paper text with figures described in prose, so exact curve shapes and point-by-point values still warrant checking against the original figures.

Sources