LoopCD treats a looped Transformer's earlier iteration as a weak reference, lifting Ouro-2.6B's AIME 2024 pass@1 from 61.88% to 73.33% and cutting 22.5%–48.2% of forward FLOPs at half the loops
Synopsis
The work introduces LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting a looped Transformer's final prediction with an earlier recurrent pass, in either logit space with one extra output pass (LoopCD-Logits) or hidden-state space with zero output overhead (LoopCD-Hidden); across four looped Transformer families it delivers consistent gains at full recurrent depth, for example raising Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33% and Huginn's HumanEval pass@1 from 22.56% to 31.71%, and these gains let it match or exceed full-depth unguided baselines at half the recurrent loops while reducing forward FLOPs by 22.5% to 48.2%.
Interpretation
LoopCD turns a looped Transformer's intermediate recurrent states into a built-in weak reference, enabling contrastive decoding without auxiliary models or extra training. Prior contrastive decoding needed an auxiliary smaller model, a perturbed context, or an intermediate layer to obtain an aligned weak predictor; recurrence itself repeatedly applies the same parameter block in the same representation space, inherently supplying aligned weak-and-strong prediction pairs for the same prefix. The paper validates this across four families (Ouro, Huginn, Parcae, Looped-Qwen3) and gives mechanistic evidence: decoding the first iteration alone trails the final converged prediction by 4.0 to 23.7 points on the seven-benchmark mean, and the gain grows with the disagreement between the early reference and the settled prediction.
The two guidance spaces trade off differently: logit space costs one extra output pass, while hidden-state space combines representations before the output layers with zero extra output overhead. The paper treats the guidance space as an explicit design axis and shows LoopCD-Hidden combines states before the coda layers, preserving the baseline's single pass through those layers. On multiple-choice scoring, LoopCD-Hidden raises Huginn's seven-benchmark mean by +0.83 and +0.75 points, exceeding fixed LoopCD-Logits (+0.74 and +0.62); on code generation the four-column mean improves by +5.13 and +3.76 points, and an appendix analyzes how coda depth affects hidden-state gain retention.
The gains concentrate on decisions the model is already uncertain about, and the mechanism is essentially re-ranking close decisions. The paper decomposes the logit contrast vector into a parallel component (a temperature-like rescaling) and an orthogonal component (pure re-ranking), showing the orthogonal re-ranking component alone matches or exceeds the full update. On HellaSwag, the isolated orthogonal re-ranking component yields +1.72 points for Ouro-1.4B and +2.27 for Ouro-2.6B, above the full update's +0.76 and +1.56; grouped by confidence, the least confident fifth gains +6.4 to +13.3 points while the most confident fifth gains at most +0.4, with only 5.8% to 14.9% of answers changed overall.
Those gains can be traded for compute: at half the recurrent loops, LoopCD still matches or exceeds full-depth unguided baselines, cutting forward FLOPs by 22.5% to 48.2%. The paper uses the decoding-quality gains directly to reduce the recurrent iteration budget and provides an architecture-level FLOP accounting showing savings are bounded by the loop's share of the full forward pass. Halving loops costs the unguided seven-benchmark mean 0.17 to 1.29 points, while applying LoopCD at that depth adds 0.75 to 1.29 points across six settings spanning Huginn, Parcae, and Looped-Qwen3; Huginn-0125 at sixteen of its thirty-two iterations beats its full-depth unguided baseline by 1.02 points (Logits) and 0.51 points (Hidden).
Perspective
The result targets inference and generation with looped Transformers: whenever a model repeatedly executes a shared parameter block and its intermediate states can be decoded by the same output head, LoopCD can be applied directly at decoding time, with no auxiliary model, no extra training, and no prompt changes. It applies to evaluation settings such as mathematical reasoning, code generation, and multiple-choice scoring, and still matches or exceeds full-depth unguided baselines when recurrent loops are halved; for fully recurrent architectures (such as Ouro and Huginn) and for architectures with a deeper coda (such as Parcae and Looped-Qwen3), the logit and hidden-state forms each suit different cases.
Reference-iteration choice depends on how the architecture is initialized: Huginn starts from Gaussian noise, so its first hidden state is dominated by noise and its hidden-state reference peaks only at the sixth step, whereas the logit reference is generally best at the first iteration. Guidance strength tolerates a wide band for multiple-choice scoring but needs smaller values for autoregressive generation, where early token shifts compound. The paper also notes that the halved-depth evaluation is run only on the multiple-choice suite, that the FLOP figures are theoretical arithmetic workloads rather than end-to-end wall-clock speedups, that GSM8K results are mixed, and that Looped-Qwen3 shows pass@1 declining while pass@10 rises on mathematical reasoning; these are areas to keep watching in relation to specific tasks and architectures.
