Looped language models gain reasoning but lose knowledge when unrolled beyond the training horizon, and history-state injection with timestep conditioning mitigates the trade-off
Synopsis
Through controlled pretraining experiments, this study systematically examines when recurrence helps, where it should be applied, and how it should be conditioned in looped language models (LoopLMs), finding that extra loops beyond the training horizon improve reasoning (e.g., a reasoning score rising from 28.52 to 31.49) while degrading knowledge (from 62.80 to 52.11), that non-recurrent output layers improve robustness to under-unrolling, and that the proposed history-state injection (especially channel-wise) combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness across inference budgets.
Interpretation
Additional recurrence beyond the training horizon has divergent effects: reasoning performance continues to improve while knowledge performance degrades, and harder reasoning instances do not consistently benefit more. Prior work mainly evaluated LoopLMs at the fixed training loop count or focused on early exiting, leaving test-time recurrent scaling beyond the training horizon unclear; this study uses controlled experiments covering inference budgets below, at, and beyond the training horizon, with knowledge and reasoning task groups. Pretraining from scratch on Llama3.1-1B and Qwen3-0.6B backbones while matching physical depth, training effective depth, inference effective depth, and training pipeline; reports concrete numbers such as reasoning rising from 28.52 to 31.49 and knowledge dropping from 62.80 to 52.11, plus roughly 7.5 percentage point absolute gains at deeper proof depths on ProofWriter.
Recurrence effectiveness depends on how physical depth and loop count are allocated, not on effective depth alone; the placement of non-recurrent boundary layers also varies with inference budget. Prior designs often evaluated LoopLMs by effective depth or fixed architecture; this study shows that configurations with the same training effective depth but different physical-depth/loop-count splits scale markedly differently, and distinguishes the roles of Prelude/Coda in BaseLoop versus CoreLoop. Compares multiple configurations under matched effective training depth and token budget; finds that a shallow physical core with many loops peaks early and degrades, a deep physical core with few loops stays relatively flat during extrapolation, Coda layers mitigate knowledge decay under under-unrolling, and Prelude-heavy allocation better supports reasoning extrapolation.
Conventional initial-state injection offers limited robustness to varying recurrence depth, and high-capacity dense injection causes knowledge collapse under deep unrolling; the proposed history-state injection conditions on relative state differences and substantially rescues deep extrapolation. Initial-state injection (e.g., Huginn-style concatenation plus linear projection) is widely used, but this study systematically compares scalar, channel-wise, residual channel-wise, and dense parameterizations, identifies their limitations, and proposes an alternative based on relative differences of historical states. Evaluates multiple injection variants on Qwen3-0.6B and Llama3.1-1B; dense initial-state injection collapses on knowledge during extrapolation, while scalar history injection with window 1 achieves the highest overall and knowledge scores at a given depth, exceeding BaseLoop by 1.64 and 6.89 percentage points, respectively.
History-state injection and timestep conditioning form the strongest complementary pair, and channel-wise history-state injection plus timestep conditioning improves both knowledge and reasoning across extended inference budgets at low cost. Beyond proposing two conditioning mechanisms, the study uses pairwise combination experiments and geometry analysis (computational interaction) to show that their complementarity comes from enhanced cross-loop interaction rather than merely larger state transformation magnitude. Evaluates all pairwise and triple combinations on Qwen3-0.6B BaseLoop up to 3x training depth; H+LG reaches 37.30 overall accuracy at depth 84, exceeding I+H (36.91) and I+LG (36.33), with no additional gain from adding initial-state conditioning; channel-wise history injection uses 2,048 weights versus 2,097,152 for dense history.
Perspective
This work targets designers of from-scratch pretrained looped language models, applicable to settings up to roughly 1B parameters, with a fixed token budget and manually specified loop counts; its findings inform choices of non-recurrent boundary-layer allocation, state conditioning, and timestep conditioning under inference budgets below, at, and beyond the training horizon. The combination of history-state injection and timestep conditioning, under channel-wise parameterization, achieves robust gains across budgets with few additional parameters, making it a low-cost starting point for LoopLM conditioning design.
Experiments are limited to models up to roughly 1B parameters under a fixed token budget, extrapolation goes only to a fixed multiple of the training horizon, and loop counts are manually specified rather than adaptive; the relationship between geometry analyses (such as computational interaction) and downstream performance still needs further study. Moreover, the optimal history window and gating type for history-state injection and timestep conditioning vary with configuration and task, and no variant consistently dominates across all configurations and metrics.
