TAPS adapts recurrent step size to the trajectory, lifting Sudoku accuracy to 91.39% and Maze to 79.90% while cutting the loops needed for matched quality
Synopsis
The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
Interpretation
The paper identifies update scale as a third control axis of recurrent inference, alongside what computation is repeated and how many times, and proposes TAPS: exponential moving averages track persistent progress and centered fluctuation across adjacent updates, and their balance sets the relaxation factor online, over-relaxing when progress dominates and damping when fluctuation dominates. Prior recurrent-reasoning work mainly designs the learned update and the recurrent depth, while standard inference implicitly applies unit-scale updates throughout; this work makes that implicit choice explicit and provides an online scheduling rule. Propositions 1 and 2 show that the temporally averaged gradient of terminal loss with respect to the relaxation factor admits an exact decomposition into persistent-progress and centered-fluctuation contributions; several TAPS instantiations are given (GD, P/S-Sign, Momentum, RMSProp, Adam, BB).
The theoretical analysis gives sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. It connects observable trajectory statistics to task-dependent sensitivity that depends on the target, showing that under proxy–oracle and temporal alignment conditions the controller's chosen sign agrees with the loss-decreasing direction. Theorem 4 provides a one-factor gain bound and Theorem 5 a cumulative gain and loop-speedup result; Appendix B.3 discusses the assumptions and a two-mode fixed-point model shows the conditions can hold exactly in that model.
On Sudoku and Maze, TAPS improves terminal accuracy without retraining and shortens wall-clock time to reach the unit-step baseline's performance; incorporating the progress–fluctuation principle into training yields further gains. Fixed scales such as constant 0.5 or 1.5 fail to improve accuracy and speed consistently across tasks, whereas all six TAPS variants improve on both tasks, indicating the benefit comes from adapting along the trajectory rather than uniformly changing update magnitude. Table 1 reports unit-step accuracy of 89.67% on Sudoku and 78.80% on Maze; after co-design the best controller reaches 91.39% on Sudoku and 79.90% on Maze, with corresponding speedups.
TAPS transfers across recurrent architectures and inference strategies: adaptive exit, hierarchical recurrence, parallel recurrent inference, fixed-point inference, language-model recurrence, and intermediate-layer recurrence. It composes with existing depth control, convergence criteria, and width scaling rather than replacing them. Adaptive exit reaches 91.2% Exact with 316.5 effective updates on Sudoku-Extreme versus 91.1% for FPRM; in fixed-point inference, the accuracy fixed-point inference attains after 1,000 updates is matched within roughly 260; average accuracy improves on Ouro-1.4B, Ouro-2.6B, and Huginn-0125; on Qwen3-4B-Instruct intermediate-layer loops all controllers select a scale above 1.
Perspective
The result targets deployment and training settings that use recurrent or looped inference: at inference time it adapts step size online using only observed trajectory statistics, without target labels and without modifying model weights; during training the progress–fluctuation principle can be added as an auxiliary objective to shape dynamics. It applies to latent-state recurrence (TRM), recurrent language models (Ouro, Huginn), and intermediate-layer recurrence (Qwen3-4B-Instruct), and composes with adaptive exit, hierarchical recurrence, parallel recurrent inference, and fixed-point inference. The theoretical guarantees hold under the proxy–oracle and temporal alignment conditions stated in the text, and the two-mode fixed-point model in the appendix illustrates a class of local dynamics where the conditions hold exactly.
The theoretical conclusions rely on comparison conditions between observable energies and oracle contributions plus temporal alignment conditions, and whether these hold for broader recurrent dynamics remains open; TAPS uses local trajectory statistics as proxies, and whether richer signals yield better control is left for future study; training–inference co-design is an initial step, with deeper training integration and broader architectures still to explore; update frequency shows a non-monotonic accuracy–efficiency trade-off, where coarser control granularity can offset saved controller overhead; strong perturbations to high-level hidden states reduce the benefit of adaptation, indicating reliance on sufficiently reliable update signals; this reading is of the full text, but some table and figure values appear as placeholders, so specific numerical details should be checked against the original tables.
