A decision model stops answering in one pass: eight loops of the same layers let a 1.4B model approach a single-pass model with three times the parameters
Lead
Turning a language model pre-trained to loop into a typed decision model and reading it after eight loops raises accuracy from 58.4% to 72.0% on 10,027 test decisions, at 7.7 times the computation of a single pass.
Story
A typed decision model can trade looping the same layers for accuracy instead of adding parameters or generating text. The caller declares the options and the model returns one probability per option; applying the same layers recursively to their own output lets the model revise its hidden state before it commits. Such models previously fixed their answer in a single forward pass, giving every decision the same amount of computation; handling decisions that need several dependent steps meant a larger model or generated reasoning, which gives up the typed contract and adds latency. On 10,027 test decisions from 59 sources, the looped model reaches 72.0%, 13.5 points above a non-looped model of the same shape trained with the same recipe and data, 5.3 points above a newer non-looped model of its size, and 1.8 points below a single-pass model with three times its parameters.
Every loop is read and trained, so one model serves every budget from one loop to eight. Training only the last loop collapses the early loops, whose first loop reaches 35.6% accuracy, whereas training every loop lets the first loop reach 58.4%. Training uses the sum of cross-entropy and the Brier score, and only LoRA adapters and a small per-loop readout are trained, 61M parameters in total, with the backbone frozen; on the 1.4B backbone trained with eight loops, reading after three loops already gives 88% of the gain from loop 1 to loop 8.
What to watch
A next step is a rule that chooses the loop for each item: 11.7% of the items are right at some loop but wrong at loop 8, which bounds any loop-selection rule. Teams making automated decisions can read different loops for accept, escalate and route calls according to their computation budget, and three loops already give most of the gain.
The findings rest on the Ouro family of backbones pre-trained to loop, and most analyses use the 1.4B model, so how far they carry to other and larger looped backbones is an open question. Looping buys accuracy with computation: eight loops take about 3.3 times the GPU time of Qwen3.5-4B and four loops about 1.6 times. The reasoning-depth analysis rests on two program-generated tasks, and the verifier case study uses one dataset and one generator. The suite is in English and assembled from existing datasets, its mix of item types is the authors' choice, and it cannot be excluded that a backbone has seen some of these public datasets during pre-training.
