REST adds four differentiable losses to latent thoughts, lifting accuracy by up to 7.5 points across 7 benchmarks
Synopsis
The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
Interpretation
The paper identifies four failures of CE-only training of latent thoughts, namely substitution divergence, information loss, thought collision and producer uncertainty, each of which lowers the probability of the correct answer. Prior work defined supervision against the text a thought replaces, as CODI does by matching hidden states induced by that text and SIM-CoT by decoding the text back from the thought; here the problem is located in the relation between thoughts and in the output distribution itself. The paper provides theoretical analysis and observes, in trained systems, CE-only thoughts collapsing across distinct questions and retaining input content irrelevant to the output.
REST writes each of causality, minimality, separability and stability as a differentiable loss term added to cross-entropy, with no architectural change and no added parameters at inference. The authors describe it as the first to turn the theoretically motivated properties of a valid thought representation into loss terms for a latent recursive LLM system, and they extend that recursive construction from the multi-agent setting to a single agent that recurs on its own hidden states. Each term is derived from the property definition, with appendix proofs relating the surrogate to the axiom up to bounded error; training updates only the outer link while base LLMs and the inner link stay frozen.
Across 7 benchmarks, with the same training data, compute and latent budget, REST improves accuracy over the CE-only baseline in both the single-agent and multi-agent settings. Averaged over configurations it gains 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, with best configurations reaching 7.5 and 6.5 points; in the Light system, causality or minimality matches or surpasses the strongest frozen agent on all three mathematical benchmarks. Results span Light and Scaled systems built from Qwen, Llama and Gemma models, averaged over three training seeds, and are compared against CODI and SIM-CoT adapted as auxiliary loss terms.
REST thoughts encode more of what is needed to reproduce the producer's output and decode back to the agent's intended output more faithfully, making latent communication easier to audit. CE-only thoughts encode relatively more of the producer's input than its output, whereas REST reverses this, and the effect grows further along the chain of latent thoughts. The paper reports that REST decodes 15.4% more tokens on average and reaches a 95% boxed-answer rate against 73% for CE; a case study finds the saved tokens are repeated text, while extra tokens in the multi-agent setting are genuine derivation.
Perspective
The work targets researchers and engineering teams already running latent recursive or multi-agent latent communication systems, in a setting where base models are frozen and only the link that transfers thoughts between agents is trained, with training data, compute and latent budget matched to the baseline. It makes the thought representation itself a trainable object and makes decoded thoughts easier to inspect, which matters for systems that need to audit inter-agent communication. The paper also extends the recursive construction from the multi-agent setting to a single agent that recurs on its own hidden states, indicating that the same loss terms apply under both topologies.
The authors note that computational cost prevented sweeping all combinations of properties and hyperparameters, and that a cheap approximation to stability was used; training updates only the outer link, leaving extension to the inner link or the agents themselves to future work. The representation analyses inspect one representative run per arm rather than the full seed and weight grid, so how robust the geometric conclusions are still needs more runs. In addition, the observation that the CE-only baseline loses accuracy on some benchmarks as the latent budget grows while REST holds its level comes from a single run per arm and warrants further watching.
