LLoCoT refines a compact latent workspace with a looped transformer, matching explicit-CoT accuracy on HumanEval and MBPP while cutting time to first answer token by about 36x
Synopsis
The authors introduce LLoCoT, which replaces left-to-right latent generation with a few iterative refinement passes of a shared transformer over a compact latent workspace, then samples latent tokens in parallel from a probabilistic head to condition an autoregressive decoder; on HumanEval and MBPP it matches the accuracy of the explicit-CoT baseline Reasoning SFT and outperforms the base model, answer-only SFT and NF-CoT, while reducing time to the first answer token by about 36x, reasoning-phase latency by about 42x, and raising end-to-end throughput by 9.2%.
Figure 1: Overview of the LLoCoT architecture. A shared transformer repeatedly refines K loop hidden states, with each non-final Gaussian mean providing feedback to the next loop. The final Gaussian head generates all final latent thoughts jointly for autoregressive answer decoding; VAE-derived targets and auxiliary intermediate supervision are used only during training.
arXivInterpretation
LLoCoT changes latent reasoning from generating latent vectors left to right into iteratively refining a compact latent workspace: one shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Most prior latent-reasoning methods retain a left-to-right dependency among latent vectors, while explicit CoT expresses extra computation as an autoregressively generated token sequence; LLoCoT replaces serial thought generation with parallel latent-slot refinement. The design is presented as a framework description with a training objective: training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states.
On HumanEval and MBPP, LLoCoT achieves the highest mean among the evaluated methods, performing on par in accuracy with the explicit-CoT baseline Reasoning SFT and outperforming the base model, answer-only SFT and NF-CoT. This provides evidence of accuracy parity with an explicit-CoT baseline while shifting reasoning computation from serial token generation to parallel latent refinement. The evidence is a mean comparison across two code benchmarks, stated at the summary level without per-benchmark numbers, variance, or sample sizes.
Relative to Reasoning SFT, LLoCoT reduces time to the first answer token by approximately 36x and reasoning-phase latency by approximately 42x, while increasing end-to-end throughput by 9.2%. These efficiency figures tie parallel latent-slot refinement directly to measurable latency and throughput gains rather than accuracy parity alone. The text reports approximate speedups and a throughput change relative to Reasoning SFT, without stating measurement hardware, batch size, or statistical uncertainty.
The design replaces serial thought generation with parallel latent-slot refinement while retaining probabilistic latent modeling and autoregressive answer decoding. It indicates that non-autoregressive latent reasoning can coexist with probabilistic sampling and an autoregressive decoder rather than dropping either. This is a descriptive conclusion about the framework's components and training supervision, forming summary-level evidence alongside the benchmark results.
Perspective
The result targets code-generation settings that have explicit-CoT supervision and care about reasoning-phase latency and throughput: on HumanEval and MBPP, LLoCoT trades accuracy parity with Reasoning SFT for the efficiency gains of parallel latent refinement, fitting deployments that can accept a small number of refinement iterations and keep autoregressive answer decoding. For researchers and engineering teams seeking to reduce serial thought generation while retaining probabilistic latent modeling, the framework offers a reusable design template; because training depends on continuous representations derived from explicit CoT, it applies to tasks that already have CoT annotations or can generate CoT data.
At the summary level, per-benchmark accuracy, variance, and sample sizes for HumanEval and MBPP are not given, nor are the measurement hardware, batch size, and statistical uncertainty behind the roughly 36x, roughly 42x, and 9.2% efficiency figures; key configurations such as the number of refinement iterations, latent-slot count, and training-data scale are also not expanded in the summary. Open questions a reader might track: the stability of parallel latent sampling across tasks and model scales, how likelihood supervision affects latent-state quality, and whether the method preserves the same accuracy-efficiency trade-off beyond code generation.
