HC-DLM makes a continuous latent the only persistent generative state and beats both discrete and continuous diffusion baselines on Sudoku, Countdown and LM1B
Synopsis
The work proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), in which a continuous latent is the only state persisting across steps while tokens are read out from it at every step and fed back as a scaffold conditioning the next latent update, with a training objective derived from a variational bound on the token likelihood; at matched model size it improves over discrete and continuous diffusion baselines in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B.
Interpretation
HC-DLM organizes the reverse process as a two-level chain: the continuous latent trajectory is the only state persisting across steps, and at each step tokens are read out from the current latent through a readout distribution, re-noised through the discrete forward kernel, and fed back as the condition of the next latent transition. Unlike recent methods that attach continuous context to a self-contained discrete chain, the token state here is itself a variable of the generative model with a forward kernel of its own, carries no transition chain of its own, and every position is read out anew at each step. The paper gives the generative factorization (Eq. 3) and the inference algorithm (Algorithm 2), and states that this organization determines both training and sampling; the loaded text is full text including main body and appendices.
The authors derive a variational lower bound on the token likelihood for this hierarchical coupling chain and turn it into a practical objective: a reconstruction term, an encoder entropy regularizer, and, at every timestep, both a discrete token prediction loss and a continuous denoising loss. The paper notes that the discrete token prediction sub-term has the same functional form as the token-prediction KL used in categorical discrete diffusion, so HC-DLM retains discrete-diffusion supervision while augmenting it with a continuous denoising signal; the continuous part is implemented with conditional flow matching. Proposition 1 states the ELBO and its decomposition, and Appendix A provides the derivation steps: Jensen's inequality, telescoping of continuous terms, the chain-rule KL decomposition, and the practical parameterization.
On Sudoku, Countdown and LM1B, HC-DLM improves over discrete and continuous diffusion baselines at matched model size: 72.41 versus CCDD's 70.73 on the Hard Sudoku split, 37.52 versus 25.35 on Countdown CD5, and a generative perplexity of 75.5 on LM1B, better than all evaluated diffusion models. The paper emphasizes that the advantage is most pronounced beyond the solution strategies represented in training (Hard Sudoku) and widens on the longer-horizon CD5; on LM1B it outperforms all discrete diffusion baselines and both continuous models that decode only once at the end. Results are given in tables, with baseline numbers taken from Kim et al. (2025), Ye et al. (2025) and LangFlow (Chen et al., 2026), while CCDD and MDM are reproduced by the authors at a matched 6M-parameter scale under the same protocol; LM1B reports GPT-2-Large perplexity over 1,024 samples at 128 sampling steps.
Ablations show that both the continuous latent and the discrete scaffold are needed: removing the token level (Latent DM) drops Hard Sudoku accuracy to 24.74 and removing the latent level (MDM) gives 49.88, against 72.41 for the full model; the uniform-state kernel also works clearly better than the absorbing one. An information-emergence analysis indicates that conditioning the continuous denoiser on the discrete scaffold locks in structural constraints sooner, with intermediate-step decoded accuracy rising earlier and reaching a higher plateau; HC-DLM also degrades less sharply than purely discrete MDM as the number of denoising steps is reduced. Ablations are run on Sudoku and include a framework-choice table, a noise-type and token-ordering table, a latent sequence length sweep, and accuracy-versus-steps plus wall-clock comparisons.
Perspective
The result targets discrete sequence generation that needs global constraint satisfaction or bidirectional computation, such as logic puzzles, arithmetic planning and unconditional language modeling; in these settings a reader can use it as a reproducible design in which a shared continuous latent carries cross-token dependence and tokens feed back as a scaffold at every step. The reported experiments are at moderate scale on benchmarks whose correctness can be verified exactly, and the authors note that all modules are standard DiT blocks and the latent channel adds only tokens, so the design carries no scale-specific components and scaling to larger pretrained backbones is mainly a question of compute. What a practitioner can borrow directly is the two-level reverse chain, the three training signals derived from the variational bound, and the alternating inference loop of latent denoising, token readout and re-noising.
Open questions remain: whether the gains from hierarchical coupling hold on larger pretrained backbones and a wider range of tasks, which the paper names as the most promising direction for future work; the latent sequence length shows an empirical optimum (4 works best in the sweep, while 8 and 16 degrade), and how that optimum varies with task and scale is not established; on LM1B the authors follow ELF in not reporting validation perplexity, since likelihood evaluation for flow-based models can require additional likelihood-specific training; and training jointly optimizes an encoder, a continuous denoiser and a token predictor, making each training step more expensive than a purely discrete masked diffusion baseline, an overhead confined to training. In addition, this summary is based on the paper's full text, in which formulas and tables appear as text, so fine typographic details of individual numbers should be checked against the original.
