Skip to main content
Back to timeline
arXivSource publication:

ELF-REG pushes continuous diffusion language models into math reasoning and code generation: 55.96% pass@1 on GSM8K and MATH-500 up from 10.55% to 13.39%

Synopsis

The work scales Embedded Language Flows (ELF) to mathematical reasoning and code generation and introduces ELF-REG, which uses a frozen autoregressive teacher to supervise intermediate denoiser features and to supply a global representation jointly denoised with the response (REPA+REG); evaluated on GSM8K, MATH-500, HumanEval, and MBPP, ELF-REG-L reaches 55.96% pass@1 on GSM8K at 64 NFE and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE, outperforming the evaluated comparable-scale dLMs on GSM8K and code and improving MATH-500 from the ELF-L baseline of 10.55% to 13.39%, while the same task-specific checkpoints support strong low-NFE performance through early-stop without few-step training, reaching 41.21% HumanEval pass@10 at 16 NFE.

Source-provided article image: ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks
Figure 1 ·

Figure 1: ELF-REG improves generation with fewer task-training tokens. (a) A frozen AR teacher supervises intermediate features through REPA and a global REG token denoised jointly with the response. (b,c) ELF-REG-L and ELF-REG-B with early-stop ( ρ = 8 \rho=8 ) show superior performance across NFE, compared with PlaidQ [ Peng et al., 2026b ] , FMLM+ [ Agarwal et al., 2026 ] , S-FLM [ Deschenaux & Gulcehre, 2026 ] , and DBTM [ Tang & Wang, 2026 ] . Code uses pass@10 on MBPP-378; GSM8K uses mean pass@1. (d) Non-padding training token budgets. Appendix B gives more details regarding benchmark mappings, model sizes, NFE accounting, budget estimates, and evaluation differences.

arXiv

Interpretation

It introduces ELF-REG, which improves learning in continuous diffusion language models via representation alignment and entanglement (REPA+REG): a frozen AR teacher supervises intermediate denoiser features and supplies a global representation jointly denoised with the response. Relative to the prior ELF baseline, it adds teacher supervision of intermediate features and joint denoising of a global representation, extending continuous dLM training from purely self-supervised denoising toward guidance by AR teacher representations. Results are reported as pass@1/pass@10 on four benchmarks (GSM8K, MATH-500, HumanEval, MBPP) and compared with evaluated comparable-scale dLMs; the MATH-500 gain from 10.55% for ELF-L to 13.39% for ELF-REG-L provides a direct baseline comparison.

ELF-REG-L achieves concrete results on math reasoning and code generation: 55.96% pass@1 on GSM8K at 64 NFE, and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. Continuous dLMs previously lagged AR LLMs and masked dLMs on challenging reasoning tasks; this result extends the evaluation of continuous dLMs to math reasoning and code generation with comparable numbers. Results are reported as pass@1 on four named benchmarks with the corresponding NFE noted; the text states it outperforms the evaluated comparable-scale dLMs on GSM8K and code.

Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory; at 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10. It decouples low-NFE inference capability from few-step training, indicating that a single checkpoint can serve both high-NFE and low-NFE settings. Supported by the concrete figure of 41.21% HumanEval pass@10 at 16 NFE, compared with recent continuous dLMs of comparable scale.

Perspective

The results target researchers and engineering teams studying the reasoning capability of continuous diffusion language models, in the settings of benchmark evaluation for math reasoning (GSM8K, MATH-500) and code generation (HumanEval, MBPP), and of inference deployment that needs parallel decoding at low NFE. The method presupposes a frozen AR teacher and task-specific checkpoints, so its conclusions apply to these tasks and the evaluated scale range; the early-stop finding offers a path for using the same checkpoint at low NFE.

The loaded text is a summary and does not include ablations, training scale, teacher-model choice, or statistical significance, so how much REPA and REG each contribute and how sensitive the results are to the teacher and checkpoints remain questions that require the full text. In addition, the absolute scores on MATH-500 and HumanEval are clearly lower than on GSM8K, so whether the gap between continuous dLMs and AR LLMs on hard reasoning tasks has substantially narrowed still needs validation across more tasks and scales; the specific MBPP result is not given numerically in the summary.

Sources