Skip to main content
Back to timeline

First systematic empirical study of diffusion LLMs for code generation: 7 diffusion models reach 79.1% on MBPP+ versus 73.3% for the best autoregressive baseline, yet open-source diffusion models still do not consistently match strong autoregressive models

Synopsis

This study empirically evaluates 7 representative diffusion LLMs (including the closed-source Mercury-Coder-Small) against 4 open-source autoregressive baselines on HumanEval/HumanEval+, MBPP/MBPP+, HumanEval-X, LiveCodeBench, RepoQA, RepoBench-C, and SWE-Bench Verified, finding promising but uneven code-generation ability: the closed-source diffusion model leads on most benchmarks (e.g., 79.1% on MBPP+ versus 73.

Source-provided article image: Exploring the Potential of Diffusion Large Language Models in Code Generation
Figure 1a

model generates tokens strictly from left to right, as illustrated in Figure 1a. However, there is an open question of whether the mainstream autoregressive generation paradigm is always the best choice for code generation. In practice, current autoregressive LLMs (AR LLMs) reveal two important limitations that may restrict their use in some code generation scenarios. ❶Next-Token Prediction →Sequential-Decoding Cost. Most AR LLMs generate code strictly one token at a time, with each forward pass typically producing only a single token. Because each forward pass typically advances generation by one token, this sequential dependency can accumulate substantial computation and latency for long outputs. Such sequential latency can be particularly costly for repository-level generation [4, 31, 52] and can reduce responsiveness in interactive coding assistants [29, 46, 54]. ❷Left-to-Right Generation Order → Limited In-Sequence Revision. AR LLMs decode in a strictly left-to-right manner, with each new token conditioned on the preceding context. This process can be less flexible than the non-sequential refinement common in programming

· Page 2

Interpretation

Diffusion LLMs show promising but uneven code-generation ability, with the closed-source diffusion system leading on most evaluated benchmarks and code-oriented open-source diffusion models outperforming general-purpose ones. Prior work lacked a code-generation-centered evaluation spanning multiple diffusion families and autoregressive baselines; this study jointly examines cross-model effectiveness, inference-setting sensitivity, efficiency trade-offs, and long-context behavior. pass@1 on HumanEval/HumanEval+, MBPP/MBPP+, HumanEval-X, and LiveCodeBench, repeated five times with mean and standard deviation; e.g., Mercury-Coder-Small reaches 85.4% on HumanEval, 93.7% on MBPP, and 79.1% on MBPP+, and Dream-Coder-v0-Instruct-7B improves over Dream-v0-Instruct-7B by 19.5 percentage points on HumanEval.

The selected diffusion and autoregressive model groups solve partially non-overlapping problem sets, indicating complementary problem coverage in the evaluated settings. Prior comparisons focused mainly on aggregate accuracy; this study characterizes coverage differences with a Venn diagram and two illustrative success/failure cases. On pooled HumanEval, MBPP, and LiveCodeBench 2410–2504 problems, 367 problems are solved by both groups, 149 uniquely by the selected autoregressive models, and 105 uniquely by the selected diffusion models; the cases are illustrative and the authors explicitly avoid attributing them to the paradigm.

Over the common 1k–8k input range, the three evaluated diffusion models degrade more gracefully than one comparable-scale autoregressive baseline on RepoQA and RepoBench-C. Adds long-context code understanding to diffusion-model evaluation, using both retrieval (RepoQA) and cross-file completion (RepoBench-C). Llama-2-7B-chat-hf drops on RepoQA from 46.2% at 3k to 9.0% at 4k and 1.2% at 8k, and on RepoBench-C from 22.1% to 3.4% EM and 62.8% to 19.2% ES between 4k and 8k, while the three diffusion models maintain similar scores over that interval; the authors note nominal context windows differ, so the finding is model-specific.

Diffusion steps, remasking strategy, generation length, block length, and temperature have configuration-dependent effects on effectiveness and efficiency; confidence-based remasking consistently beats random remasking, while more steps and longer lengths reduce efficiency. Systematically sweeps diffusion-specific inference factors and provides practical configuration guidance, with McNemar tests supporting the remasking conclusion. Confidence-based remasking improves LLaDA-1.5 and Dream-Coder-v0-Instruct-7B by 33.0 and 35.4 percentage points on HumanEval with McNemar p<0.01; reducing DiffuCoder-7B-cpGRPO steps from 512 to 256 roughly halves FLOPs/token (1.17×10^13 to 5.88×10^12) and cuts average decoding time on HumanEval from 39.5s to 19.9s per question.

Perspective

The study targets researchers and practitioners using diffusion LLMs for code generation, covering function-level synthesis, multilingual generation, recent competitive-programming problems, and repository-level long-context understanding in the evaluated settings; its configuration guidance (generation length, diffusion steps, remasking strategy, block length, temperature) applies to the evaluated models and benchmarks and serves as a tuning starting point rather than a universal constant. For teams evaluating or deploying diffusion code models, or designing hybrid autoregressive-diffusion systems, the study offers comparable baselines and efficiency-effectiveness trade-off data.

Almost all evaluated models are under 10B parameters, and the closed-source Mercury-Coder-Small's scale and training pipeline are undisclosed, so its lead cannot be attributed to a single factor; the long-context comparison is not context-window matched, so the finding is model-specific; on SWE-Bench Verified all six open-source diffusion models achieve a 0% resolved rate under zero-shot evaluation, which the authors treat as a boundary case rather than a paradigm conclusion; setting sweeps do not cover all model-benchmark combinations, so trends may be sensitive to the tested configurations; static benchmarks such as HumanEval and MBPP carry potential data-leakage concerns and should be read alongside newer benchmarks such as LiveCodeBench.

Sources