Under matched compute on Dream-7B and LLaDA-8B, ORM reranking beats deterministic PRM guidance by up to 82.71% versus 70.02%
Synopsis
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
Interpretation
Under a shared forward-pass budget, deterministic PRM guidance loses to ORM reranking on math and code: on GSM8K, 75.13% versus 65.18% at N=8 and 82.71% versus 70.02% at N=32; on MATH, 30.65% versus 20.80%; on MBPP, 63.04% versus 50.88%. PRM guidance had been effective for autoregressive reasoning, but dLLM snapshots are scattered visible positions rather than prefixes; this work is the first matched-compute comparison of guided versus reranked test-time compute in that setting. Evaluated on the GSM8K test set with Dream-7B and strict answer extraction, a best-of pool of 42,208 complete trajectories, forward-pass counts validated against wall-clock ratios within about 10%, and paired bootstrap confidence intervals.
The gap decomposes into pool damage and terminal selection: PRM ROC-AUC falls monotonically with mask ratio, top-1 pruning cuts the Oracle ceiling from 81.05% to 67.30%, and an SMC sampler that restores the ceiling to 77.89% still selects at 65.48%. The decomposition turns 'does guidance help' into two independently measurable stages, with an offline counterfactual showing a top-1 cut at the initial state removes every eventually-correct lineage in 46.06% of cases. Combines mask-bucket ROC-AUC curves, the PRM Hybrid Oracle ceiling, an ESS-tempered SMC control, and a step-resolved pruning-risk table, each with 95% confidence intervals.
On one shared candidate pool, a PRM trained only on final states matches ORM reranking (75.40% versus 75.13%, and 82.79% versus 82.71%), while the cross-mask PRM trails badly (42.84% and 65.35%); on MBPP the PRM already ranks finished programs on par with the ORM, so its guided shortfall there is the cost of guidance itself. Shows that cross-mask training spreads the scorer's discrimination thin, and that high pooled ROC-AUC does not guarantee within-problem top-1 ranking. Same-pool reranking controls, a task-specific MBPP verifier control, and formal propositions separating pooled discrimination from within-problem ranking.
Bidirectional PRMs beat causal PRMs, and most of that gap comes from the readout: last-token pooling raises the causal PRM's final-state ROC-AUC from 0.6789 to 0.7274, closing about half the gap; on LLaDA-8B-Base bidirectional guidance averages 31.64%, about 10.87 points above causal guidance. Separates architectural attention from readout design, showing that the last-token readout standard in causal reward models is the better choice on diffusion snapshots. Reproduced on two backbones, Dream-7B and LLaDA-8B-Base, with longer-training and ORM-protocol retrain controls; the bidirectional advantage holds in all ten mask buckets.
Perspective
The result speaks to researchers and engineers who use discrete diffusion language models for math and code reasoning and plan to spend extra test-time compute on reward guidance; the setting is deterministic top-k PRM guidance compared against task-matched verifier reranking under matched compute. The released snapshot corpus and evaluation toolkit let others measure pool damage and terminal selection under the same protocol, and the decomposition extends to any search that prunes on intermediate scores and then selects with the same scorer.
The paper lists open questions: domain scope covers only math and code, with creative or open-ended text untested; an LLaDA-specific ORM comparison has not been run; supervision uses binary final-correctness labels only, and whether step-level supervision changes the signal-decay curve is unclear; readouts tested are mean pooling and last-token, leaving [CLS] and learned query pooling untested; guidance algorithms cover only the segmental top-k pruning family, with adaptive branching, lookahead scoring, and stochastic beam search left for follow-up; and sampler sensitivity of the snapshot distribution remains an empirical question. The mechanism behind the bidirectional advantage is also not uniquely identified: the paper proposes joint unordered modeling, readout design, and training-distribution shift as residual hypotheses, and suggests a permutation-LM scorer as the falsification test.
