Skip to main content
Back to timeline
arXivSource publication:

DRM replaces the scalar reward head with a diffusion head: matching or beating same-data, same-backbone baselines on five benchmarks while its reward distribution turns multimodal as human disagreement grows

Synopsis

The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.

AI-generated editorial illustration: Diffusion Reward Models

Interpretation

The paper identifies a shared limitation across scalar, multi-attribute, and parametric-distributional reward models: all commit to a fixed output-distribution family, whereas human preference is itself multimodal; it therefore recasts reward modeling as density estimation over the conditional reward distribution. Where DPL and URM predict a Gaussian, DPRM a categorical distribution, and QRM a fixed quantile grid, DRM imposes no parametric form on the output shape and represents the distribution implicitly through iterative denoising. This judgment rests on the paper's survey of existing reward-modeling paradigms and its citations of the multimodal-preference literature; it is an argument at the problem-formulation level rather than a single experimental result.

On top of a frozen FsfairX-LLaMA3-RM-v0.1 encoder, DRM replaces the original scalar head with an approximately 12.0M-parameter Diffusion Transformer head; the same architecture trains on multi-attribute regression with a masked denoising loss and on pairwise preference data with a distributional Bradley–Terry objective. Prior multi-attribute reward models require a predefined attribute schema and collapse a vector through a gate, while generative judges shift the burden to long-form reasoning; DRM covers both supervision regimes with a single diffusion head and, at inference, samples an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. The paper provides a full description of the architecture, training objectives, and inference procedure, and reports two variants, DRM-Multi-8B and DRM-Pref-8B, that share the same architecture and diffusion schedule and differ only in reward dimension and a few optimization settings.

Across five benchmarks (RewardBench v2, PPE, RMB, RM-Bench, JudgeBench), DRM-Multi-8B averages 66.2, exceeding the same-data, same-backbone scalar head ArmoRM (62.3) and the parametric-quantile head QRM (64.1), with the largest gain on RMB (78.0); DRM-Pref-8B averages 65.8. In the controlled comparison where data and backbone are held fixed, only the reward head changes, so the paper attributes the improvement to the diffusion head itself; in absolute terms DRM also remains competitive with larger discriminative, distributional, and generative reward models. The controlled comparison uses the same ArmoRM corpus and the same frozen backbone, which the paper explicitly describes as the most controlled setting; the cross-family comparison involves different data and larger backbones and is a competitiveness demonstration rather than a controlled conclusion.

On HelpSteer2-Disagreements and MultiPref, which contain repeated human annotations, DRM's sampled distributions are closer to empirical human rating distributions (Wasserstein distance 0.804 on helpfulness, versus 1.030 for the empirical prior, 1.032 for a global Gaussian, and 1.032 for the pointwise baseline), and its outputs become more frequently multimodal as human disagreement grows (37.6% to 63.2% on helpfulness). Prior reward models output only a point estimate or a fixed-family distribution and cannot express multiple defensible judgments for the same input; DRM demonstrates a correspondence between input-conditional distributional modeling and the structure of human disagreement. The conclusion rests on distributional distance metrics over repeated-annotation datasets and on multimodal-ratio statistics grouped by disagreement level; the paper also notes that its single-example qualitative illustration is only illustrative.

Perspective

This work targets reward modeling for general-domain alignment and suits readers who want to handle multi-attribute annotations and pairwise preferences within one architecture and who need uncertainty or risk-sensitive aggregation. It enables follow-up work to explore distributional reward heads on top of a frozen encoder at modest training cost (the paper reports about 1.54 GPU-hours to train the RewardDiT checkpoints) and to use reward distributions for selective prediction, LCB ranking, and RLHF training signals. The paper positions itself as a first exploration of this new paradigm, and its conclusions apply to the reported data scale, the single encoder used, and the chat, instruction-following, math, code, factuality, and safety tasks covered by the five benchmarks.

The paper states that DRM is trained on only a few hundred thousand open-source samples, far below the tens of millions of curated preferences used by industrial reward models, and that it does not yet match the strongest open-source scalar reward models in absolute terms; the study fixes a single encoder and a fixed training-data scale, so behavior across backbone sizes, model families, and larger data regimes remains to be characterized. In addition, like other reward models, DRM may inherit biases from its training data and could assign high scores to outputs that are stylistically persuasive but factually incorrect or socially harmful; if used as an optimization objective without human oversight or safety constraints, it may amplify such biases. Readers should also note that in the text underlying this summary some table values appear as placeholders, so the complete numbers in a few appendix tables cannot be verified, and the relevant conclusions rest on the values stated explicitly in the main text.

Sources