RWTD lifts a one-step generator's GenEval from 0.73 to 0.80 via reward-weighted transport distillation
Synopsis
The work introduces Reward-Weighted Transport Distillation (RWTD), a post-training method for one-step generators that needs only generated samples and scalar reward evaluations, builds an adaptive target mixing separately reward-tilted current and reference distributions, realizes it through feature-space optimal transport and fixed-point regression, and raises the one-step SANA Sprint 1.6B GenEval score from 0.73 to 0.80 while showing cross-reward generalization in preference alignment experiments.
Interpretation
RWTD builds an alignment target that evolves with training: Qβ,ρ[pθ]=(1−ρ)Tβ,r[pθ]+ρTβ,r[pref], mixing separately reward-tilted current and reference distributions weighted by the reference mass ρ. Earlier one-step alignment methods mostly align to a fixed reward-tilted reference distribution, or require differentiable rewards, likelihood surrogates, or multi-step denoising trajectories; the authors state RWTD is the first one-step alignment method to explicitly target a mixture of separately normalized reward-tilted current and reference distributions. The paper gives a particle realization: separately normalized softmax reward weights over current and reference particles (Equation 6) form the empirical target measure, requiring only generated samples and scalar rewards, with no likelihoods or reward gradients.
The target is realized by feature-space optimal transport plus partial transport regression: an entropic OT coupling is computed with Sinkhorn in the frozen encoder's feature space (DINOv2 for SANA Sprint, MAE for SDXL Turbo), each source feature is moved partway toward its conditional mean by step size η, and the result is distilled back into the generator with a stop-gradient feature regression loss. Unlike direct reward backpropagation, Stein variational gradient descent with kernel density estimation, or rank-based dipole preference fields, RWTD uses transport geometry to produce sample-specific assignments and requires no reward gradients, no KDE, and no multi-step diffusion teacher. Algorithms 1 and 2 give the full procedure; the transport-target ablation shows barycentric and sampled variants perform similarly, while independent reward-weighted regression (RWR) without transport geometry degrades CLIP and PickScore relative to the base model, which the authors read as evidence that the benefit comes mainly from the sample-specific assignments induced by feature-space OT.
The theoretical analysis characterizes the mixed reward-tilting fixed points: for 0<ρ≤1 fixed points correspond one-to-one with solutions of H(λ)=1/ρ and are unique, taking the form p∞(x)=ρ/Zref·pref(x)exp(βr(x))/(1−λexp(βr(x))); at ρ=1 this recovers the standard reward-tilted reference, and decreasing ρ increases λ and sharpens amplification of high-reward regions, while the purely on-policy endpoint ρ=0 has no unique fixed point and concentrates on reward-maximizing regions. The authors note that mixed targeting is not equivalent to interpolating model parameters or merely changing the reward temperature; it defines a distinct fixed-point family with a reward-dependent rational amplification factor. Proposition 5.1 provides the derivation and uniqueness argument, and Appendix A.2 gives existence conditions plus the implicit-differentiation result dλ/dρ<0; the authors explicitly state these results operate under idealized conditions and provide intuition rather than strict guarantees.
Empirically, RWTD raises the one-step SANA Sprint 1.6B GenEval score from 0.73 to 0.80, surpassing comparable-scale multi-step models (Playground v3 0.76, FLUX.1 Dev 0.66, SD3.5-L 0.71), and improves SDXL Turbo from 0.55 to 0.61; trained with HPSv2 it is the only method to improve all five preference metrics and the out-of-domain GenEval score (0.75, versus 0.72, 0.72, and 0.62 for FAV, DrPO, and DRaFT). On SANA Sprint 1.6B, DrPO and FAV reach only 0.75 and 0.73; gradient-based methods optimize HPSv2 more strongly but push CLIP below the base model, which the authors describe as stylization and semantic degradation consistent with reward overoptimization. GenEval uses the official codebase with four images per prompt; preference alignment evaluates on PartiPrompts with five images per prompt reporting PickScore, HPSv2, Aesthetics, CLIP, and ImageReward; training uses LoRA on 8 GPUs for SANA and 4 for SDXL Turbo, with GPU-hour costs reported (GenEval: RWTD 197 versus DrPO 194 and FAV 431).
Perspective
The result targets research and engineering settings where an existing one-step generator must be aligned to a new scalar reward, including non-differentiable rewards, provided a reference generator, a frozen feature encoder, and samplable reward evaluations are available; the authors emphasize that the method assumes no likelihoods, score functions, reward gradients, or multi-step denoising trajectories, making it suitable for black-box rewards and implicit generators. The theory characterizes fixed points and transport-step invariance at an idealized population level, while the practical version approximates those dynamics with finite minibatches, entropic OT, and barycentric projections. Experiments cover two backbones (SDXL Turbo and SANA Sprint 1.6B), a compositional GenEval reward and the HPSv2 preference reward, and report GPU-hour costs so readers can judge feasibility on their own hardware.
The authors state that the theoretical results operate under idealized conditions and offer intuition rather than strict guarantees, while practical training approximates them with finite minibatches, entropic OT, and barycentric projections. The hyperparameter sweeps show that increasing the reward temperature β raises HPSv2, Aesthetics, and ImageReward while CLIP gradually declines, and that purely on-policy training at ρ=0 substantially reduces diversity (average pairwise LPIPS falling from 0.652 to 0.414), so the β and ρ trade-off remains an open question to weigh per setting. All aligned models remain below the base model in diversity, and reference mixing mitigates rather than removes that loss. Experiments are also confined to image generation and image reward models, leaving behavior on video, 3D, or other modalities and more complex rewards to be tested.
