Skip to main content
Back to timeline
arXivSource publication:

iADD uses early-timestep updates and Feynman-Kac branching to improve reward alignment, diversity, and rare-prompt success in image and 3D scene generation

Related research and updates

Synopsis

The authors propose iADD, a reinforcement-learning post-training method for discrete-time diffusion models that theoretically shows updating early denoising timesteps preserves diversity better than updating only late timesteps, and combines an incremental sparse-timestep curriculum with discrete-time Feynman-Kac path pruning and branching; across rare-prompt image generation, vanishing-point correction, and 3D indoor scene synthesis, iADD improves reward, Inception Score, rare-event success, and AUC over DDPO, B2-DiffuRL, and related baselines, and its components compose with GRPO-style optimization.

AI-generated editorial illustration: iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

Interpretation

The paper theoretically analyzes how timestep selection affects diversity, with propositions showing that under a fixed optimization budget, uniform timestep sampling yields a strictly greater total perturbation magnitude than a late-stage strategy, and that sparse updates preserve the prior distribution's probability volume better than full updates. Prior work such as B2-DiffuRL argued for updating only late timesteps; this paper provides the opposite theoretical prediction and actionable curriculum design principles. Propositions 1 and 2 with proofs in the supplement; direct training-time measurements show cumulative output perturbation remains larger for uniformly sampled timesteps than for late-only updates, and at matched compute sparse updates improve the full-generator log-volume measure by several nats (paired difference with SEM across 24 prompt-seed pairs).

The paper proposes iADD, which integrates an incremental sparse-timestep curriculum with discrete-time Feynman-Kac branching, pruning low-reward-expectation paths to improve alignment and branching promising ones to encourage diversity. Prior Feynman-Kac guidance was mainly used for inference-time sample improvement; this work devises it for training and combines it with a timestep curriculum. Experiments on image generation, vanishing-point correction, and 3D indoor scene synthesis, plus ablations for each component; ablations show naive combinations such as branching with late-timestep updates or FK alone improve IS only modestly, whereas the complete incremental FK-branching recipe produces a substantially larger gain.

On image generation, iADD improves reward, IS, rare-event success, and AUC over DDPO and B2-DiffuRL; on vanishing-point correction the reward and rare-event success gains are especially large; on 3D indoor scenes the collision rate is lower than baselines. Baselines tend to sacrifice IS or increase collisions when raising reward, whereas iADD maintains better diversity, rare-event success, and physical plausibility at the same reward target. Table 3 reports image generation reward 0.3553, IS 1.3241, rarity 66.85%, AUC 77.7%, versus DDPO's 0.3417, 1.2786, 24.81%, 53.8% and B2-DiffuRL's 0.3452, 1.2728, 24.22%, 58.7%; vanishing-point reward 0.7763 and rarity 95.11%; 3D collision 59.20% versus DDPO's 63.39% and B2-DiffuRL's 64.06%.

iADD's trajectory guidance and sparse curriculum compose with GRPO-style optimization, further improving CLIPScore and rare-event success. This indicates the method is not tied to a specific DDPO objective but is complementary to the underlying policy optimizer. Table 4 shows iADD+GRPO at CLIPScore 0.4126, rarity 88.33%, LPIPS 0.7778, above DanceGRPO's 0.3886, 61.67%, 0.7624 and BranchGRPO's 0.3854, 56.67%, 0.7447.

Perspective

The work targets settings that use discrete-time diffusion models and online reinforcement-learning fine-tuning with sparse terminal rewards, applied to image generation, vanishing-point correction, and 3D indoor scene synthesis; the method builds on DDIM sampling and LoRA fine-tuning, using Feynman-Kac particle sampling and an incremental timestep curriculum during training. For researchers and practitioners who want to improve reward alignment while preserving diversity without additional preference data, the results offer reusable timestep-selection principles and trajectory-guidance components, and show these components can be composed with GRPO-style optimization.

The theoretical propositions rest on idealized optimization and first-order sensitivity assumptions; the authors state they establish comparative predictions rather than exact training-time bounds, and training-time measurements test these predictions under nonlinear fine-tuning dynamics. Reward estimation from noisy intermediate states can be unreliable at earlier timesteps, and in 3D scenes geometric attributes such as object size and position are often not accurately recoverable midway through denoising; the authors suggest future work on more accurate terminal-reward estimators. Training cost is higher than DDPO, mainly due to Feynman-Kac particle sampling, although the incremental timestep strategy partially reduces it. The CLIP reward has known failure modes, illustrated by the prompt "chicken playing chess" potentially scoring high for images with weak semantic alignment. These are scope and open questions rather than faults of the work.

Sources