Skip to main content
Back to timeline
arXivSource publication:

Deferring particle expansion to late transformer layers and adding time-dependent scoring-rule schedules lets a DDM trained from scratch in one stage reach 4.48 FID at 4 steps and 2.38 at 50 steps on ImageNet-256²

Synopsis

To address two obstacles in scaling Distributional Diffusion Models (DDMs)—multi-particle training overhead that grows with the number of particles, and globally fixed scoring-rule hyperparameters that force a single trade-off across sampling budgets—the work defers particle expansion to late transformer layers and introduces time-dependent scoring-rule schedules informed by the dynamical regimes of Biroli2024; combined with a DiT-based latent setup, this makes DDM training practical on class-conditional ImageNet-256², reaching 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2 from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs, and the same recipe transfers to text-to-image generation.

AI-generated editorial illustration: Improved Distributional Diffusion Models

Interpretation

The work defers particle expansion to late transformer layers, mitigating the multi-particle training overhead that scales with the number of particles. Previously, DDM multi-particle training overhead scaled with particle count, one of the two obstacles to scaling DDMs to modern image-generation settings; moving particle expansion to late layers changes how that overhead is distributed. An architectural change stated at the abstract level, supported by training results on ImageNet-256² with a DiT-based latent setup; the loaded text does not provide layer-by-layer ablation detail for this change.

The work introduces time-dependent scoring-rule schedules, replacing globally fixed hyperparameters with ones that vary with the sampling budget, mitigating the single-trade-off problem. Previously DDMs used globally fixed scoring-rule hyperparameters, forcing one trade-off across all sampling budgets; time-dependent schedules allow different trade-offs at different times, motivated by the dynamical regimes of Biroli2024. A methodological change stated at the abstract level, with non-degrading FID across 4 to 50 NFE as indirect evidence; the specific schedule form and ablations are not expanded in the loaded text.

With a DiT-based latent setup, the recipe reaches 4.48 FID at 4 steps and 2.38 at 50 steps on class-conditional ImageNet-256², from a single model trained from scratch in one stage. DDMs previously struggled to scale to modern image-generation settings; this result indicates usable few-step generation quality on ImageNet-256² without a teacher model, self-distillation or JVPs. Specific FID numbers and training conditions given in the abstract (DiT-XL/2, one stage, from scratch, no teacher, no self-distillation, no JVPs); the loaded text does not provide full experimental tables or statistical detail.

The resulting generator is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Few-step generation methods commonly behave differently as sampling steps increase; here the reported behavior is non-degrading FID with increasing NFE, plus evidence of cross-task transfer. Abstract-level result statements covering the 4 to 50 NFE range and text-to-image transfer; specific text-to-image metrics are not given in the loaded text.

Perspective

The result is framed for the DiT-XL/2 latent setup on class-conditional ImageNet-256², trained in one stage from scratch without a teacher, self-distillation or JVPs; the same recipe is reported to transfer to text-to-image generation. For a reader, this means few-step stochastic generators can be trained without a distillation pipeline, suited to settings where usable image quality at few sampling steps is wanted and extra training-time cost is acceptable.

The loaded text is abstract- and metadata-level evidence, lacking full experimental tables, layer-by-layer ablations and schedule details, so the concrete implementation of the time-dependent scoring-rule schedules, the overhead curve from deferring particle expansion to late layers, and the specific text-to-image transfer metrics remain open questions a reader would need the original to confirm; moreover, the 4.48 FID at 4 steps and 2.38 at 50 steps are reported for the specific DiT-XL/2 class-conditional ImageNet-256² setting, and generalization to other resolutions, datasets and backbones remains to be seen.

Sources