Skip to main content
Back to timeline
arXivSource publication:

PerF makes pixel-space diffusion Transformers specialize by feature: FID on ImageNet 256 drops from 1.86 to 1.63 with about 3.6% more parameters

Synopsis

The work introduces heterogeneous refinement, giving different feature groups distinct refinement budgets across depth in pixel-space diffusion Transformers, which spontaneously yields persistent features that encode global structure and active features that encode high-frequency detail; building on this, Persistence Forcing (PerF) and Persistence Guidance (PG) reduce FID on ImageNet 256 from 3.66, 2.36, and 1.86 to 2.81, 1.91, and 1.63 for JiT-B/L/H, and on ImageNet 512 from 1.94 to 1.76 for JiT-H/32.

AI-generated editorial illustration: Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

Interpretation

Heterogeneous refinement alone induces an ordered feature specialization: sparsely refined features predominantly encode coherent global visual structure, while more frequently refined features increasingly specialize toward localized high-frequency details, which the authors call persistent and active features. Prior pixel-space DiTs largely propagate the full hidden representation through every block, giving all features a uniform refinement schedule; this work uses layer-wise width variation (a contraction-expansion schedule) as a way to create heterogeneous refinement histories in pixel-space diffusion, and observes that the resulting roles are not explicitly assigned but emerge through denoising. Decomposing the final RGB prediction into additive contributions of different feature groups shows an ordered coarse-to-fine progression as refinement budget increases; in the frequency domain, the normalized spectral centroid is consistently low for groups with small refinement budgets (about 0-128) and substantially higher for more refined groups, with the same overall trend across several diffusion timesteps. Depth-wise trajectories show a representative persistent group updated only at Blocks 1 and 12, whose RGB readout and feature-space PCA stay stable during the bypass from Blocks 2 to 10 after only the first update.

Heterogeneous refinement alone is not enough to improve generation quality; the specialization must be explicitly exploited, which PerF does by using persistent features as structural context for active refinement (P-to-A conditioning) and amplifying their influence at sampling time through Persistence Guidance. The authors first show a parameter-matched heterogeneous-refinement model does not outperform vanilla JiT (FID 42.7 vs 42.1 at 100 epochs, 33.8 vs 33.1 at 200 epochs), separating 'revealing roles' from 'using roles'; they then turn persistence from passive preservation into an explicit interaction and further into a controllable internal conditioning variable that defines a guidance direction. Ablations start from the heterogeneous backbone: adding P-to-A conditioning improves FID from 42.71 to 40.35 without CFG and from 8.02 to 7.19 with CFG; adding PG further reduces FID to 24.92 and 5.81. P-to-A is implemented with a low-rank bottleneck (default rank 128, about 5.7M extra parameters); raising rank from 128 to 256 changes FID only from 40.35 to 40.06 without PG, but from 24.92 to 22.16 at PG=4.0.

Persistence Guidance supplies an internal structural signal complementary to classifier-free guidance rather than a stronger semantic signal. CFG builds its prediction contrast by removing an external class condition, whereas PG stochastically disables the P-to-A pathway so the same model predicts with and without persistent structural context; the contrast captures how the sample's already-formed global structure redirects ongoing refinement. Across all tested CFG scales, a moderate amount of PG consistently improves FID, with the best region around PG=3.0; applying PG over the full denoising interval gives the best overall performance (FID 16.03 without CFG and 3.12 with CFG, versus 21.05 and 3.37 without PG). The authors read this as PG not merely reproducing stronger semantic guidance.

PerF consistently improves over the corresponding JiT baselines across model scales and resolutions with limited overhead. The gains do not rely on a particular backbone size or resolution: on ImageNet 256 the FID reductions for PerF-B/L/H are about 23%, 19%, and 12% relative to JiT, with higher IS and recall and comparable precision; on ImageNet 512, PerF-H/32 improves FID from 1.94 to 1.76, IS from 309.1 to 335.3, and recall from 0.61 to 0.64. Main experiments train for 600 epochs on ImageNet 256 and 512 with the Heun solver at 50 sampling steps; parameter overhead is +4.6%, +2.6%, and +3.6% for Base, Large, and Huge, with single-forward FLOPs increasing by roughly 3.3%-5.1%. Without any CFG, PerF-B/L/H reach FID 12.53, 5.60, and 3.48, still better than the corresponding JiT values of 25.42, 13.85, and 7.15.

Perspective

The results target class-conditional pixel-space image generation with the JiT backbone, validated on ImageNet 256 and 512 with 600-epoch training and the Heun solver at 50 sampling steps; P-to-A conditioning uses a low-rank bottleneck (rank 128 for PerF-B/L, 192 for PerF-H), and class conditioning and P-to-A conditioning are each dropped with probability 0.1. Main results additionally apply REPA to the persistent feature subspace for the first 100 epochs. The setting is directly usable by researchers and practitioners who want to improve pixel-space DiTs without changing the underlying generation framework; the authors also list studying heterogeneous refinement in broader diffusion architectures and applications as future work.

In the frequency analysis the authors note the ordering is not strictly monotonic at every individual timestep, only that larger refinement budgets shift representations toward higher-frequency content overall; the observation that persistent features show coherent global structure after the first update comes from visualizing a representative feature group, and how general it is remains to be characterized more systematically. Performance degrades again when PG becomes too large, indicating a balance point between guiding active refinement and over-constraining it. Main results also rely on REPA alignment for the first 100 epochs while full-schedule alignment is worse, leaving room to clarify how this training recipe interacts with heterogeneous refinement.

Sources