Removing residual connections in the DDT condition encoder and fusing multi-depth features lets DDT-RFE improve four visual understanding tasks and ImageNet FID with fewer encoder blocks
Related research and updatesSynopsis
The authors remove the residual connections around Self-Attention and MLP in the condition encoder of the Decoupled Diffusion Transformer (DDT), stabilize training with zero-initialized output gates and a mixed orthogonal attention initialization, and form the encoder output by aggregating the input patch embedding with the outputs of all encoder blocks; under ImageNet-1k training at the stated resolution, patch size 2, 80 epochs and batch size 1024, DDT-RFE improves overall over DDT on image classification, semantic segmentation, object discovery and semantic correspondence while using fewer encoder blocks (12 vs. 20 at L, 14 vs. 22 at XL), and lowers ImageNet generation FID from 8.98 to 7.68 at L and from 7.27 to 6.81 at XL.
Figure 1: Left: Overall architectures of DDT and DDT-RFE (our method). Right: Encoder representations from shallow to deep layers at timestep t = 0.5 t=0.5 . Features are projected into three dimensions using PCA and visualized as RGB images. Our method produces more abstract representations with clean backgrounds and salient objects. More visualizations are shown in Figs. 5 , 6 and 7
arXivInterpretation
Removing residual connections inside the condition encoder and moving multi-level feature aggregation to the encoder output yields better semantic representations with a shallower encoder. DDT lets shallow features accumulate into deeper blocks through residual updates (the expansion in Eq. 8); the authors instead make each block a gated transformation with no identity pathway and aggregate the input patch embedding with all block outputs at the encoder output (Eq. 11). Trained from scratch on ImageNet-1k at the stated resolution for 80 epochs with batch size 1024 at L (458M) and XL (675M); linear-probe top-1 is 60.9 (L) and 62.4 (XL) versus 59.6 for DDT-L and 61.8 for DDT-XL, with encoder blocks reduced from 20/22 to 12/14.
The residual-free encoder needs a matching initialization to converge stably, and the proposed mixed orthogonal attention initialization has fewer hyperparameters to tune than mimetic initialization. Mimetic initialization requires tuning two scalar constants; the authors instead mix a shared random orthogonal matrix with the default query/key initializations via a coefficient (Eq. 13), and initialize MLP weights with the SUO distribution so layers preserve input vector RMS with condition number 1. Ablation shows that removing this initialization raises FID from 7.68 to 9.86, drops classification top-1 from 60.9 to 38.6, drops segmentation mIoU from 57.02 to 30.62, and drops correspondence PCK@0.1 from 37.24 to 13.89, which the authors attribute to representation collapse from network conditioning issues.
Aggregating multi-depth features at the encoder output is critical for both generation and understanding, and it narrows the semantic gap between the encoder output and the best intermediate block. DDT's residual accumulation already mixes multi-level information into a single block, whereas DDT-RFE separates the two: abstraction arises within the hierarchy and multi-scale context is recovered by aggregation. The top-1 drop from the best block to the encoder output is 7.8 and 4.7 percentage points for DDT-L and DDT-XL, versus only 1.2 and 1.1 points for DDT-RFE-L and XL; multi-scale probing benefits DDT-RFE more on ADE20K (+2.9 vs. +1.8 mIoU at L), and with multi-scale features DDT-RFE leads on every dataset at both scales.
Better representation quality and better generation quality appear together rather than trading off. Under the same training budget, DDT-RFE attains the lowest FID among the compared models at both L and XL, and still improves over a same-depth DDT-L (12en12de, FID 8.82). Each model generates 50,000 samples with EMA weights and a 250-step Euler ODE sampler without classifier-free guidance; at XL, FID is 6.81 vs. 7.27, sFID 5.50 vs. 5.89, precision 0.740 vs. 0.724, IS is equal (125.6), and recall is slightly lower (0.610 vs. 0.624).
Perspective
The result targets researchers and engineers training DDT-style models from scratch on ImageNet-1k at the stated resolution, patch size 2, 80 epochs and batch size 1024, and applies to settings where a shallower condition encoder should serve both generation and downstream visual understanding. Reusable elements include removing Self-Attention and MLP residuals inside encoder blocks, zero-initializing output gates, the mixed orthogonal attention initialization and SUO MLP initialization, and aggregating the input patch embedding with all block outputs at the encoder output. For evaluation, classification uses DINO-style kNN and a 30-epoch BatchNorm plus linear probe, object discovery uses LOST with CorLoc, semantic correspondence reports PCK@0.1 on SPair-71k, segmentation uses a linear probe with mIoU, and generation reports FID/sFID/IS/precision/recall from 50,000 samples with EMA weights and a 250-step Euler ODE sampler.
The authors note that compute limits restricted experiments to ImageNet at the stated resolution for 80 epochs, leaving longer training schedules, higher resolution, and other diffusion architectures to future work. The ablation shows that removing feature fusion causes large drops, but additionally removing the output gates partially recovers performance, indicating an interaction between gating and fusion that is not fully separated. DDT-RFE is also 0.3 mIoU below DDT on ADE20K single-block probing at XL and ties on COCO-Stuff, which the authors attribute to scene complexity and the need for multi-scale information; how far that explanation extends to more complex datasets remains open. Some equations, table values, and figure captions are incomplete in the parsed text, so exact hyperparameters and per-item numbers should be checked against the original figures and tables.
