Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Removing residual connections in the DDT condition encoder and fusing multi-depth features lets DDT-RFE improve four visual understanding tasks and ImageNet FID with fewer encoder blocks

The authors remove the residual connections around Self-Attention and MLP in the condition encoder of the Decoupled Diffusion Transformer (DDT), stabilize training with zero-initialized output gates and a mixed orthogonal attention initialization, and form the encoder output by aggregating the input patch embedding with the outputs of all encoder blocks; under ImageNet-1k training at the stated resolution, patch size 2, 80 epochs and batch size 1024, DDT-RFE improves overall over DDT on image classification, semantic segmentation, object discovery and semantic correspondence while using fewer encoder blocks (12 vs. 20 at L, 14 vs. 22 at XL), and lowers ImageNet generation FID from 8.98 to 7.68 at L and from 7.27 to 6.81 at XL.