Public articles linked to the same research event.
arXiv The authors remove the residual connections around Self-Attention and MLP in the condition encoder of the Decoupled Diffusion Transformer (DDT), stabilize training with zero-initialized output gates and a mixed orthogonal attention initialization, and form the encoder output by aggregating the input patch embedding with the outputs of all encoder blocks; under ImageNet-1k training at the stated resolution, patch size 2, 80 epochs and batch size 1024, DDT-RFE improves overall over DDT on image classification, semantic segmentation, object discovery and semantic correspondence while using fewer encoder blocks (12 vs. 20 at L, 14 vs. 22 at XL), and lowers ImageNet generation FID from 8.98 to 7.68 at L and from 7.27 to 6.81 at XL.
The authors remove the residual connections around Self-Attention and MLP in the condition encoder of the Decoupled Diffusion Transformer (DDT), stabilize training with zero-initialized output gates and a mixed orthogonal attention initialization, and form the encoder output by aggregating the input patch embedding with the outputs of all encoder blocks; under ImageNet-1k training at the stated resolution, patch size 2, 80 epochs and batch size 1024, DDT-RFE improves overall over DDT on image classification, semantic segmentation, object discovery and semantic correspondence while using fewer encoder blocks (12 vs. 20 at L, 14 vs. 22 at XL), and lowers ImageNet generation FID from 8.98 to 7.68 at L and from 7.27 to 6.81 at XL.
The authors remove the residual connections around Self-Attention and MLP in the condition encoder of the Decoupled Diffusion Transformer (DDT), stabilize training with zero-initialized output gates and a mixed orthogonal attention initialization, and form the encoder output by aggregating the input patch embedding with the outputs of all encoder blocks; under ImageNet-1k training at the stated resolution, patch size 2, 80 epochs and batch size 1024, DDT-RFE improves overall over DDT on image classification, semantic segmentation, object discovery and semantic correspondence while using fewer encoder blocks (12 vs. 20 at L, 14 vs. 22 at XL), and lowers ImageNet generation FID from 8.98 to 7.68 at L and from 7.27 to 6.81 at XL.
The authors remove the residual connections around Self-Attention and MLP in the condition encoder of the Decoupled Diffusion Transformer (DDT), stabilize training with zero-initialized output gates and a mixed orthogonal attention initialization, and form the encoder output by aggregating the input patch embedding with the outputs of all encoder blocks; under ImageNet-1k training at the stated resolution, patch size 2, 80 epochs and batch size 1024, DDT-RFE improves overall over DDT on image classification, semantic segmentation, object discovery and semantic correspondence while using fewer encoder blocks (12 vs. 20 at L, 14 vs. 22 at XL), and lowers ImageNet generation FID from 8.98 to 7.68 at L and from 7.27 to 6.81 at XL.