Turning the residual stream from passive summation into active retrieval: mirror routing cuts DiT-XL/2 FID from 18.85 to 15.07 and REPA-XL/2 from 5.9 to 4.34
Synopsis
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
Interpretation
The paper systematically analyzes routing inside DiTs and observes two behaviors that deviate from standard transformer behavior: early-layer representations receive high importance across the entire network depth (encoder-side dependency), and when given the freedom to route across layers the model spontaneously prefers mirrored/symmetric layer pairs. Symmetric skip connections in U-Nets were largely treated as a hand-crafted inductive bias, and prior work such as U-ViT and U-DiTs reintroduced long skips into transformers with static projection fusion; here symmetry is presented as a structural tendency the model itself seeks out under learnable routing. Based on visualization of routing weights across all layers (Figure 4) and CKA representation similarity analysis (Figure 5(a)-(b)); this is observational evidence at the internal-representation level rather than downstream performance evidence.
It proposes structured residual routing: the DiT is divided evenly along depth into encoder and decoder phases, encoder-side representations (including the patch embedding output) serve as differentiable residual sources that stay attached to the computation graph, and each decoder sublayer uses softmax scoring to select a single mirrored source from the candidate set for adaptive fusion, inserted before both the self-attention and MLP sublayers. It shares the softmax source-scoring mechanism with Attention Residuals in LLMs but restricts the source pool to encoder-side representations rather than all preceding layer outputs, and each decoder layer retrieves a single stage-matched mirror source instead of routing densely over the pool; it resembles the mirror pairing of U-ViT/U-DiTs but its fusion weights adapt to token representations, spatial positions, sublayers, and diffusion timesteps rather than being a static projection. Ablation on DiT-S/2: over the same mirror pairs, adaptive fusion (Fused-Mirror w/o Enc., FID 64.83) beats a static skip (U-ViT-style skip, FID 67.90, which also uses about 1.476M more parameters); adding encoder-internal routing further improves FID to 62.48.
On ImageNet class-conditional generation, the method consistently improves DiTs of different scales under an identical training protocol, with relative FID reduction growing with model size and added parameters staying below 0.1%. At 400K iterations, DiT-S/2 goes from 69.81 to 62.48, DiT-B/2 from 44.21 to 36.33, and DiT-XL/2 from 18.85 to 15.07; across matched-FID comparisons the method requires up to 1.73 times fewer training iterations, and this advantage tends to grow as training progresses. Controlled comparison in Table 2, with the same training protocol across the three scales; DiT-XL/2 reaches 11.40 versus the baseline's 13.71 at 800K iterations, while the baseline reaches 11.76 at 1300K.
The method works as a fine-tuning module on top of an already strong pretrained model: REPA-XL/2 improves from 5.9 FID to 4.94 after 0.28M fine-tuning steps and to 4.34 after 0.35M steps, whereas matched continued fine-tuning yields essentially no gain (5.87 and 6.01). This indicates the late-stage bottleneck is not simply insufficient model capacity but leaves room in connectivity and gradient structure; on efficiency, replacing dense routing with a single stage-matched source cuts decoder-side source-stack activation from 1.51 GB to 289 MB. The REPA fine-tuning experiment starts from the released 4M-iteration checkpoint, adds only the routing modules, and uses the denoising objective without the REPA projection loss; the continued fine-tuning baseline uses an identical setting without routing. With classifier-free guidance the model reaches 1.39 FID.
Perspective
The result targets diffusion transformers trained on ImageNet class-conditional generation under latent diffusion protocols, applying to from-scratch DiT-S/2, DiT-B/2, and DiT-XL/2 as well as REPA-XL/2 fine-tuning from a released checkpoint. The method is implemented with lightweight normalization and projection layers and can be layered onto an existing model as a module, so the cost of changing an existing training pipeline is low. The paper positions mirror routing as a structural prior: it directly supplies the symmetric pairing the model already prefers under dense routing, reducing routing ambiguity and simplifying optimization. For a reader, this offers a reusable idea, namely reorganizing cross-depth information and gradient pathways to unlock the expressive potential of an existing model while keeping the DiT single-scale architecture unchanged, without adding model capacity.
The paper itself notes that broader validation on larger text-to-image and text-to-video models remains future work, since those models involve more complex conditioning mechanisms, larger-scale datasets, and different training recipes that may affect the optimal routing structure; it is also unclear whether mirror routing is always the best choice once stronger multimodal conditioning or temporal dependencies are introduced. In addition, the gradient alignment induced by mirror pairing is partly caused by the mirror connections themselves, and the paper presents it as a characterization of the optimization behavior its routing induces rather than an independent causal conclusion. A reader might also watch how routing weights vary with diffusion timestep, how mirror routing behaves at greater or non-uniform depth divisions, and how the method interacts with other representation-enhancement strategies, none of which the main text develops.
