Skip to main content
Back to timeline
arXivSource publication:

FuseReg trains decoders and DiTs on random encoder-layer subsets, letting one decoder reconstruct from full, sparse, and single-layer fusions and lowering gFID with the generator unchanged

Synopsis

The work introduces FuseReg, which replaces heuristic layer selection in representation autoencoders (RAEs) with training over random subsets of encoder layers: on ImageNet-256 with a frozen DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining and achieves higher PSNR than decoders specialized to fixed fusions, while decoder replacement alone (with an unchanged RAEv2 DiT-XL generator) reduces unguided gFID and joint regularization of decoder and generator further reduces unguided gFID on DiT-Base.

AI-generated editorial illustration: FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Interpretation

Introduces FuseReg, a layer-fusion regularizer that trains the decoder and generator on normalized fusions of randomly sampled encoder layers, so downstream models are no longer bound to one fixed fusion. RAEv2 commits to a fixed heuristic fusion by summing the last encoder layers, and DRoRAE and attentive multi-layer probing learn deterministic fusion weights; FuseReg does not search for a better single fusion but instead keeps downstream models effective across a family of fusion configurations. The paper provides the method construction and a theoretical analysis, plus controlled experiments on ImageNet-256 with a frozen DINOv3-L over 23 candidate layers; decoder variants are matched in architecture, parameter count, optimization hyperparameters, and training-step budget.

Proves that normalized random-subset sampling preserves the full-layer mean while turning cross-layer disagreement into an explicit regularization penalty, and establishes a second-order separation from any deterministic global fusion. The authors place fixed subsets and learned global gates in a common fusion-weight formulation to highlight their shared limitation, then show that subset sampling varies weights in every layer-contrast direction with weight moments no deterministic global fusion can match. Mean preservation, covariance identities, and the full-support bound do not assume a linear consumer; the objective decompositions and closed-form optimizers assume homogeneous linear predictors and squared loss, which the authors describe as mechanism-level intuition rather than quantitative prediction for nonlinear models.

One FuseReg decoder reconstructs reliably across full, sparse, and single-layer fusions, whereas fixed-fusion RAEv2 decoders degrade sharply off their training fusion. A fixed-fusion decoder specializes to its training fusion: RAEv2K=23 reaches 27.04 dB and rFID 0.18 on its own fusion but drops to 12.51–14.31 dB on the others; FuseReg (p=.95) reaches 23.77, 27.52, and 25.13 dB across the three fusions with rFID below 0.61. Table 1 reports PSNR, SSIM, and rFID on 50k images with light-blue cells marking each row's training support; supplementary diagnostics cover leave-one-out sensitivity, Shapley ordering, and marginal gains from subset size.

Decoder replacement alone improves generation with the generator and sampled latents fixed, and joint regularization of both stages on DiT-Base exceeds the additive gains of the single-axis changes. Swapping the plain RAEv2 decoder for a FuseReg decoder lowers unguided gFID from 3.01 to 2.21 on the native fusion and from 27.73 to 1.92 under the shifted fusion; on DiT-Base, decoder-only regularization moves 17.28 to 12.37 and generator-only to 13.52, while joint regularization reaches 12.09. The decoder-swap study keeps the DiT, fusion space, guidance setting, and sampled latents fixed; the DiT grids enforce the same controls within each of the DiT-Base and DiT-XL scales over 40 training epochs and are reported as point estimates.

Perspective

The result targets representation autoencoder pipelines that pair a frozen visual encoder with a learned decoder, with experiments on ImageNet-256, 23 candidate layers of DINOv3-L, and DiT-Base and DiT-XL generators under matched training budgets. It lets one decoder serve full, sparse, and single-layer fusions without retraining per configuration, and lets generation benefit from a better readout while the generator and sampled latents stay fixed; for researchers and practitioners who want to improve the decoder–generator interface without touching the pretrained encoder, this offers a directly applicable training principle. The two drop rates must be chosen separately, and the authors accordingly recommend selecting them separately in a new setting.

The objective decompositions assume homogeneous linear predictors and squared loss, and the authors state their role is to identify the training signal induced by subset sampling rather than to replace the nonlinear experiments; the reconstruction and generation benefits in nonlinear models are established empirically. The decoder and generator respond differently to regularization, with preferred rates depending on model scale, evaluation metric, and guidance: on DiT-XL, regularizing the generator alone fails to improve gFID while joint regularization unlocks further gains, and IS prefers a different rate pair than gFID. Under guidance, the full and REPA predictors enter a cross-term, so reducing one predictor's sensitivity alone does not determine the sensitivity of the guided combination. Transfer to other encoder hierarchies, resolutions, data domains, and longer training schedules remains open.

Sources