Skip to main content
Back to timeline
arXivSource publication:

HiRAE fuses all 24 DINOv3-L layers with depth-grouped residual budgets, cutting ImageNet-256 reconstruction FID from 0.299 to 0.209 and lifting post-fine-tuning GenEval to 87.70

Synopsis

HiRAE introduces a hierarchical representation autoencoding framework that groups all 24 layers of a frozen DINOv3-L encoder by depth into shallow, middle, and deep groups, learns residual corrections to the deepest representation under group-wise norm caps with tighter budgets for shallower groups, and jointly trains the fusion module and decoder while preserving the latent token count and channel dimension, reducing ImageNet-256 reconstruction FID from RAEv2's 0.299 to 0.209, raising PSNR from 22.667 to 26.377 dB and lowering LPIPS from 0.074 to 0.043, lowering guided generation gFID from 1.060 to 1.038, and improving text-to-image alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning, with post-fine-tuning GenEval rising from 84.

AI-generated editorial illustration: HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

Interpretation

HiRAE replaces manual layer selection or staged optimization in encoder-hierarchy fusion with learnable full-hierarchy fusion plus depth-dependent residual budgets: the 24 layers are split into groups 0-7, 8-15, and 16-23, each learning a residual correction to the deepest feature with norm caps of 0.025, 0.075, and 0.150 and residual dropout of 0.50, 0.25, and 0.10, so the total correction is bounded within 0.250 by the triangle inequality. RAEv2 uses fixed aggregation over a selected 7 layers, DRoRAE trains fusion before adapting the decoder, and DecQ relies on extra detail-query tokens; HiRAE learns spatially varying contributions from all 24 layers within the same latent shape and removes the separate fusion-only adaptation phase. The paper provides the full architecture and training configuration (DINOv3 ViT-L/16, ViT-XL decoder, 16 Stage-1 epochs, 80 Stage-2 epochs) and ablations: without residual regularization rFID can reach 0.023, but guided gFID rises to 7.905 and 14.722 at epoch 20, indicating the budget constraint is what balances reconstruction and generation.

On ImageNet-256, HiRAE-24 reduces reconstruction FID from RAEv2's 0.299 to 0.209 (about 30%) while preserving the latent token count and channel dimension, raising PSNR from 22.667 to 26.377 dB and lowering LPIPS from 0.074 to 0.043 on a matched 5,000-image subset, with lower Sobel and Laplacian errors as well. The gain occurs with the same frozen DINOv3-L encoder and the same latent shape as RAEv2, so it comes from the fusion scheme rather than a larger latent space or a stronger encoder. Reconstruction metrics use the ImageNet validation split and a fixed 5,000-image matched cohort, and improvements hold across all four quartiles of original-image texture strength (absolute LPIPS gains rising from 0.022 to 0.039).

The reconstruction gain does not come at the cost of generation: after 80 generator epochs HiRAE-24 reaches guided gFID 1.038 and IS 257.823 versus RAEv2's 1.060 and 255.300, and unguided gFID 2.129, better than RAEv2 K=23's 3.010 but above RAEv2's 1.650. The paper compares under one shared generator-training and evaluation protocol and shows reconstruction fidelity alone is insufficient to choose a fusion method: DRoRAE-style full-depth fusion adapted to RAEv2 achieves lower rFID (0.065 versus 0.209) but worse guided gFID (1.551 versus 1.038). Generation metrics use 50,000 class-conditioned ImageNet-256 samples per configuration, 100 Euler ODE steps, and epoch-80 EMA generators, and report the six-representation aggregate FDr (1.856 versus RAEv2's 2.170).

On text-to-image generation, HiRAE-7 and HiRAE-24 exceed RAEv2 both after pretraining and after supervised fine-tuning: after pretraining HiRAE-24 improves GenEval by 4.52 points and DPG-Bench by 1.51 points, and after fine-tuning GenEval rises from 84.86 to 87.70 (+2.84), DPG-Bench by 1.45 points, and GenAI-Bench by 0.97 points. This extends the evidence for learned full-hierarchy fusion from class-conditioned generation to text-conditioned generation and supports the full-depth configuration without prior layer-subset selection. Evaluation uses frozen tokenizers, a shared CFG=6 sampling protocol, and 553 GenEval, 1,065 DPG-Bench, and 1,600 GenAI-Bench images, with HiRAE-7 also exceeding RAEv2 on every available benchmark.

Perspective

The work targets representation autoencoding and latent diffusion generation built on frozen pretrained vision encoders: it applies where one wants higher reconstruction fidelity and maintained generation quality without changing the latent token count or channel dimension and without manually selecting a layer subset, including class-conditioned ImageNet-256 generation and text-to-image generation. The method is validated on the 24 layers of DINOv3-L, and HiRAE-7 offers a compact variant over an established 7-layer subset for reuse in existing subset configurations. The analysis (class neighborhoods, spatial effective rank, decoding sensitivity, generator prediction error) offers a reusable diagnostic lens for understanding why better reconstruction need not hurt generation.

The number of depth groups, norm caps, and dropout values come from ablations under a specific configuration (three groups give the best guided gFID, two groups reconstruct slightly better), so how they map to other encoder depths or latent shapes remains to be tested. Class-neighborhood fractions in the latent analysis shift with the coordinate standardization used (under shared anchor standardization RAEv2 moves from 74.454% to 66.024% and HiRAE-24 from 79.536% to 81.466%), so those conclusions depend on the chosen coordinate convention. Decoding sensitivity and generator prediction error are measured at fixed checkpoints and specific perturbation directions, making them diagnostic evidence. Text-to-image evaluation uses each benchmark's established scorer and sampling protocol, so behavior under other scorers or prompt sets remains to be seen.

Sources