PixelDense splits SAM2, Depth Anything v2 and Metric3D v2 into semantic and geometric alignment streams, lifting PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093
Synopsis
The work introduces PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream and Depth Anything v2 and Metric3D v2 through a geometric projection stream, adding a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces; with a single recipe on PixelGen and DeCo it improves GenEval, DPG-Bench and HPS v2.1, raising PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093 and beating every single-teacher and unfactored multi-teacher variant.
Interpretation
The paper uses dense-prediction foundation models as REPA alignment targets and finds that SAM2, Depth Anything v2 and Metric3D v2 each outperform the DINOv2-only GenEval baseline in pixel-space diffusion, with the two geometric teachers leading the segmentation teacher. Prior REPA-style alignment relied almost exclusively on semantic encoders such as DINOv2 and CLIP, while iREPA pointed to spatial structure as the carrier of the alignment effect; this work follows that lead and systematically uses dense-prediction models, which are explicitly trained to expose spatial structure, as alignment teachers. Fine-tuning on the PixelGen text-to-image backbone, single-teacher results in Table 4 are SAM2 0.8020, DA2 0.8069 and M3D 0.8060, all above the DINOv2-only baseline of 0.7927.
The paper identifies a negative-transfer effect in dense multi-teacher alignment: an equal-weight sum of all four teachers lands below the best single geometric teacher because semantic and geometric gradients compete for one denoiser projection. This reframes teacher composition from adding more teachers to a representation-routing problem, noting that semantic encoders reward viewpoint invariance and instance identity while geometric encoders reward metric layout and surface orientation, so their invariances are incompatible. In Table 4 the unfactored sums score 0.8036 (equal weight), 0.7970 (L2-normalized) and 0.8009 (reduced weights), all below the best single teacher DA2 at 0.8069; the training-dynamics diagnostic in Appendix C.2 shows the naive sum has higher cross-teacher step-delta cosine and a larger semantic-versus-geometric reduction gap.
PixelDense implements the factorization with a semantic stream, a geometric stream and a weight-space orthogonality penalty, reaching 0.8093 and outperforming all single-teacher and unfactored variants, while either stream alone scores lower. Relative to the unfactored sum, factoring lets each teacher's signal be absorbed more evenly: mean relative loss reduction rises from 72.13% to 73.40%, the semantic-versus-geometric reduction gap falls from 0.0719 to 0.0006, and the across-teacher coefficient of variation falls from 0.0973 to 0.0629. In Table 4 the semantic stream (DINOv2+SAM2) scores 0.8022, the geometric stream (DA2+M3D) 0.7960 and full PixelDense 0.8093; leave-one-out shows every teacher contributes (w/o DINOv2 0.7940, w/o SAM2 0.8027, w/o DA2 0.8025, w/o M3D 0.8037).
Structure preservation appears in independent probes and editing: in partial-noise reconstruction panoptic PQ gains up to 53.1% and depth AbsRel drops up to 36.0% at τ=0.5, while SDEdit editing raises background PSNR by up to 2.2 dB. Text-to-image metrics are agnostic to whether a generation preserves the geometry of a specific input scene, so the work adds probes independent of the REPA teachers (OneFormer and Marigold) plus background metrics outside the PIE-Bench edit mask. Table 2 shows PixelDense better than the matched PixelGen baseline across every dataset, reference type, probe and noise level on COCO and Flickr30K, with the largest gap at τ=0.5 narrowing at τ=0.9; Table 6 shows SDEdit better on all six metrics at all seven noise levels.
Perspective
The result targets pixel-space diffusion transformers whose denoiser tokens lie on the same RGB patch grid that dense teachers consume, so each token aligns directly with the teacher feature at the matching spatial location; the paper instantiates it on PixelGen-XXL and DeCo with one recipe, indicating the recipe ports across pixel-space backbones. Alignment is applied at a single mid-network block, and the four teachers (DINOv2-B/14, SAM2 Hiera-L, Depth Anything v2 ViT-L, Metric3D v2 ViT-L) are frozen during training and dropped at inference, so the deployed sampler reads only three RGB channels and sampling cost and downstream filtering pipelines are unchanged. Editing evidence is limited to SDEdit on PIE-Bench, with inversion-based and instruction-based editing left to future work.
The text presents many numbers in tables and appendices, but several cells in the homepage evidence bundle are missing values (for example some percentages, dB figures and speedup factors are empty in the captured text), so those specific magnitudes can only be restated from the values readable in Tables 1, 2, 4, 5, 6 and the appendix tables. Alignment currently sits at a single mid-network block, leaving multi-scale and multi-block alignment, scaling of backbone and training corpus, and teacher banks beyond the four perception models as open directions; editing covers only SDEdit, with inversion-based and instruction-based editing and post-edit identity and 3D geometry metrics not yet evaluated. The teachers are perception models trained on natural images, so PixelDense inherits whatever biases are present in their training distributions on top of the biases of the pixel backbone.
