Skip to main content
Back to timeline
arXivSource publication:

EagleDepth pairs latent diffusion guidance with a pixel diffusion decoder for faster, finer monocular depth at 4K

Related research and updates

Synopsis

EagleDepth combines depth-adapted latent diffusion with pixel-space generation: it first fine-tunes FLUX.2-klein-4B on paired RGB-depth data to predict a coarse depth latent from resized low-resolution RGB, then freezes that branch and fine-tunes the pretrained pixel diffusion decoder PiD to generate the depth map directly at the target resolution in a single step, conditioned on the coarse latent and the original high-resolution RGB, bypassing the VAE decoder; it reaches competitive overall accuracy on five common depth datasets and Synth4K, leads on high-frequency-masked metrics and boundary F1 across all five Synth4K subsets, and is the fastest among evaluated methods at 4K inference.

Source-provided article image: EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder
Figure 1 ·

Figure 1: EagleDepth excels at predicting fine-grained depth maps from high-resolution images. Compared with state-of-the-art methods, our model achieves better depth estimation results with faster inference, as shown in the radar chart (quality metrics averaged over the five Synth4K ( Yu et al., 2026a ) subsets; larger radii indicate better performance)

arXiv

Interpretation

It introduces EagleDepth, which separates depth-adapted latent diffusion priors from pixel-space generation: the latent branch supplies scene-level geometric guidance at low resolution, while the pixel branch generates depth directly at the target resolution. Prior generative depth methods rely on latent-space modeling and VAE reconstruction, where latent compression can obscure narrow structures and closely spaced boundaries and VAE decoding adds cost as image size grows; this work retains latent priors without running the latent backbone at output resolution or using its VAE decoder. The paper describes the full framework and sequential training: FLUX.2-klein-4B is adapted with rank-256 LoRA for 30,000 steps on 536K RGB-depth pairs, then frozen while PiD is fully fine-tuned for 50,000 steps, with 20,000 further mixed-resolution steps, trained on 8 NVIDIA GPUs with AdamW and total batch size 16.

It shows that depth-aware features from downsampled RGB can guide depth generation at higher resolutions, and that pixel diffusion decoding improves fine-detail recovery across different latent priors. In ablations, Lotus-2 and two Flux.2-depth variants are compared standalone and with an adapted PiD decoder; PiD reduces overall AbsRel on all four datasets and improves both high-frequency metrics on Synth4K-1/2 for all three priors, with the pixel-supervised Flux.2-depth giving the lowest overall AbsRel and best high-frequency metrics after pixel decoding. This rests on a controlled decoding ablation across three latent priors; the paper also reports that the latent model always runs at roughly 512-pixel input, with its predicted latent resized to the target patch grid before injection into PiD.

On Synth4K evaluated at native 4K, Ours (MR) leads all compared methods on high-frequency-masked AbsRel, high-frequency-masked delta1, and boundary F1 across all five subsets, while overall accuracy stays comparable to Ours (1024). Most compared methods are trained at lower resolution, whereas this work extends output resolution to 4K through mixed-resolution continuation training, in which 25% of batches keep the earlier data settings and 75% are drawn from high-resolution datasets, preserving existing capability. Evidence is the Synth4K five-subset tables and the boundary F1 table; for example, Ours (MR) reaches boundary F1 31.1 on Synth4K-1 versus 22.5 for Ours (1024) and lower values for the other listed methods. On the five real benchmarks, both variants outperform compared methods in AbsRel and delta1 on ETH3D and DIODE.

It adds a continuity-promoting module that interleaves progressive unpatchification with small-kernel convolutions before the output projection to suppress depth discontinuities at patch boundaries. The unpatchification head of diffusion transformers maps each patch token to a pixel block without local spatial mixing, producing patch-interval artifacts at abrupt depth transitions; the module mixes hidden features at multiple spatial scales, with each Conv-GELU-Conv block jointly initialized as an identity mapping to preserve the pretrained readout. The paper contrasts the decoder with and without the module, reporting that the variant without it shows discontinuities at patch-size intervals while the module suppresses them through spatial feature mixing; component-wise efficiency analysis shows modest overhead (e.g., 77 ms and 14.7 GB at 4K).

Perspective

The result targets monocular depth estimation settings that need high-resolution, fine-grained geometry, such as small-obstacle avoidance and precise robotic manipulation, and it fits 3D scene understanding pipelines that take multi-megapixel images as input. The method extracts geometric priors from roughly 512-pixel low-resolution RGB and generates depth at the target resolution in the pixel branch, so users can choose between the 1024-resolution model and the mixed-resolution extension supporting up to 4K according to detail needs and compute; the paper also notes that patch size 16 prioritizes detail while patch size 32 is faster at similar overall accuracy, suiting applications with lower depth-detail requirements. The sequential training recipe, first the latent branch and then pixel-branch fine-tuning with the latent branch frozen, plus mixed-resolution continuation, offers a reproducible way to extend output resolution on top of existing generative priors.

The reported comparisons use a unified resolution setting: real-benchmark images are resized to roughly 1024 pixels while Synth4K is evaluated at native 4K, so the magnitude of 4K-level detail gains on more native high-resolution real data remains to be observed. The high-resolution stream of mixed-resolution continuation draws only on UnrealStereo4K, UrbanSyn, Spring, and Rawmantic3D, so how well it transfers to broader scene types is worth watching. In addition, the latent branch is fixed at roughly 512-pixel input at inference, leaving open whether low-resolution guidance still suffices for detail recovery when target resolution rises further or scenes contain extremely fine structures. The trade-off between patch size and detail quality also suggests that latency-sensitive deployments with lower detail requirements may need to recalibrate the configuration.

Sources