Public articles linked to the same research event.
arXiv EagleDepth combines depth-adapted latent diffusion with pixel-space generation: it first fine-tunes FLUX.2-klein-4B on paired RGB-depth data to predict a coarse depth latent from resized low-resolution RGB, then freezes that branch and fine-tunes the pretrained pixel diffusion decoder PiD to generate the depth map directly at the target resolution in a single step, conditioned on the coarse latent and the original high-resolution RGB, bypassing the VAE decoder; it reaches competitive overall accuracy on five common depth datasets and Synth4K, leads on high-frequency-masked metrics and boundary F1 across all five Synth4K subsets, and is the fastest among evaluated methods at 4K inference.
EagleDepth combines depth-adapted latent diffusion with pixel-space generation: it first fine-tunes FLUX.2-klein-4B on paired RGB-depth data to predict a coarse depth latent from resized low-resolution RGB, then freezes that branch and fine-tunes the pretrained pixel diffusion decoder PiD to generate the depth map directly at the target resolution in a single step, conditioned on the coarse latent and the original high-resolution RGB, bypassing the VAE decoder; it reaches competitive overall accuracy on five common depth datasets and Synth4K, leads on high-frequency-masked metrics and boundary F1 across all five Synth4K subsets, and is the fastest among evaluated methods at 4K inference.
EagleDepth combines depth-adapted latent diffusion with pixel-space generation: it first fine-tunes FLUX.2-klein-4B on paired RGB-depth data to predict a coarse depth latent from resized low-resolution RGB, then freezes that branch and fine-tunes the pretrained pixel diffusion decoder PiD to generate the depth map directly at the target resolution in a single step, conditioned on the coarse latent and the original high-resolution RGB, bypassing the VAE decoder; it reaches competitive overall accuracy on five common depth datasets and Synth4K, leads on high-frequency-masked metrics and boundary F1 across all five Synth4K subsets, and is the fastest among evaluated methods at 4K inference.
EagleDepth combines depth-adapted latent diffusion with pixel-space generation: it first fine-tunes FLUX.2-klein-4B on paired RGB-depth data to predict a coarse depth latent from resized low-resolution RGB, then freezes that branch and fine-tunes the pretrained pixel diffusion decoder PiD to generate the depth map directly at the target resolution in a single step, conditioned on the coarse latent and the original high-resolution RGB, bypassing the VAE decoder; it reaches competitive overall accuracy on five common depth datasets and Synth4K, leads on high-frequency-masked metrics and boundary F1 across all five Synth4K subsets, and is the fastest among evaluated methods at 4K inference.