Skip to main content
Back to timeline
arXivSource publication:

UltraTex pushes multi-view diffusion to 2048 resolution with background token dropping and block-sparse attention, reaching up to 91.1x training speedup

Synopsis

UltraTex is an end-to-end multi-view diffusion framework that scales image-guided 3D texturing from 512/768 to 2048 resolution through Background Token Dropping, Block-Sparse Attention, and Foreground-Aware VAE Decoding, and it builds G-buffer TexVerse covering over 268,000 3D assets, achieving 20.6-91.1x training speedup and 22.3-74.6x end-to-end inference speedup within the dataset's common foreground-ratio range of 10%-30%.

AI-generated editorial illustration: UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Interpretation

UltraTex raises multi-view diffusion texture generation to 2048 resolution, producing textures that retain more high-frequency detail. Prior multi-view diffusion texturing methods are typically constrained to low resolutions such as 512 or 768 and struggle to preserve high-frequency details from high-resolution reference images; this work completes 2K generation after compressing a unified multi-view sequence that would otherwise exceed 212K tokens. On a held-out TexVerse test set of 100 unseen objects, the paper reports FID, CLIP-FID, CMMD, CLIP-I, and LPIPS across Unshaded, Shaded, and Relighting tracks, with UltraTex best on most metrics, for example FID 125.13 and LPIPS 0.082 on the Unshaded track.

Background Token Dropping uses geometric foreground masks to remove background tokens before they enter the DiT, fundamentally shortening the sequence processed by every DiT block. Existing efficient multi-view methods mainly sparsify or compress cross-view attention inside the attention module, whereas this work drops tokens before the MM-DiT backbone and keeps each retained token's original RoPE index to preserve spatial layout. The paper reports that specific views contain only 7.9% to 24.9% valid foreground pixels and provides absolute per-foreground-ratio timing tables on a single H200 GPU, showing BTD alone delivers roughly 1.0-52.98x training speedup.

Block-Sparse Attention further reduces attention computation over the compressed foreground sequence and complements Background Token Dropping. The work first analyzes attention maps from a low-resolution full-attention model trained with Background Token Dropping, observes sparse patterns in both double-stream and single-stream blocks, and then applies Top-K block selection over the retained foreground sequence. The paper reports that BSA after BTD contributes roughly 1.55-4.07x training speedup that grows with foreground ratio; the final model uses a retention ratio of 20% because smaller ratios cause noticeable texture degradation.

Foreground-Aware VAE Decoding replaces the noisy background with a canonical background latent and fine-tunes the decoder with a foreground-restricted objective, so foreground-only denoising avoids reconstruction artifacts. Foreground-only denoising leaves the background in its initial Gaussian noise state, and feeding that composite latent directly to a standard VAE decoder degrades foreground quality; this work substitutes a solid-color black image encoding as an in-distribution background latent and supervises the decoder only on foreground pixels. The paper reports reconstruction comparisons: FLUX VAE full image PSNR 47.65, SSIM 0.9937, LPIPS 0.0023; noise background PSNR 29.66; without fine-tuning PSNR 37.81; with fine-tuning PSNR 50.26, SSIM 0.9956, LPIPS 0.0038.

Perspective

The result targets object-centric 3D texture generation with geometry conditions (six per-view normal maps) and a reference image, with training and evaluation conducted under the six-view canonical camera configuration of G-buffer TexVerse and rendering resolutions up to 2K. It makes high-resolution, detail-rich texture generation feasible on a single H200 GPU and supports Test-Time Foreground Scaling to increase the valid foreground area in generated views, providing more texture evidence for downstream 3D texturing. The dataset additionally releases a 36-view sphere configuration that can serve novel-view synthesis, sparse and dense view reconstruction, PBR material estimation, and texture baking or reprojection pipelines.

The paper notes the method may struggle with objects containing highly repetitive texture patterns and is inherently bounded by the FLUX model it builds on: VAE latent compression may blur part of the high-frequency details in rendered ground truth, and FLUX's effective operating range is closer to 1K-2K, so scaling to 4K or 8K may require redesigning DiT-VAE foundation models that natively support ultra-high-resolution inputs. In addition, several speedup factors and dataset ratios appear as placeholders in the loaded main text, while the supplementary tables give complete per-foreground-ratio numbers; readers needing exact figures should consult the supplementary material.

Sources