Tex-Zero trains on roughly 11.1 million 2D images and generates textures on real 3D assets, beating baselines trained on 3D data
Synopsis
Tex-Zero introduces a data-construction pipeline that treats each 2D image as a plane in 3D space and applies patch-wise random rotations and aggregation, training a native 3D texture generation VAE and DiT without any real 3D assets and achieving high-fidelity texture reconstruction and generation on real 3D assets, outperforming baselines such as NaTex and TRELLIS.2 in six-view and front-view evaluations.
Interpretation
The paper proposes and validates a counterintuitive observation: native 3D texture generation training essentially needs high-quality, fine-grained color information, while geometric information is less critical and can be manually constructed rather than taken from real 3D assets. Prior native texture generation methods such as NaTex and TRELLIS.2 generally rely on large-scale, high-quality 3D assets for training, and acquiring such assets has long been difficult; this work is the first to shift the training data source from 3D assets to 2D images. Appendix A analyzes why: fixed-plane training leaves 3D kernel directions of sparse convolutions unsupervised, globally rotated planes remain coplanar, independently rotating patches breaks coplanarity, and aggregation brings differently oriented patches into the same receptive field; the ablation table shows LPIPS 0.2160 and PSNR-PC 11.34 for one patch without rotation, versus LPIPS 0.0355 and PSNR-PC 28.97 for 16 rotated and aggregated patches.
The Tex-Zero VAE is trained only on constructed image data yet reconstructs real 3D assets with high quality at inference time and also reconstructs 2D images faithfully, forming a shared latent space for 2D and 3D. Existing 3D VAEs perform poorly on 2D image reconstruction, whereas this VAE consistently outperforms its 3D-trained counterparts and other baselines on 2D image reconstruction metrics. In Table 1, Tex-Zero VAE-f8c16 reaches PSNR-PC 0.0154, PSNR 34.03, SSIM 42.41, LPIPS 0.987 on 3D asset reconstruction and PSNR 39.82, SSIM 0.956 on 2D image reconstruction; the same architecture trained on 3D data (VAE-f8c16) gives 0.0208, 36.55, 44.71, 0.989 and 36.71, 0.944 respectively.
The Tex-Zero DiT is likewise trained only on constructed image data, representing six-view conditioning images as planes in 3D space and encoding them with the same VAE, thereby reducing the representation gap and generating fine-grained textures on real 3D assets. Conditioning images were previously encoded with DINO or a separate VAE, creating a representation gap; this work puts conditions and target textures in one shared latent space and randomly samples view subsets during training to improve generalization. In Table 2, Tex-Zero achieves LPIPS 0.0340, PSNR 35.68, SSIM 0.983 in the six-view setting and 0.0294, 35.31, 0.985 in the front-view setting, outperforming NaTex (six-view 0.0754, 27.74, 0.949) and TRELLIS.2 (front-view 0.1187, 21.38, 0.900); Table 5 shows LPIPS 0.0481, PSNR 32.54, SSIM 0.970 when conditioning is encoded by the Tex-Zero VAE, better than DINO, a separate VAE, and a shared 3D-trained VAE.
When real 3D assets are available, image-derived data still helps: joint 2D and 3D training improves both VAE reconstruction and DiT generation. This indicates the data construction is not merely a substitute when 3D data is scarce but can also complement existing 3D data pipelines. In Table 3, VAE-f16c32 trained on 3D data gives LPIPS 0.0343, PSNR 41.84, SSIM 0.981, improving to 0.0140, 42.85, 0.988 with joint 2D+3D training; the DiT improves correspondingly from 0.0302, 36.23, 0.980 to 0.0216, 37.68, 0.983.
Perspective
The result targets native 3D texture generation, i.e., directly predicting colors in 3D space given geometry and multi-view reference images, and applies to research and engineering settings that want to avoid costly 3D asset acquisition by leveraging existing 2D image collections to scale training data. The paper trains on SA-1B, BLIP3o-60k, and ShareGPT-4o, about 11.1 million images in total, resized to the training resolution; the VAE is evaluated on an internal test set of 200 real 3D assets, the DiT on an internal set of 160 samples, and 2D image reconstruction on the ImageNet validation set. For downstream users, this means one can first train a shared 2D-3D VAE and DiT on 2D images, then feed real geometry and multi-view images at inference time; when 3D data is already available, joint 2D+3D training as in Table 3 is also an option.
The quantitative conclusions rest mainly on internal test sets (200 3D assets, 160 generation samples) and the ImageNet validation set, and the relationship between these evaluation sets and the public image data used for training is not elaborated, so behavior on broader public benchmarks is worth watching. In Table 1, Tex-Zero VAE-f8c16 is slightly lower than the same architecture trained on 3D data on 3D asset reconstruction PSNR and SSIM (34.03 vs 36.55, 42.41 vs 44.71) while better on PSNR-PC and LPIPS, so the trade-off across metrics deserves further observation. The ablations are mostly short 20K-step runs, leaving longer-training trends an open question. In addition, the DiT randomly samples view subsets and its conditioning images are orthographically rendered from point clouds without an illumination model, so the effect of these settings under real captured images also merits follow-up.
