Position Forcing recovers token positions from VecSet latents and quantizes them progressively, reaching best or tied-best ULIP and Uni3D scores for single-stage 3D generation
Synopsis
Observing that VecSet latent tokens retain recoverable spatial correspondences, the work proposes Position Forcing: during denoising it recovers token positions from the predicted clean latent, quantizes them at progressively finer resolutions according to the noise level, and feeds them back to the diffusion Transformer through 3D RoPE, providing coarse-to-fine spatial guidance for single-stage 3D generation without a separate position generation stage and achieving best or tied-best results on all four ULIP and Uni3D metrics.
Figure 1 : Our observation and positional conditioning designs. (I) Latents can be decoded into shape geometry and token positions. (II) (a) VecSet uses no explicit positional conditioning; (b) VoxSet and two-stage pipelines use positions established beforehand. (c) Our naive design recovers positions from the current noisy state. (d) Position Forcing recovers positions from the predicted clean state and applies progressive quantization to provide coarse-to-fine spatial guidance.
arXivInterpretation
The authors find that VecSet latent tokens retain a decodable correspondence with their spatial queries, so query positions can be recovered from the latents alone. Previously, single-stage VecSet generative models had to infer token positions implicitly during denoising, while multi-stage methods relied on a separate structure-generation stage for explicit positional guidance; this observation shows the positional information is already present in the latents. The paper illustrates (Fig. 1(I)) that query positions can be recovered from latent tokens alone and designs a position decoder accordingly; Table 1 shows limited position recovery with a frozen VAE, rising to 99.9 at 0% noise and 67.3 at 30% noise after joint finetuning.
It proposes Position Forcing, a self-conditioning framework combining position quantization with position recovery from the predicted clean latent to provide increasingly precise spatial guidance. Compared with predicting positions from the current noisy latent and building full-resolution positional encodings, the framework adjusts quantization granularity to the noise level and recovers positions from the predicted clean latent. Ablations (Fig. 4) show configuration (b) with fixed resolution 128 exhibits fragmented structures around the house and irregular character surfaces, while progressive quantization in (c) reduces these artifacts; configuration (f), recovering positions from the predicted clean latent, improves all four metrics over (e), which predicts from the noisy latent (ULIP-T 0.077 vs 0.058, ULIP-I 0.130 vs 0.095, Uni3D-T 0.257 vs 0.225, Uni3D-I 0.321 vs 0.265).
The method guides generation along a coarse-to-fine spatial trajectory, enabling high-quality single-stage 3D generation without a separate position generation stage. Multi-stage methods first determine where geometric content resides with an independent structure-generation model before generating details, whereas this method continually updates spatial guidance within a single denoising trajectory. In Table 3, Position Forcing achieves best or tied-best results on all four metrics (ULIP-T 0.077, ULIP-I 0.130, Uni3D-T 0.257, Uni3D-I 0.321), outperforming multi-stage methods TRELLIS, TRELLIS 2, Direct3D-S2, and Hi3DGen as well as single-stage methods CraftsMan 1.5, UniLat3D, and Hunyuan3D 2.1; in Table 2 the largest latent configuration reaches the lowest CD of 5.39 and the highest F1 of 95.38.
Jointly finetuning the VAE and the position decoder supports both geometry reconstruction and position recovery and improves downstream generation quality. Training the position decoder on frozen latents alone yields limited position accuracy that degrades quickly with noise, whereas joint training makes the latent representation both geometry- and position-aware. Table 1 shows position recovery dropping from 67.4 at 0% noise to 6.7 at 30% noise with a frozen VAE, versus 99.9 to 67.3 after joint finetuning; in the ablation, configuration (g) using the Stage-I VAE and retraining the DiT scores lower than the full configuration (f) on all four metrics.
Perspective
The result targets single-stage image-conditioned 3D geometry generation with VecSet/VoxSet latent representations: in that setting, positions can be recovered from latents and fed back to the DiT under progressive quantization, yielding coarse-to-fine spatial guidance within a single denoising trajectory. For researchers and implementers using this representation who want to avoid an extra structure-generation stage, it offers a path that plugs into an existing diffusion Transformer; the reported inference configuration uses 12288 latent tokens, 50-step Euler sampling, and a CFG scale of 5.0.
Position recovery accuracy drops at higher noise levels (67.3 at 30% noise in Table 1), and training uses ground-truth query positions while inference uses the predicted clean latent; how this train-inference difference affects more extreme noise ranges or longer sampling trajectories remains an open question. The specific contributions of the progressive quantization schedule, perturbation strength, and reference coordinate system are presented through ablations and figures, so readers reproducing on their own data should check the implementation details in the appendix. In addition, qualitative comparisons rely on figures, while quantitative conclusions come mainly from ULIP and Uni3D similarity metrics.
