SILSA generates 3D from a single image with 384 sliding-window slice latents, lifting PSNR 8.7% and cutting Betti error 9.2%
Synopsis
SILSA is a topology-aware high-resolution 3D generation framework that replaces voxel tokens with a fixed set of overlapping slice latents along the three canonical axes, combining a SliceVAE, a Volumetric Anchor Lattice, and slice-level persistent-homology supervision to raise PSNR from 30.12 to 32.74, coverage from 73.12% to 79.08%, and reduce Betti error by 9.2% on image-to-3D generation while using only 384 tokens, cutting training memory by 40.4% and inference time by 58.5%.
Interpretation
SILSA represents shapes as fixed overlapping slice latents along the x, y, and z axes, where each token summarizes a local depth window, preserving cross-sectional continuity while compressing the generation sequence to 384 tokens. Dense voxels, sparse voxels, octrees, and hierarchical tokenizers have token counts that vary with surface area or occupancy, while triplanes and set-based latents weaken the correspondence between a token and the local geometry it must preserve; SILSA's fixed-length multi-axis slice layout keeps both spatial grounding and a bounded token count. The paper reports SILSA uses only 384 tokens, over 98% fewer than Trellis, XCube, and SparseFlex and 70.0% fewer than the next-most compact baseline; ablations show single-axis variants reach Betti error of 5.92 to 8.07, two-axis variants 2.61 to 2.74, and the three-axis design 1.58.
The SliceVAE encodes oriented surface samples with sliding windows and decodes through sparse volumetric upsampling, adding slice-level topology supervision that matches persistence diagrams within each cross-section and aligns Betti transitions across neighboring slices. Prior topology losses were mostly used for 2D segmentation or implicit reconstruction, and full volumetric persistent-homology supervision is too costly inside high-resolution generative training; SILSA confines supervision to within-slice and between-slice terms, making it tractable. In reconstruction evaluation SILSA lowers CD from 0.61 to 0.59, raises F-Score@0.01 from 96.18 to 96.79 and IoU from 92.54 to 93.01, and reduces Betti-Err from 1.743 to 1.582; ablations show removing the topology loss gives Betti-Err 4.43 and CD 0.78, persistence-only 2.91, and transition-only 2.87.
The Volumetric Anchor Lattice (VAL) gives the rectified-flow transformer a shared 3D workspace where slice tokens from each axis read from and write to their depth planes, turning cross-axis consistency into a spatially grounded communication mechanism. Independently denoised axis-wise slice streams lack coordination, and direct cross-axis attention improves consistency only weakly; VAL's gated read-write lets evidence from different axes accumulate in one shared space. Ablations show no cross-axis communication gives CD 5.87, IoU 49.26, and Betti-Err 17.91; direct cross-attention gives CD 0.60 and Betti-Err 1.64; VAL gives CD 0.59 and Betti-Err 1.58. For the write mechanism, overwrite gives CD 0.66 and Betti-Err 1.91, additive write 0.63 and 1.74, and gated write is best.
On image-to-3D generation SILSA achieves the best FD, PSNR, coverage, and MMD while matching the best KD and LPIPS, and substantially reduces training and inference cost. Compared with strong baselines such as SparseFlex, SILSA improves geometric and distributional metrics while replacing the multi-stage pipeline that first predicts active voxels and then synthesizes local geometry with a single-stage generator. The paper reports PSNR rising from 30.12 to 32.74 (8.7% relative), coverage from 73.12% to 79.08% (5.96 absolute points), FD 10.16 and MMD 14.02; training memory drops from 14.6GB to 8.7GB, training time from 0.38s/iter to 0.21s/iter, and inference from 0.82s/shape to 0.34s/shape.
Perspective
The work targets high-resolution 3D asset generation from a single image, suited to content creation, design, education, and simulation where 3D models must be produced quickly, and its efficiency gains make high-resolution 3D generation more accessible to researchers with limited compute. The method assumes a sufficiently diverse and large training distribution and an informative input image; the paper trains on Trellis-500K, evaluates on 200 Toys4K assets and 50 in-the-wild images, and additionally validates SliceVAE generalization on the open-surface DeepFashion3D dataset. For readers who want to reproduce or extend it, the paper gives adjustable ranges and defaults for slice count, window width, VAL resolution, and topology loss weight.
The paper states that performance depends on the diversity and scale of the training distribution, that objects with structural patterns far outside it may be reconstructed with reduced fidelity, and that ambiguous or low-information input views may yield less faithful 3D structure, with scaling to broader data sources listed as future work. Topology metrics also saturate on topologically simple data such as DeepFashion3D, so their discriminative power is limited there. Increasing slice count beyond the default still slightly improves CD and IoU but worsens Betti error, indicating a trade-off between slice granularity and topological consistency that needs further clarification. This evidence bundle is full text, but figures and tables appear as text, so some visual details cannot be fully recovered from the text alone.
