MeshCarve flow-matches in spatially compressed latents at 0.13 tokens per face, reaching the lowest Chamfer distance on Objaverse
Synopsis
MeshCarve introduces a flow matching method that runs entirely in spatially compressed latents: a VAE generates coarse vertex occupancy in one forward pass as an anchor, then a vertex flow and an edge flow generate vertex positions and edge connections separately, with both latents learned by a VertexVAE and an EdgeVAE sharing a hierarchical sparse transformer backbone with spatial-aware compression and vertex-link encoding; its latent holds 0.13 tokens per face on average, achieving the lowest Chamfer distance and highest normal consistency on Objaverse and generalizing to Toys4K.
Figure 2: Overview of MeshCarve. Top: VertexVAE and EdgeVAE share one sparse transformer backbone, reaching the compressed 64 3 64^{3} latent through the spatial down projection of Eq. 1 and decoding the vertex voxels and their connections through the mirrored up projection. Bottom: end-to-end generation from a dense point cloud through the anchor generator and two flows.
arXivInterpretation
MeshCarve encodes vertex positions and edge connections into two separate compact latents, avoiding the significant reconstruction drop that prior work suffers when geometry and topology are jointly encoded. Prior methods such as LATO and MeshFlow place vertex positions and connections in one latent, and the paper reports that this joint encoding substantially degrades reconstruction; MeshCarve uses a VertexVAE and an EdgeVAE that share one backbone but handle the two signals separately. Ablations show that using the same VAE backbone to encode vertex position and edge connection simultaneously degrades performance noticeably, and the degradation is amplified by inaccurate vertex subdivision and pruning because error accumulates fast during inference.
The paper proposes a hierarchical sparse transformer VAE backbone with spatial-aware compression, so the token count follows local vertex density rather than grid resolution, shortening the latent without costing reconstruction. Prior vertex-anchored designs report a large reconstruction loss when tokens are spatially down-sampled, and MeshFlow's TokenMerge trades reconstruction significantly; MeshCarve instead stacks the eight children of a parent voxel channel-wise and learns a linear mapping. Replacing spatial-aware compression with naive mean pooling and broadcast drops EdgeVAE F1 from 0.96965 to 0.88111; VertexVAE reconstruction F1 reaches 0.998 on Toys4K and 0.996 on Objaverse.
The paper proposes vertex-link encoding, which turns arbitrary connectivity into a fixed-length, order-free per-vertex continuous embedding and recovers complex artistic topology faithfully. Existing constructions assign neighbor coordinates to ordered slots and pad to a prescribed maximum degree, which is sensitive to neighbor ordering, caps vertex degree, and reserves dimensions for padding; vertex-link encoding encodes relative positions with random Fourier features and sums over neighbors, with no padding, no degree cap, and no order dependence. Removing vertex-link encoding drops F1 to 0.70793, swapping in slotted vertex neighbor encoding gives 0.96167, and removing RoPE causes the largest drop to 0.40872; given reference vertices, the MeshCarve edge flow leads the LATO.2 topology flow by 22% to 30% of the mean Chamfer and Hausdorff distance on Objaverse.
MeshCarve's latent holds 0.13 tokens per face on average, the shortest among compared methods, affording a larger flow transformer and achieving the lowest Chamfer distance and highest normal consistency on Objaverse while generalizing to Toys4K. Prior autoregressive methods use 1.0 to 9 tokens per face and flow matching methods 0.52 to 6.5 tokens per face; MeshCarve anchors on coarse vertex voxels and compresses further, generating a mesh in about 11 seconds end to end. Evaluated on 185 held-out Objaverse meshes and 500 Toys4K meshes, MeshCarve reaches Chamfer distance .00649 and normal consistency .8529 on Objaverse and is level with LATO.2 in Chamfer distance on Toys4K; its vertex and edge flows have 1.27B and 2.37B parameters.
Perspective
The result targets artisan mesh generation conditioned on a point cloud and controlled by a user-given vertex budget, applies to evaluation distributions such as Objaverse and Toys4K, and is trained on 558K Objaverse meshes with the maximum vertex count capped at 20,000. Its value lies in placing every generative stage in a spatially compressed latent, so the token budget grows only sublinearly as meshes size up and a larger flow transformer runs at acceptable cost; this offers shorter generative sequences for downstream tasks that need editable topology and efficient rendering. The vertex budget enters as a condition, letting generation adapt to user preference, and the anchor generator is a VAE, so anchors can also be sampled with different seeds.
Several open questions remain for a careful reader: VAE reconstruction is close but not exact, and everything downstream inherits that ceiling; above 10K vertices, quantization can reduce mesh quality; and the edge flow is teacher-forced on reference vertices but receives generated ones at inference, creating a domain gap, so it does better given the reference. Surface metrics do not see connectivity, and the end-to-end non-manifold edge share remains higher than that of autoregressive methods, which the paper locates in the vertices the edge stage is run on. Decoding the anchor generator with the posterior mean scores about 2% better in Hausdorff distance and 3% to 4% better in Chamfer distance than a posterior sample, and sampling adds variation to the anchor.
