Skip to main content
Back to timeline
arXivSource publication:

VolCo uses volumetric contact and volume-wise latent diffusion to generate tighter human grasps with shallower penetration

Synopsis

The authors introduce VolCo, a hierarchical contact representation that expands object-surface sample points into local 3D volumetric grids whose points encode contact likelihood and Continuous Surface Embedding correspondences, and build VolCoDiff on it: a VolumeVAE models feasible local hand configurations while a prior-guided latent diffusion model captures global hand geometry, yielding the lowest penetration depth and highest stability on GRAB and HO3D with tighter, larger-contact-area grasps.

Source-provided article image: VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation
Figure 1 ·

Figure 1 : 3D contact volumes vs. 2D contact maps. Previous contact maps query points only on object surfaces, which can retain only a small fraction of hand surface. Thus it requires nearest-neighbor (NN) queries to avoid penetrations, which can easily stick in local minima. Contact volumes, however, sample points around the object surfaces and is able to retain the whole local hand configuration in the volume. Consequently, contact volumes can deterministically and precisely reconstruct local hand part and guide the pose fitting without instable penetration losses.

arXiv

Interpretation

VolCo extends contact from the object surface to local volumetric grids around it, where each grid point encodes both contact likelihood and a 4-D CSE correspondence vector, allowing deterministic recovery of hand vertex positions via Eq. (3). Earlier contact representations (per-point likelihood, part/anchor labels, surface contact maps) stay confined to the object surface and lack explicit modeling of penetration depth, leaving detail recovery to hand-crafted rules; VolCo scatters contact points around the surface so contact covers a larger hand area and preserves exact penetration depth. On the GRAB test set for reconstruction from contact, switching from surface contact maps to the volume representation reduces EPE by 40.1% at the same point count, and the Normal configuration with CSE correspondence reaches EPE 6.33, F@5 0.752, F@15 0.971 and AUC 0.875, outperforming ContactOpt, ContactGen and ManiDext.

VolCoDiff follows VolCo's hierarchy: VolumeVAE compresses each volume into a 128-dimensional latent conditioned on the local object SDF to model feasible local hand configurations, while a Scene Diffuser-based latent diffusion model models combinations of volume latents, i.e. the global hand prior. Prior latent-diffusion grasp methods focus on global hand pose and ignore contact detail, or recover contact first and fit the hand afterwards, relying on non-differentiable nearest-neighbor penetration losses that risk local minima; here local detail and global composition are decoupled so the diffusion model only handles compressed volume latents. Ablation shows simulation displacement drops from 1.72 to 1.28 when point-based contact is replaced by VolCo; adding hand prior guidance lowers penetration depth from 0.42 to 0.23 and raises contact area from 19.5 to 26.5; adding initialization further lowers displacement to 0.78 and raises contact area to 28.5.

Because VolCo jointly encodes SDF and contact likelihood, contact force can be inferred explicitly through a linear spring-damper model, enabling a force-aware stability loss with slack variables that loosen the constraint on penetration depth. The stability loss improves on the existing force-aware formulation by adding slack variables, so unavoidable random errors in penetration depth in motion-capture data are not penalized as incorrect. In the ablation the stability loss clearly reduces displacement only when prior guidance is present (1.63 to 1.10), which the authors read as geometric plausibility being a prerequisite for physical stability.

On both GRAB and HO3D the method improves stability, penetration depth and contact area over compared methods, and its computation grows more gently with batch size. On in-domain GRAB it exceeds the previous state of the art by 14.8%, 32.5% and 1.6% in simulation displacement, penetration depth and contact area respectively; on out-of-domain HO3D by 12.3%, 28.5% and 18.2%, indicating generalization to unseen objects. Results are reported with 1-sigma error bars; the method has more parameters (74.04M vs. 24.41M) and higher peak memory (358.1M vs. 191.4M) than the point-contact baseline but substantially lower TFLOPs (2.01 vs. 13.2), and inference time barely changes from 5.28 to 5.47 s as samples go from 1 to 10, whereas the baseline rises from 2.90 to 19.57 s.

Perspective

The work targets human grasp synthesis for a given object template, in a setting that uses MANO as the hand representation and GRAB-style motion-capture data for training, with the authors reporting that the volumes roughly cover objects in the YCB benchmark. Its value lies in retaining fine-grained contact detail from training data into generated results, thereby reducing severe penetration and improving geometric plausibility, which can help improve AR/VR experiences and dexterous robotic deployment. The authors also note the method currently prefers only tighter grasps and that the fixed number and scale of volumes limit coverage of object scales; future work will explore controlling VolumeVAE latents and applying adaptive volume numbers and scales.

The numeric entries of the main tables are not fully rendered in the provided text, so the specific SD, PD, IV and CA values for each method on GRAB and HO3D can only be understood through the percentage improvements stated in the prose and cannot be checked item by item. Key hyperparameters such as the slack-variable limit in the stability loss, the default volume half-size and grid resolution, and the VolumeVAE latent dimension appear as symbols or placeholders in the main text and would need the appendix and code to confirm. In addition, the contact area metric depends on a distance threshold and an opposing-normal criterion, and how those are set affects CA comparisons.

Sources