Skip to main content
Back to timeline
arXivSource publication:

PhysVGGT: Feed-Forward Dense Physical Property Estimation from a Single Image

Synopsis

PhysVGGT formulates physical property estimation as dense per-pixel prediction: a visual geometry transformer extracts geometry-aware tokens, a dense branch predicts per-pixel maps of friction coefficient, Shore hardness, Young's modulus and density, and a global branch predicts object-level mass, all from a single RGB image in one forward pass, achieving state-of-the-art mass estimation on ABO-500 and state-of-the-art friction and hardness on the out-of-distribution NeRF2Physics benchmark with 0.13 s inference latency, 27 times faster than the previous state of the art.

Source-provided article image: PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
Figure 2 ·

Figure 2: The overall framework of PhysVGGT.

arXiv

Interpretation

It proposes PhysVGGT, a feed-forward framework that outputs dense property maps of friction coefficient, Shore hardness, Young's modulus and density together with object-level mass from a single RGB image in one forward pass, eliminating per-object reconstruction and test-time optimization. Prior methods typically rely on per-object NeRF or 3DGS reconstruction followed by material reasoning, or still query test-time VLMs, material retrieval and property lookup within a feed-forward path; this work recasts physical property estimation as dense per-pixel prediction and needs no auxiliary semantic model at inference. On ABO-500 (100 held-out objects) mass ADE is 4.73 kg versus 5.96 kg for the strongest prior result, GaussianProperty; on the NeRF2Physics 13-scene real benchmark friction ADE is 0.102 and Shore hardness ADE is 7.05, best in the table; inference latency is 0.13 s versus 3.6 s for VoMP.

It designs a Cross-Property Coupling (CPC) module that applies self-attention across the property-specific latent representations at the encoder bottleneck, so each property prediction is informed by the others rather than decoded independently. Independent output heads cannot explicitly exchange information, yet friction, hardness, stiffness and density are governed by shared material and geometric characteristics and are strongly correlated; CPC attends across the four properties (not across spatial locations) at the lowest-resolution, most semantic bottleneck, making its cost independent of image resolution and adding only 0.9 M parameters. Ablation shows that relative to the no-CPC variant, the two-layer attention module improves friction MnRE by 3.6%, reduces hardness ADE by 22.4% and improves mass MnRE by 38.2%; alternative couplings such as cross-stitch and FiLM capture only limited interactions.

It adopts a shared decoder with lightweight per-property heads and introduces a global mass head that regresses mass directly from pooled geometry tokens. A single shared upsampling trunk replaces four independent decoders, exploiting the common spatial structure that physical properties share material boundaries; mass is no longer derived from predicted density and reconstructed volume but regressed directly as log-mass from the global geometry-aware representation, avoiding error accumulation from intermediate quantities. The shared decoder cuts decoder parameters from 40.9 M to 31.9 M (a 75.2% reduction in decoder parameters) while improving friction MnRE from 0.808 to 0.831 and reducing hardness ADE from 8.62 to 7.05; the dense ρ×V route gives mass ADE 10.3 kg and MnRE 0.517, whereas the global direct head gives 4.73 kg and 0.655.

It proposes a scalable pseudo-label generation pipeline that produces dense physical-property supervision from multi-view product imagery through part segmentation, material and hardness prediction, and cross-view refinement, without direct physical measurements. Existing annotation pipelines either assign physical attributes to 3D assets, limiting scalability to curated asset collections, or infer properties through per-object reconstruction and optimization, which is expensive and yields no reusable corpus; this pipeline is fully offline, needs no 3D assets, and produces per-pixel supervision over more objects. The training set ABO-8k contains 8,213 objects rendered from 15 calibrated views each (about 123k images), with catalogue mass available for 7,326 objects; cross-view refinement raises friction MnRE from 0.813 to 0.831, hardness from 0.856 to 0.908 and dense mass from 0.450 to 0.655 relative to raw per-view labels.

Perspective

The result targets estimating physical properties of objects and surfaces from a single RGB image, suited to robot grasp and manipulation planning and to physics-based applications that need physical property priors; training and the main quantitative evaluation build on synthetic multi-view renders of the ABO product catalogue with catalogue-derived mass supervision, quantitative friction and hardness validation is performed on real measurements in NeRF2Physics, and terrain generalization is only observed qualitatively on RUGD and RELLIS-3D. The method itself uses visual information only, and the authors note it can be misleading when appearance does not accurately reflect the underlying physical properties; they plan to incorporate additional sensory modalities and to integrate PhysVGGT into robotic systems to evaluate its impact on perception and manipulation.

Young's modulus and density have no measured per-pixel ground truth, so their evaluation is qualitative and the absolute accuracy of these two properties remains an open question; pseudo-labels are derived from part segmentation, VLM material and hardness prediction, cross-view refinement and a material property table, and how their errors propagate to final predictions deserves further observation; mass MnRE is essentially flat as training data grows from 500 to 7225 objects and dips at 2000 objects, while out-of-distribution friction and hardness improve monotonically with data scale, so the mechanism of the scaling effect is not yet clear; the hyperparameter sensitivity study trains for 2 epochs on a 2000-object subset and the authors state it compares relative trends rather than absolute performance, with visible variation across random seeds; in addition, this is a full-text parse and visual details in figures were not checked item by item, so specific figure-level conclusions are best verified against the original figures.

Sources