Skip to main content
Back to timeline
arXivSource publication:

Bidirectional Part–Handle Coupling: Letting a Static 3D Scene Say What Moves, How, and Where to Operate It

Synopsis

The work presents Segment–Snap, which uses three independently trained predictors to recover movable parts, their motion parameters, and operable handles from a static RGB point cloud, coupling them through two one-pass information transfers—handle locations guiding hinge selection, and part context supplying and correcting handles; on Articulate3D public validation, handle guidance raises motion-gated AP from 13.74% to 40.98%, appending part-associated handle candidates raises handle AP from 24.63% to 29.65%, and part-based class correction brings it to 30.99%.

AI-generated editorial illustration: Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Interpretation

Handle locations resolve hinge-side ambiguity in part motion: with masks, classes, scores, and axes fixed, replacing the centroid origin with a box-derived hinge line selected by a handle raises motion-gated AP from 13.74% to 40.98% (+27.25 percentage points), while mask AP and axis-gated AP remain exactly unchanged. The USDNet motion decoder does not explicitly use detected handle locations to choose a hinge, and functional scene graphs express which element controls which object only at a semantic level without a metric axis or origin; here independently detected scene-handle evidence is transferred into the part's metric motion estimate. Fixed-input ablation, with equal mask and axis columns serving as an implementation check; the class breakdown places the gain on rotational parts (1.88% to 56.37%) while translation scores are unchanged because translation origins are not evaluated; the paired scene jackknife 95% interval is positive.

Part-associated handle candidates complement dense detections: appending the joint model's query-associated child masks to the dense detections raises handle AP from 24.63% to 29.65% (+5.01 percentage points), preserving the dense detections' matching prefix and adding 53 ground-truth handle matches. The dense path judges foreground locally, whereas a part query represents a larger movable surface, so their failure modes differ; spatially permuted and random size-matched child masks return exactly 24.63%, indicating the gain comes from real localization rather than merely extending the ranked list. Three controls (dense only, spatially permuted, random size-matched) give the same baseline; across three child-model training draws the append increment is 3.94–5.01 percentage points, and the reference configuration's scene interval is positive under every leave-one-scene-out evaluation.

Part semantics can correct the motion class of appended handles: part context alone adds 0.98 percentage points, and the full rule with the dense-handle fallback reaches 30.99% (+1.34 percentage points), changing only the appended handles' labels while leaving masks, scores, and dense detections untouched. A handle initially inherits its rotation/translation class from its joint-model parent, but another part prediction may offer better evidence; corrupting only the part-class channel lowers AP to 28.62%/28.90%, below the uncorrected 29.65%, showing that useful semantic evidence rather than the number of relabelings explains the benefit. The parts-only and dense-only cue gains overlap and must not be summed; the class-resolved result shows a +2.71-percentage-point translation-handle gain with a positive interval, while the rotation-handle change is unresolved; across three child draws the correction gain ranges from 0.19 to 1.34 percentage points, more variable than the append gain.

Motion decoding needs no trained regressor: the geometric decoder computes axes and origins directly from a planar rectangle fit, a world-vertical axis prior, and handle-guided hinge selection, reaching 40.98% versus at most 39.36% for the displayed learned heads, while continuous regression with the same handle cue reaches only 17.24–18.24%. Splitting 'where does the axis come from' and 'where is the origin' into support estimation, an axis prior, and hinge selection as fixed rules lets handle evidence enter the comparison only through the representative origin, so it can be ablated in isolation. The learned heads consume frozen query features and geometry rather than jointly trained backbones; the rule has higher point estimates on two part-model draws, but the paired intervals include zero, so what is supported is effective decoding without motion-regression training in this setting, not a statistically resolved general advantage.

Perspective

The result targets the prediction setting of recovering interaction structure from a static RGB point cloud; its outputs describe possible interaction and do not include an executed action, a motion trajectory, or a completed simulation asset. It applies to approximately planar mechanisms whose rotational axis is near world-vertical, such as upright doors and drawers. For downstream use, it turns handle location into an interpretable hinge cue, makes part semantics a correction source for handle classes, and provides a geometric decoder that needs no motion-regression training, usable as input to interaction reasoning and simulation pipelines. Evaluation is limited to Articulate3D's 195 training scenes and 42 public validation scenes, test annotations are withheld, and all component studies are development-set comparisons.

Hinge-candidate coverage reaches 94.4% on matched rotational predictions with an 88.3% conditional selection rate, but instance coverage remains the main bottleneck: of 390 ground-truth parts, 139 lack a mask meeting the IoU criterion, 14 more fail the axis gate and 14 fail the origin gate, and the handle union covers about half of 388 ground-truth handles, with the smallest size quartile least covered. The class-correction magnitude ranges from 0.19 to 1.34 percentage points across three child-model draws, so its reference gain and repeated-training gains should be read separately; the paired intervals between learned decoders and the training-free rule include zero, leaving jointly learned representations an open question; and the role of parent conditioning itself is not isolated, with the gain attributed to the implemented proposal source. This reading is full text, but figures appear here as text and tables, so verifying the details of specific qualitative cases would still require the original figures.

Sources