Skip to main content
Back to timeline
arXivSource publication:

Distilling a single-agent carrying policy into a reusable low-level skill lets a high-level policy coordinate humanoids for sofa carrying, sofa pushing, and four-agent transport

Synopsis

The work proposes a hierarchical framework that relabels successful rollouts of a single-agent physics-based human-object interaction teacher as short-horizon object-proxy displacement supervision, distilling a frozen reusable Object-oriented Motion Skill, while downstream training only learns a high-level policy that emits region-wise proxy-motion commands; in Isaac Gym with 4096 parallel environments the skill achieves better tracking than InterPhys and SkillMimic on both teacher-rollout and random proxy-motion targets, and supports two-human sofa carrying, four-human H-box and cross-box carrying, and two- and four-human sofa pushing.

Source-provided article image: From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction
Figure 1 ·

Figure 1: Our method composes object-oriented motion skills and cooperative planning to support shared object manipulation across team sizes, object geometries, and interaction types.

arXiv

Interpretation

It defines an Object-oriented Action Space of short-horizon object-proxy anchor displacements, decoupling high-level coordination from contact-rich full-body execution. Prior methods such as CooHOI adapt single-human interaction policies to cooperative tasks through policy initialization or task-specific fine-tuning, entangling low-level motor skill with high-level coordination in one task policy; here the low-level skill only realizes near-future proxy motion while the high-level policy performs region-wise proxy-motion planning in a compact object-level action space. The method formalizes the object-oriented action space and the low-level policy mapping, and the appendix specifies that the proxy observation encodes vertices, orientation, vertex linear velocities, and angular velocity, with proxy-motion targets transformed into the receiving humanoid's local frame.

Relabeling successful rollouts of a task-specific single-human teacher as object-proxy motion supervision, combined with online behavior cloning and random proxy-motion RL post-training, yields a reusable Object-oriented Motion Skill. The teacher goal is used only to query teacher actions online, while the student input contains only humanoid state, object-proxy observation, and the future proxy-motion condition; reference human motion is used only for reward computation and never as student input, preventing the student from becoming a full-body motion tracker. On teacher-rollout targets the method reports 79.8 mm position error, 38.2 mm/frame velocity error, 31.3 mm/frame² acceleration error, 100.0% Trajectory Targets Reached, and 98.9% complete trajectory-following success, better than InterPhys (106.5/47.8/50.9/82.6/68.2) and SkillMimic (91.0/40.3/44.2/90.8/81.2); on random proxy-motion targets it reports 95.8/45.2/39.2/96.3/91.8, also leading.

Ablations show the two training stages play distinct roles: online distillation supplies a feasible-interaction prior, and random proxy-motion post-training expands the executable motion range. Removing online distillation drops complete trajectory-following success on random targets to 23.8%, indicating object-proxy tracking rewards alone are insufficient to bootstrap contact-rich carrying; removing random-motion post-training keeps 97.2% TTR and 98.1% trajectory success on teacher-rollout targets but drops trajectory success on random targets to 38.7%. Both ablations are compared against the full method under the same target interface and the same reported metrics, covering both the teacher-rollout and random proxy-motion target distributions.

With the skill frozen, training only a high-level policy transfers across object geometries, interaction types, and team sizes. The skill is distilled only from single-human carrying yet supports two- and four-human sofa pushing, showing interaction-mode transfer requires only a new high-level policy; single-human armchair/table carrying and large-box pushing further confirm geometric transfer. Success rate and precision are reported over 4096 parallel environments with a 15-rigid-body, 28-PD-joint humanoid: two-human sofa carrying 88.21% versus CooHOI's 84.17%, four-human H-box 85.63% versus 80.96%, four-human cross-box 86.14% versus 82.38%; for pushing, which CooHOI does not cover, the method reports 96.24% for single-human large-box, 94.98% for two-human sofa, and 88.14% for four-human sofa.

Perspective

The framework targets compositional human-object interaction within a reusable Object-oriented Motion Skill, where a high-level policy coordinates local proxy motions through a shared low-level executor. It applies to settings where the object or a task-relevant local region is represented as a cuboid proxy and manipulation regions can be assigned in advance from object geometry and interaction role, such as furniture carrying and pushing; during skill learning the box size is also randomized around its nominal scale to improve geometric coverage. For researchers and engineering teams who want to reuse an existing single-human interaction policy while extending to different object geometries, interaction types, and team sizes, this interface can serve directly as the action space of a downstream high-level policy.

The authors note the framework does not explicitly address composition across heterogeneous skill types or long-horizon temporal task structures such as multi-stage task decomposition and dynamic role reassignment, and list skill-level composition with higher-level planning as future work. A careful reader may also watch how the procedurally generated random proxy-motion targets relate to the distribution produced by real downstream high-level policies; that pushing tasks lack a CooHOI comparison, so cross-interaction-mode comparison currently comes from the paper's own results; and that some values and symbols in the appendix tables and equations are missing in this parse, so reproducing specific hyperparameters and reward weights requires checking the original appendix.

Sources