COSMI composes 222k multi-object interaction sequences from single-object captures and trains a diffusion model that generates one to five objects
Related research and updatesSynopsis
COSMI composes multi-object human-object interaction data from single-object captures by extracting contact-consistent clips, mirroring them to balance the hands, and admitting only semantically and physically plausible pairings via a language model and geometric checks, yielding a dataset of 222k sequences and 275 hours with up to five objects (nearly thirty times the largest multi-object capture); it then trains a text-driven diffusion transformer whose object slots share weights and whose objects are predicted relative to the body joints that drive them, outperforming HIMO and multi-object adaptations of MDM and PriorMDM in text alignment and contact accuracy on a benchmark with an unseen object and unseen combinations, with the largest margin on the unseen object.
Interpretation
It introduces a data-composition method built on the observation that interactions are local: contact-consistent atomic clips are extracted from single-object captures, and the acting limb together with its object is transferred onto a body performing another interaction, so multi-object interactions are composed from single-object data. Prior full-body multi-object datasets were bounded by combinatorially expensive capture and stayed below ten hours, against 21.8 hours in the single-object InterAct collection; this work changes data growth from linear in recording time to combinatorial in clips. Thirteen interaction rules define interaction types by body parts that must and must not touch the object, pose constraints, and local reference joints; rules are evaluated per frame with contact defined as proximity below 5 cm, and contiguous windows of at least 3 s yield 1,491 atomic clips, 2,663 after mirroring; compositions are kept only after semantic and geometric feasibility checks.
It builds the COSMI dataset: 221,677 sequences and 275 hours of motion from 59 subjects and 36 objects, seven of them articulated, with up to five objects per sequence and left and right hands balanced at 49/51. This is nearly thirty times the largest multi-object capture, and nothing in the composition is specific to a source dataset, so it grows by adding a rule rather than a recording session, with hand-object recordings entering once a body is fitted. Using HUMOTO metrics on composed sequences and their parent recordings: foot sliding is unchanged, human jerk rises by a fifth to nearly a half, and object jerk and penetration change by less than the captured datasets differ among themselves.
It presents the COSMI method: a single text-conditioned diffusion transformer whose object slots share input and output projections with no parameter tied to object count, so one model generates one to five objects including articulated and worn ones, with a part-assignment head predicting the joint that drives each object and objects expressed relative to that joint. Prior multi-object generators trained one model per object count, padded a fixed state, or generated in separate stages, and left implicit which body part moves which object; this work covers a variable object count in a single stage and explicitly grounds objects to driving joints. The model has 30.4M parameters, between MDM (20.7M) and PriorMDM (39.6M), and trains in eight hours on one RTX 5090; ablations show auxiliary losses matter most (without them foot sliding doubles and temporal contact accuracy drops 17 points) and part assignment improves both text alignment and contact accuracy.
On a benchmark holding out an unseen object (the black chair) and five unseen interaction combinations, COSMI and the baselines alike generalize to the unseen combinations, while COSMI is the most accurate in text alignment and contact accuracy, with its largest margin on the unseen object. The split is by composition rather than by subject, testing what a multi-object model is actually expected to generalize to; baselines lose four to eight points of temporal contact accuracy on the unseen object, whereas COSMI stays within one point of its value on the whole test set. The test set holds 35k sequences against 187k for training; every test sequence is generated three times over its full length; retrained on the HIMO dataset, COSMI has the highest top-3 R-precision in both tiers and the best FID with three objects.
Perspective
The work targets animation, embodied AI, and augmented reality settings that need multi-object full-body interaction data, and its composition covers sustained single interactions performed at once or in sequence; the rules are stated on a body unified to gender-neutral SMPL-X at 30 fps in a common frame, so any clip on this body enters through a rule, as the AMASS walking sequences do. On the method side, COSMI addresses single-stage text-conditioned generation of one to five objects, articulated and worn included, and is trained and evaluated on the authors' benchmark, which holds out one unseen object and five unseen combinations. The authors note that composed sequences would drive any SMPL-X-anchored avatar and demonstrate rendering them on clothed scans registered as SMPL-X+D, with held objects re-attached to the hands of the scan.
Composed sequences inherit the capture quality of their parents, and the authors report human jerk rising by a fifth to nearly a half on some sources while object jerk and penetration change by less than the captured datasets differ among themselves; how these magnitudes affect downstream training awaits broader validation. Articulated objects show a smaller hinge sweep (a sweep ratio of 0.665) and the unseen black chair is occasionally placed incorrectly relative to the body, which the authors attribute to the seats seen in training having no armrests. On the method side, the contact-consistent guidance closes gaps between fingers and objects, but finger motion has yet to reach the precision of dedicated grasp synthesis, and text is a coarse interface that fixes objects and actions but not the person's path or the timing of sequential actions. In addition, the material available here is the paper text and appendix prose rather than the figures themselves, so some qualitative comparisons and failure-case details are conveyed only through the written description.
