Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

COSMI composes 222k multi-object interaction sequences from single-object captures and trains a diffusion model that generates one to five objects

COSMI composes multi-object human-object interaction data from single-object captures by extracting contact-consistent clips, mirroring them to balance the hands, and admitting only semantically and physically plausible pairings via a language model and geometric checks, yielding a dataset of 222k sequences and 275 hours with up to five objects (nearly thirty times the largest multi-object capture); it then trains a text-driven diffusion transformer whose object slots share weights and whose objects are predicted relative to the body joints that drive them, outperforming HIMO and multi-object adaptations of MDM and PriorMDM in text alignment and contact accuracy on a benchmark with an unseen object and unseen combinations, with the largest margin on the unseen object.