Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Dream4ACT unifies multi-embodiment video-action modeling with fixed-view action views, reaching 88.98% average success on RoboTwin 2.0

Dream4ACT renders target joint configurations as "action views" under four fixed virtual cameras, jointly models physical-camera observations and action views with a shared video autoencoder and diffusion transformer under masked flow matching, and recovers executable joint targets by training-free, URDF-constrained multiview matching; it reaches 90.50% (clean) and 87.46% (randomized) success on RoboTwin 2.0, averaging 88.98%, a TriWorldBench score of 65.66, and closed-loop evaluation across five simulated embodiments and four real-world platforms.