Skip to main content
Back to timeline
arXivSource publication:

Real2Gym turns human demonstration videos into executable simulation gyms and trains a failed Franka task into success

Synopsis

Real2Gym is an agentic Real2Sim2Real framework that reconstructs human and robot demonstration videos into visually aligned, natively physics-validated Blender and MuJoCo interactive environments, where an agent generates executable code and distills successes and failures into reusable skills, reaching 87.5% task success with roughly 75% fewer policy-execution tokens than GPT-6 Astra across 24 reconstructed DROID and EgoDex environments and turning a zero-shot real-robot failure on narrow-clearance plate placement into success after simulation-based evolution on a Franka arm.

AI-generated editorial illustration: Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

Interpretation

The Real2Sim pipeline builds executable digital twins from human or robot demonstration videos: geometry is initialized with MoGe-3 or Pi3X, SAM2 parses objects and support surfaces, calibrated cameras and metric cues constrain global scale, and iterative self-inspection across five diagnostic dimensions (primary discrepancy, camera alignment, relative object placement, contact/penetration, appearance fidelity) at event keyframes drives localized correction, with native MuJoCo execution validating grasp, support, and release. Compared with manual pipelines such as RialTo that span scene scanning, mesh repair, articulation modeling, and physics parameterization, and with Agentic Real2Sim's episodic twins, this work embeds physical feasibility validation directly in the reconstruction loop and adds task-conditioned scene augmentation. The paper reports evaluation on 24 MuJoCo environments, 12 scenes each from DROID and EgoDex with four easy, four medium, and four hard scenes per dataset; reconstruction metrics are scored by GPT-6 Astra at high reasoning effort using the Appendix E rubric, with viewpoint alignment on DROID improving by 40.83 points over GPT-6 Astra and a simulation success score of 85.76 on EgoDex.

An experience-driven manipulation agent alternates observation, code generation, execution, and feedback at the subtask level, where one program can coordinate approaching, closing the gripper, and testing grasp retention; after an episode, each code round is labeled positive, negative, or unknown, and the extraction process distills applicability conditions, task procedures, effect checks, object-relative motion rules, and recovery guidance. Relative to Direct Mode baselines that decide individual actions, stage-level code generation reduces model interactions; relative to ASPIRE and Agent as Policy, which emphasize skill accumulation or program reuse, the skill representation here is shared between simulation and the real robot while model weights stay fixed. Table 2 shows 79.2% average success before skill accumulation with 67.3% fewer tokens than GPT-6 Astra; with skills, DROID reaches 91.67% and EgoDex 83.33%, with further token reduction, based on per-task statistics over 12+12 tasks.

Real-robot deployment on a 7-DoF Franka Emika Research 3 with a Robotiq 2F-85 gripper and two Intel RealSense RGB-D cameras covers four tasks, and the zero-shot agent matches or exceeds the Direct Mode baseline on all four, including a pick-block-into-plate task where the baseline's horizontal grasp jams because the block width slightly exceeds the gripper span while the agent executes a feasible top-down grasp. This supplies closed-loop evidence from a simulation skill library to physical execution rather than simulation-only success comparisons. The paper reports success rates and completion times across the four tasks and describes how, after zero-shot failure on narrow-clearance plate placement, an uncalibrated third-person human video was reconstructed into a digital twin, the agent failed twice in simulation, acquired a stable skill by the third evolution iteration, and succeeded after deployment back to the physical Franka.

The framework treats simulation as a scalable trial-and-error setting, using task-conditioned augmentation across object geometry, pose, support height, material properties, distractors, background, and lighting to grow the environment pool while preserving physical feasibility, thereby converting human manipulation experience into reusable robot capabilities. This responds to the cost, safety risk, and low sample efficiency of trial-and-error directly on hardware by completing the self-improvement loop in simulation before transfer. The paper supports this with 24 reconstructed environments, per-task success checks, and four real-robot tasks, and notes that failed candidate environments are routed back for scene or action refinement.

Perspective

The work targets robot learning settings where interactive simulation environments must be built quickly from human or robot demonstrations and reusable manipulation skills accumulated within them; it applies when a URDF/MJCF description and multi-view or single-view RGB input are available and the task can be carried out with stage-level code and a parallel-jaw gripper. It lets researchers build environment pools from public datasets (DROID, EgoDex) and self-recorded videos, run trial and error plus skill distillation in simulation, and deploy the skill library to hardware, which is especially relevant for teams seeking to cut real-robot trial costs and reuse human manipulation experience.

Multi-view misalignment between wrist-mounted and static cameras can yield ghosting or layered surfaces in noisy Pi3X point clouds, leaving residual geometric and textural discrepancies, so cross-view registration and appearance refinement remain future priorities; stage-level code generation improves interaction efficiency but lacks the local fine-grained reactivity required for highly constrained manipulation, and adaptive switching between stage-level and action-level control is left to explore; real-world experiments are currently confined to Franka arms with parallel-jaw grippers, so extension to dexterous hands and humanoid platforms is unknown. In addition, this reading covers the full paper and appendices, but some tables appear in the text as numeric sequences, so a few specific real-robot success rates and times could not be fully matched to their tasks; readers needing exact values should consult the original figures and tables.

Sources