Skip to main content
Back to timeline
arXivSource publication:

RoboCoach turns imagined failures into targeted demos: 150 extra subtask demonstrations lift Franka success from 13.3% to 75.0%

Synopsis

The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.

AI-generated editorial illustration: RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Interpretation

It couples what to teach with where to update by making the subtask-expert pair the basic unit, so reusable skills become the shared target of both data acquisition and policy adaptation. Prior active and corrective imitation learning largely decides where additional supervision is needed, while modular policies largely decide which component to update; this work explicitly binds both decisions to the same subtask-expert pair. Four conditions (Single VLA + Uniform, Single VLA + WM-targeted, Modular + Random, RoboCoach) are compared on LIBERO and RoboTwin under matched demonstration budgets, LoRA training settings, and evaluation protocols, with per-round and per-task results reported.

The RIDI loop shifts the world model from simulating more experience to deciding where scarce real experience should go: closed-loop imagination in CoachWorld, progress-based diagnosis of the first timed-out subtask, and scorecard ranking by task-balanced first-timeout mass to request demonstrations. World models have mainly been used for policy evaluation, policy optimization, or generating additional supervision; here the model localizes the bottleneck skill expert and allocates demonstrations. Imagined and deployed success correlate across 22 frozen task-policy checkpoint pairs (Spearman rho = 0.840 as reported); the progress judge reaches a Switch MAE of 0.414 s, Recall@1.0s of 82.73%, and outcome F1 of 88.21% on reference-video replay, and 0.815 s, 62.73%, and 83.84% respectively on CoachWorld-generated observations.

Under a fixed demonstration budget, sending the same targeted demonstrations to the corresponding skill experts outperforms sending them to a shared global adapter. The comparison separates targeted acquisition from update location, isolating the effect of shared versus expert-specific updates on identical acquired data. With a shared adapter, targeted acquisition changes final success relative to uniform acquisition by several percentage points on LIBERO and RoboTwin; applying the same targeted demonstrations to selected experts adds further percentage points; within the modular system, world-model targeting exceeds random target selection by several percentage points (exact values in the main text and appendix tables).

Coached experts can be recombined along unseen routes: four held-out compositions (longer continuation, cross-task combination, skill reordering) average 35.0% success versus 0% for the shared-policy baseline. The held-out routes and their semantic paraphrases are excluded from policy and judge training, coaching data, router examples, threshold calibration, and checkpoint selection, so the evaluation isolates reuse of previously learned skills along unseen routes. Twenty real-robot trials per route: Case A 13/20 (65%), Case B 3/20 (15%), Case C 5/20 (25%), Case D 7/20 (35%), with the baseline at 0/20 in all four cases.

Perspective

The result targets long-horizon manipulation tasks built from reusable atomic skills, in settings that have a shared VLA backbone, a library of skill experts, usable camera calibration, and human operators able to collect requested demonstrations; simulation findings come from ten LIBERO-Long tasks and five RoboTwin 2.0 dual-arm tasks, and real-robot findings from seven tasks on Franka Research 3 and AgileX dual-Piper, with 50 trials per task in simulation and 20 on hardware. It enables follow-up work along two lines: extending the imagine-failures-to-targeted-demonstrations-to-expert-updates loop to more embodiments and longer routes, and treating world models as schedulers of where supervision goes rather than merely as simulators.

Worth watching: when the manipulated object is occluded, or when contact and collision occur outside the camera's view, physical modeling and prediction in that region become difficult, and RoboCoach relies on action-conditioned predictions that preserve task-relevant evidence; agreement across generation seeds reduces sensitivity to stochastic variation but cannot rule out systematic world-model bias; and the expert library is predefined by skill semantics, so learning its structure and reducing the need for human demonstrations remain open directions. This summary is based on the paper full text and appendix tables and does not include additional material from the project page or code repository, so questions about implementation detail and reproduction cost still require the original paper and project page.

Sources