Skip to main content
Back to timeline
arXivSource publication:

Robo-COP co-evolves robot orchestrators and VLA policies during deployment, lifting simulated held-out success from 64.8% to 73.8% and real-world success from 38.3% to 50.0%

Synopsis

Robo-COP co-evolves a robot system's VLM orchestrator and VLA policy during deployment: it curates skill-level demonstrations from its own executions (keeping successful skills even from failed episodes), fine-tunes the policy only when the data can address recurring failures, verifies each candidate on target skills, overall task success, and previously admitted skills before adoption, and revises memory built around the old policy; it raises mean held-out success from 64.8% to 73.8% across ten simulated RoboLab tasks and from 38.3% to 50.0% on three real-world tasks.

Source-provided article image: Co-Evolving Robot Orchestrators and Policies through Deployment
Figure 1 ·

Figure 1 : Two coupled loops for deployment-time improvement. The inner loop adapts the harness through episode reflection. The outer loop uses collected experience to update the policy, then revises memory so the harness can reassess the policy’s capabilities.

arXiv

Interpretation

Reframes deployment-time learning as joint improvement of an agentic orchestrator and its policy rather than one side alone. Prior harnesses assume a frozen policy and can only route around its failures; Robo-COP updates the policy and revises the memory formed around the previous checkpoint. A framework-level contribution, exercised as a full system in simulation and on real hardware.

Turns deployment executions into skill-level demonstrations without manual labels, retaining successful skills from failed episodes. Curation evaluates segments independently of episode outcome: a text judge proposes and labels skills and a video judge confirms success or clear partial progress on sampled frames, so a successful grasp can enter training even if the block is later dropped. Supported by the two-stage text and video judge acceptance rules, with segmentation rules and judge outputs in the appendix.

Verifies each candidate policy before adoption and revises policy-dependent memory after every update. Verification is performed at the skill level rather than only at the episode level, checking target skills, task success, and regressions on admitted skills; after adoption or reversion, advice that discourages using the VLA is marked unverified for the current checkpoint rather than deleted. Verification rejected 5 of 12 candidates in simulation; on the four tasks whose first candidate was reverted, the fixed schedule fell below the frozen-policy baseline on three.

Improves held-out success in simulation and the real world without sacrificing deployment-time success. Mean held-out success rises from 64.8% to 73.8% in simulation, where fixed-schedule training without verification reaches only 65.8%, and from 38.3% to 50.0% on real-world tasks; deployment success is 73.2% versus 68.7% for the frozen-policy baseline and 70.3% for the fixed schedule. Ten simulated tasks with 100 deployment trials and 50 held-out initializations each, and three real-world tasks with 50 learning trials and 20 held-out tests each, under a single-learning-run protocol.

Perspective

The result targets robot systems that learn while attempting their assigned tasks during deployment: the orchestrator uses scripted motion primitives and a VLA policy as tools, and all training data comes from its own executions, so it applies where the policy and scripted tools can at least partially demonstrate the target skills. Simulation covers ten RoboLab tasks spanning three regimes: tasks the pretrained VLA solves alone, tasks scripted primitives solve, and tasks that need both. Real-world evaluation covers three tasks with 40-, 60-, and 90-second limits. For a reader, this offers a reusable deployment-time improvement path: skill-level curation, training-timing control, candidate verification, and memory revision can each be assessed separately, and it also indicates that program-search methods retain an advantage where scripted primitives already cover a task.

Open questions remain: the method can learn only skills its policy and scripted tools already demonstrate at least partially; physical deployment requires human resets, and each policy update adds curation, training, and evaluation cost; where scripted primitives already cover a task, ASPIRE's program search outperforms Robo-COP; and the authors note that scalability still needs validation on larger experiments. Training a single policy on skills pooled from many tasks is proposed as future work toward cross-task generalization and is not yet tested here.

Sources