Skip to main content
Back to timeline
arXivSource publication:

A strong agent writes its intervention experience into a playbook for a light agent, lifting real-world manipulation success from 37.3% to 64.0%

Synopsis

The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.

AI-generated editorial illustration: Recursive Harness Distillation across Agents for Robot Manipulation

Interpretation

It introduces RHD, a framework that externalizes knowledge about how to intervene on a deployed VLA into a reusable playbook that a strong agent recursively revises. Prior work uses language models to generate robot programs or treats learned policies as callable tools, but does not explain how effective intervention strategies can be accumulated and transferred across agents; RHD shifts the distillation target from model parameters to external guidance, and revises it based on feedback from the recipient agent's actual execution. The paper provides a problem formulation, an interface definition for three intervention sites (instruction, attention, and action output), and an objective characterization based on the finite-horizon performance-difference identity; revisions are evaluated through complete recipient rollouts and their task outcomes rather than by reproducing the teacher's decisions on its own histories.

On SimplerEnv Bridge, the light agent with the playbook outperforms the strong agent without one, and the same playbook also improves the strong agent. The paper reports 41.7% for GR00T alone, 43.8% for adding the light agent without a playbook, 66.7% for the light agent with the refined playbook, and 79.2% for the strong agent with the same playbook, indicating the gain comes from the playbook rather than from merely adding an agent. Four WidowX tasks (spoon-on-towel, carrot-on-plate, cube stacking, eggplant-in-basket), with 12 initial configurations reserved for development and 12 for evaluation, yielding 48 instances per split; evaluation uses initial configurations excluded from playbook construction, and each evaluated playbook is frozen.

On real robots, the playbook raises the light agent's overall success from 37.3% to 64.0%, while the light agent without a playbook scores 0%. The paper notes that access to intervention tools alone does not ensure effective control of a physical robot: without a playbook the light agent struggled to compensate for recurring spatial offsets between the intended interaction location and the robot's executed motion, and changes to instructions or attention did not reliably resolve these errors; the real-world playbook was adapted by Astra from the GR00T-derived playbook to the physical tasks and fine-tuned policy. A 7-DoF Franka Panda across three tasks (Cube-to-Tray, Cube Stacking, Button Pressing), with 50 human-teleoperated demonstrations per task used for fine-tuning and 25 trials per task at evaluation, 75 trials in total, with object positions randomized before every trial.

Recursive refinement itself is the source of the gain, and the teacher's capability tier matters: when the light agent teaches itself, the benefit disappears. The initial playbook actually reduces the light agent's success from 43.8% without a playbook to 31.3%, while the refined playbook raises it to 66.7%; with Luna as teacher and Luna as recipient the refined playbook achieves only 22.9%, below the 43.8% without a playbook, compared with 66.7% when Astra is the teacher. The two comparisons use the same development split, interface, playbook length limit, and limits on development rollouts and revision attempts; the paper also reports that refinement improves success across all four tasks and reduces mean tokens per episode in most tasks relative to the initial playbook.

Perspective

The result targets robot manipulation settings where a frozen VLA must be deployed repeatedly across changing scenes: it presupposes an intervention interface exposing instruction, internal attention, and action output, plus a strong agent that performs the initial distillation and later revisions. The paper validates on four WidowX tasks in SimplerEnv Bridge and three real tasks on a Franka Panda, where the real tasks use a policy fine-tuned from 50 teleoperated demonstrations per task and are evaluated over 25 trials per task. For teams that want to cut inference cost without retraining the policy, this offers a route of externalizing experience into a playbook and reusing it across agents; the paper also reports that refinement reduces mean tokens per episode in most tasks.

The playbook's benefit depends on having a strong agent as teacher: when the light agent teaches itself, the refined playbook reaches only 22.9%, below the 43.8% without a playbook, so the role of the teacher's capability tier warrants continued observation. The acceptance rule takes the first candidate meeting the target and the paper states it does not guarantee termination, leaving the relationship between revision rounds and final playbook quality an open question. The real-world playbook was adapted by Astra from the GR00T-derived playbook, so the scope of transfer across policies and platforms remains to be tested. In addition, this load contains the full paper text and the external story but not the per-task numeric details in the figures, so task-level differences cannot be checked one by one here.

Sources