RoboFoundry treats the agent system itself as the policy to evolve, lifting GPT-5.5 by 27.8% on EmbodiedBench and transferring zero-shot to real robots
Synopsis
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
Interpretation
The paper redefines the executable embodied policy as a system rather than an isolated model, and proposes RoboFoundry so that this system itself becomes the object of evolution across two complementary surfaces: a context system managing active internal context and persistent file-system memory, and a hierarchical skill system organizing atomic skills, reusable compositions, and failure-conditioned recovery. Prior work typically refines a single component of the agent stack (interaction harness, memory, skills, or action interface), whereas RoboFoundry lets execution history decide which part of the system should change and how far that change should propagate, advancing from code-as-policy to system-as-policy. The method section formalizes an inner execution–outer evolution loop and describes the model inspecting the system via file-system operations such as cat and grep and editing it via add, modify, and delete, while the foundation model parameters remain frozen.
The paper develops a capability-guided context–skill co-evolution mechanism: the task-level system policy attributes a failure or inefficiency to one of six capabilities (the first three characterizing decision-making, the latter three memory management), localizes the gap, and edits only the corresponding surface; a successful repair is validated in place, and only when the same improvement is verified against the same capability gap on held-out traces from other tasks is it promoted into a general system rule. This turns self-evolution from unconstrained self-rewriting into a controlled loop of diagnosis, intervention, validation, and promotion, with a capability-conditioned transfer score rather than committing a change on a single task success. The paper gives two-level optimization objectives for task-level and general-level improvement and reports an ablation: RoboFoundry-Lite, which keeps only task edits, scores below the full version on Qwen3.7-Plus, leading the authors to identify general-scope evolution as the critical ingredient.
The paper introduces a shared semantic interface that separates embodiment-invariant decisions from embodiment-specific execution: decisions are expressed once as semantic units and realized by replaceable embodiment bindings through robot-specific primitives, so a system evolved on one platform can be reused on another by swapping bindings, and the same evolved support can attach to different foundation models. This keeps evolved capability from being bound to one robot or one backbone, providing the mechanism for transfer across heterogeneous robots and across foundation models. The paper reports consistent gains across five foundation models (GPT-6 Astra, GPT-5.5, Qwen3.7-Plus, Qwen3.8-27B, GLM5.3-Flash) and, on real robots, zero-shot transfer of the evolved folding skill from a blue towel to held-out green and pink towels with different appearance, geometry, and material properties.
The paper reports effects of system-level evolution across multiple benchmarks and real deployments: on EmbodiedBench, GPT-5.5 improves by 27.8% and Qwen3.7-Plus reaches 70.3% versus GPT-5.5's 72.7%; on RoboMemArena, TSR and CSR beat all baselines by at least 39.0%; on LIBERO-PRO, it outperforms Cap-Agent0 by 243.8%–679.7%; and on real robots it completes nesting-doll manipulation, zero-shot towel-folding transfer, semantic navigation on a Unitree G1, and a six-subtask chemistry experiment. These results extend the system-as-policy claim from a single benchmark to four settings—decision-making, long-horizon memory, distribution-shift robustness, and real robots—and show open-source backbones brought close to frontier-model performance. EmbodiedBench covers 1,128 tasks in four environments; RoboMemArena contains 26 tasks in four memory categories; LIBERO-PRO follows the CaP-Agent0 protocol on 30 tasks from the Object, Goal, and Spatial suites with 50 trials per task and perturbation; real-robot results report success over 10 trials per task.
Perspective
The framework targets embodied agents that use a foundation model as the decision core, and applies to simulation and real-robot settings that provide a writable file-system workspace, recorded execution traces, and the ability to reset task state between trials while persisting validated revisions; its gains come from persistent changes to the context system and hierarchical skill system rather than from updating foundation-model parameters. For a reader, this means it can be understood as a system-level route to continued improvement after deployment: on simulation benchmarks it improves decision-making, long-horizon memory, and robustness to perturbations at once, and on real robots it supports zero-shot transfer and online evolution, with the same evolved system reusable across robots by swapping embodiment bindings and attachable to different foundation models. The paper also notes that perception remains a bottleneck: on LIBERO-PRO, providing privileged object poses improves success across all six settings, reaching 100% on Object and Goal task perturbations.
Several open questions remain for a careful reader: the promotion of general-level improvements relies on capability-conditioned scoring over held-out traces, and how stable that validation stays under larger task-distribution differences is unclear; the paper reports that a larger context budget does not by itself consistently improve success (for example, GLM5.3-Flash drops from 62.0% at 4K to 61.3% at 16K), so the relation between context size and evolution gains still needs clarification; real-robot evaluation uses 10 trials per task, leaving longer-term online evolution over more tasks and longer cycles to be observed; and because this evidence bundle is a full-text parse in which some figures and tables appear as text, a few numbers differ in phrasing between body text and tables, so specific figures should be checked against the original tables.
