Skip to main content
Back to timeline
arXivSource publication:

EmbodiedRSI uses a hypothesis graph and value-of-information experiment selection to lift a frozen VLA to 77.0% on RoboCasa365 and 86.8% on LIBERO-Pro, transferring zero-shot to a real robot at 71.3%

Synopsis

EmbodiedRSI is a self-evolving agentic harness that, around a frozen robot foundation model, maintains code and skill hypotheses in a Hypothesis Graph, uses Value-of-Information Experiment Selection to pick the physical experiments that best distinguish them, and turns each interaction into reusable improvement through Code-Skill Co-Evolution and Reward-Grounded Memory Learning; it reaches 77.0% overall success on RoboCasa365 and 71.3% on Composite-Unseen (best baseline 40.1%), 86.8% overall on LIBERO-Pro, and transfers the simulation-evolved harness zero-shot to a real robot at 71.3% overall success.

Source-provided article image: EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
Figure 1 ·

Figure 1: The EmbodiedRSI Self-Evolution Loop. Physical outcomes guide code and skill updates while the robot foundation model remains frozen.

arXiv

Interpretation

A Fast-Slow Dual-System Architecture: the Fast System selects the next physical experiment from current hypothesis weights and cost and rapidly iterates code and skill updates, while the Slow System builds Hierarchical Memory (M1 records physical execution, M2 maintains hypotheses in the Hypothesis Graph, M3 retains reusable cognition distilled from repeated evidence) that informs the next experiment choice. Prior self-evolving robotics harnesses treat physical trials as an undifferentiated resource to consume rather than a scarce resource to actively allocate, lacking an explicit mechanism for deciding what experiment to run next; this architecture makes that allocation explicit. The method section gives a problem formulation (Eqs. 1-3), a posterior mean confidence for each hypothesis (Eq. 4), an exploration-uncertainty measure (Eq. 5), and a value-of-information selection criterion (Eq. 6), with the full procedure and configuration in Appendices D and E.

Code-Skill Co-Evolution: code hypothesis, skill hypothesis, and their joint execution are evaluated from the same initial states, and the measured joint gain (Eq. 7) indicates whether code and skills reinforce each other; the Slow System records the strongest relations as bidirectional works-with edges while directed refines edges retain the refinement history. Code and skills are commonly evolved through separate processes even though long-horizon manipulation often depends on their interaction; this work makes the joint effect a traceable piece of evidence in the update decision. Ablations hold the frozen VLA, task suites, physical-experiment budget, and evaluation protocol fixed while comparing code-only evolution, skill-only evolution, and co-evolution, reporting success and learning efficiency.

Reward-Grounded Memory Learning: experience, retrieved records, and action descriptions are encoded with a frozen Qwen3-Embedding-0.6B encoder, and a trainable policy selects among Retain, Merge, Abstract, Compress, Forget, and Skip, trained with PPO on downstream value to the Fast System (later task improvement, exploration-uncertainty reduction, and physical-experiment efficiency). Memory is typically accumulated for retrieval and reuse rather than learned according to whether it improves later self-evolution; this work turns memory selection into a reinforcement-learning problem with a closed-loop reward. Ablations show Composite-Unseen success rising from 61.7% (retrieval-based) and 65.2% (PPO with task reward alone) to 71.3%, and learning AUC rising from 0.556 and 0.589 to 0.635.

On RoboCasa365 the method reaches 77.0% overall success (Atomic-Seen 94.4, Composite-Seen 63.1, Composite-Unseen 71.3) and on LIBERO-Pro 86.8% overall, leading Harness VLA (72.1%) in all eight perturbation cells; on a real SO-101 arm with a frozen SmolVLA, zero-shot transfer of the harness raises overall success from 46.0% to 71.3%. Relative to end-to-end frozen VLA policies, a world-model baseline, a leaderboard entry, and code-as-policy and harness agents built on a frozen VLA, the method achieves higher success under unseen task compositions and under instruction and position perturbations, and shows the simulation-evolved harness can be deployed directly on hardware. Both simulated benchmarks use 10 random seeds per task and 10 trials per seed, with 10 separate seeds for evolution and another 10 mutually disjoint seeds for evaluation; real-robot results use 30 trials per condition per task with SmolVLA frozen throughout.

Perspective

The results apply to execution-layer adaptation around a frozen robot foundation model using a limited budget of physical trials: the simulated benchmarks are RoboCasa365 (Atomic-Seen, Composite-Seen, Composite-Unseen) and LIBERO-Pro (instruction-redirection T and position-swap S perturbations), and the real-robot setting is an SO-101 arm with SmolVLA across multi-step manipulation, semantic and arithmetic reasoning, and precision grasping. The method suits researchers and engineering teams who want to turn physical interaction into reusable code, skills, and memory without retraining the underlying model; its gains are most visible when previously learned behaviors must be reorganized for new task compositions and changed object layouts.

A careful reader would still watch: how the hypothesis graph and memory scale to longer horizons and more task families; how sensitive value-of-information selection is to the physical-experiment cost weights; what explains the Glasses Bridge Grasp result moving from 60.0% to 56.7%; and the limitation the paper itself notes, that code-as-policy and agentic harness methods learn executable behavior from physical trajectories without updating the robot foundation model's parameters, with future work able to use harness-selected trajectories for online model updates. In addition, this parse is based on the main text and tables; the formulas and algorithm details sit in Appendices D, E, F, and I, which are not included in the readable text, so specific threshold, cost-weight, and training-hyperparameter values cannot be checked here.

Sources