Skip to main content
Back to timeline
arXivSource publication:

InterEvolve evolves reward programs at test time for a fixed humanoid controller, lifting simulated success from 34.6% to 86.5% and running autonomously on a Unitree G1

Synopsis

InterEvolve introduces a test-time evolution framework in which an object-aware forward-backward (FB) behavioral foundation model executes reward programs, an LLM agent revises program structure in context while CMA-ES tunes its constants, and every candidate is verified across parallel simulation scenarios, letting a fixed humanoid controller solve untrained loco-manipulation tasks without retraining and raising macro-average success across eight task families from 34.6% for the best calibrated fixed program to 86.5%, with evolved skills deployed on a physical Unitree G1.

AI-generated editorial illustration: InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

Interpretation

The framework shifts task-specific compute from training to test-time evolution: an LLM agent writes and adapts reward programs in context, executed by a fixed, reusable controller for new tasks and scenes. Prior language-model reward-design work trains a new policy for each candidate reward, whereas here each candidate costs a batch of rollouts on a fixed controller, and verified programs enter a skill library for later tasks. The paper states three contributions and reports macro-average success 86.5%, Earned 95.6%, 2.1 GPU-h and 0.23 M tokens over eight task families; ablations show removing CMA-ES tuning drops success to 51.6% and forcing a single stage to 44.7%.

It contributes an object-aware FB behavioral foundation model that attaches trainable object residuals to a frozen body prior, reading object position, orientation, linear and angular velocity, and distance-decayed vectors from body links to the nearest object surface, turning reward objectives into whole-body interaction. Earlier humanoid FB models such as BFM-Zero observe only the body, so rewards differing only in the object's goal collapse into one behavior; here the actor and the forward and backward maps all receive object information, with residual branches trained on human-object interaction data. On the tracking benchmark the object-aware model reaches SR 60% and object-surface error 30.80 cm, versus SR 8% and 62.33 cm for body-only BFM-Zero; removing link geometry drops SR to 12% and removing object state to 20%, training from scratch reaches 2%, and wider residuals raise SR from 28% to 60%.

Reward programs express tasks as staged rewards with completion conditions and tunable constants, with structure revised by the LLM and constants calibrated by CMA-ES in two nested loops. Common task interfaces such as motion references, goal states, and skill labels struggle to express contact-rich, multi-stage interaction; here the reward itself is the editable task strategy, with structure and constants explicitly separated. Ablations show the largest loss from forcing a single stage (SR 44.7%), then from removing numerical tuning (51.6%), with dropping targeted edits (68.9%) and multi-scenario evaluation (68.4%) costing similarly, and scene context mattering least (78.7%).

Evolved experience can be reused and composed: the skill library records each completed task's scene, task text, verifier pass rates, and selected program, and composite tasks are solved when the library covers their contact modes. Prior work tends to keep experience inside policy weights; here it is stored as readable, editable, retrievable text programs, so later tasks evolve on top of earlier solutions rather than rediscovering them. On three composite tasks the full library solves 8/10 relocation, 4/10 stacking, and 5/10 carry-place-kick episodes, versus almost none without a library; running stored programs directly almost never succeeds (6.3% for kicking, 0.0% for the rest), so the gain comes from search adapting stored programs.

Perspective

The result targets humanoid controllers that already hold broad whole-body motion priors, is validated in simulation, and runs on one physical Unitree G1 with egocentric onboard perception for two evolved tasks; the applicable setting is loco-manipulation task families whose outcomes and physical constraints a fixed verifier can judge, including pushing, carrying, tipping, kicking, lifting onto a support, and pushing through gates or around obstacles, plus long-horizon tasks composed from those contact modes. For a reader, this means task knowledge can be pulled out of policy weights and written as readable, editable, retrievable reward programs that the same controller is repeatedly called on to execute at test time; for teams working on robot learning or embodied intelligence, it offers a path that trades simulation rollouts for task adaptation and deposits successful experience into a skill library.

Program search is bounded by the controller's motor repertoire, the reward-inference bank, and the available measurements, and cannot elicit behavior the controller never learned; each search costs GPU-hours of simulation and LLM latency, which rules out real-time replanning during physical execution. Cross-object transfer varies by family, with push to mark reaching at most 25% on other boxes and the small box almost never pushed around an obstacle; kicking remains the hardest family because kicks are rare in the training data. Skill-library reuse depends on covering the contact modes a task needs, and a carry-only library helps little on stacking or kicking. Success grows with test-time scaling but not monotonically, dipping from round 2 to round 3 because full success requires all criteria in the same rollout. Letting accumulated experience also update the controller is listed by the authors as a next step.

Sources