HarnessPAI's evolving code harness lifts LIBERO-PRO by 61.6 points without retraining the underlying model
Synopsis
The work introduces HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface organizing the underlying action primitive: within a rollout it executes open-loop with a fixed program, and across rollouts it evolves closed-loop by using execution feedback to revise the program and distill failures into reusable skills; across desktop robot arms, household robots, a robot vacuum, and a legged walking agent it improves on both pure action models and code-as-policy baselines without retraining the underlying model, with a 61.6-point gain over π0.5 on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks, and the converged program also serves as an expert-data collector whose data lifts π0.
Interpretation
It proposes HarnessPAI, a model- and embodiment-agnostic Harness framework that treats code as the executable and evolvable interface organizing the underlying action primitive. Whereas Physical AI has focused primarily on the action model mapping observations to low-level controls, this framework integrates perception, task understanding and reasoning, and action execution into a unified system. The text presents the framework descriptively alongside experiments across several embodiments; the summary level gives no ablation or statistical detail.
The framework separates two timescales: within a rollout it executes open-loop at the program level with a fixed program guiding and checking execution, and across rollouts it evolves closed-loop, using execution feedback to revise the program and distill failures into reusable skills. This two-level design means that once a program is selected, rollout execution requires no online high-level LLM deliberation. The text states explicitly that 'Once a program is selected, rollout execution requires no online high-level LLM deliberation', a design-level statement.
Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, it improves on both pure action models and code-as-policy baselines without retraining the underlying model. It reports a 61.6-point gain over π0.5 on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks. The text gives these two specific gain figures and states coverage of several embodiments, but the summary provides no sample sizes or variance.
The converged program is also a cheap and reliable expert-data collector, and fine-tuning π0.5 on collected expert data lifts success rate on LIBERO-PRO by 38.8 points. It turns the program produced during execution into a source of training data, forming a loop from execution to data to model improvement. The text gives the specific figure of 38.8 points and does not state data volume or fine-tuning configuration in the summary.
Perspective
The result is aimed at researchers and engineering teams who want to improve the robustness of embodied agents without retraining the underlying model, and it applies to the tested embodiments (desktop robot arms, household robots, a robot vacuum, a legged walking agent) and benchmarks such as LIBERO-PRO and RoboCasa. Its value lies in treating the code harness as a reusable, verifiable intermediate layer: once a program is selected, rollout execution needs no online high-level LLM deliberation, and the converged program can act as an expert-data collector feeding later fine-tuning.
At the summary level there is no information on sample sizes, variance, statistical tests, or ablations, so the stability of the 61.6, 27.2, and 38.8 gains remains to be confirmed in the full text; the scale of execution feedback needed for program evolution, the precise mechanism of failure distillation, and the tuning cost of transferring across embodiments are open questions a careful reader would still watch. In addition, this assessment is based only on the summary text, without figures or appendices, so the account of experimental detail is limited to what the text states.
