Skip to main content
Back to timeline
arXivSource publication:

Writing robot tasks as code: HexaAnything lifts RoboCasa365 Composite-Unseen success from 34.3% to 38.3% and turns Harness traces into a stronger HexaModel

Synopsis

The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.

AI-generated editorial illustration: Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

Interpretation

It formulates Physical Coding: Code as World records objects, relations, constraints, observations, and progress predicates, while Code as Policy organizes planning, tool calls, verification, recovery, and execution, together forming an executable interface between state and action. Earlier code-based robot systems covered mainly the policy side, with perception trusted and what was learned stored outside the model; here the world representation and the execution procedure are both written as programs that can be inspected, edited, and versioned, and an independent verifier decides whether they are supported by evidence. The paper gives a formal runtime definition (world program, policy program, Harness, verifier, recovery operator, provenance-aware memory) and illustrates the interface with LoadKebabSandwich and the Hooke's-law experiment, where 'both ingredients are inside the oven' becomes a gating predicate.

With the action model fixed to XR-1, HexaAnything improves long-horizon manipulation: Composite-Unseen from 34.3% to 38.3% (+4.0 points), Composite-Seen from 54.8% to 61.5% (+6.7 points), and overall from 56.6% to 61.1%; on three Composite-Unseen tasks with 100 seeds each the gains are 31.0, 15.0, and 17.0 points. The gain comes from moving decomposition, verification, and recovery into the Harness rather than changing the action model; the general-purpose coding agent Codex driven by the same model reaches 59.5% overall and 34.1% on Composite-Unseen, not improving on native XR-1 there. Controlled comparisons on the same tasks and random seeds, plus additional checks: doubling the VLA action budget on PortionHotDogs did not improve the native condition (21/100 versus 22/100), whereas the coding-agent condition completed 37/100.

Traces returned by the Harness can train a better model: HexaModel v0.1, Qwen3.8-27B fine-tuned on Harness data (9.3K RoboCasa365 VQA examples with chain-of-thought and 1.0K agent traces) plus general data, improves over its base on every split when placed back in the same Harness, from 37.3% to 39.5% on Composite-Unseen and from 60.5% to 61.7% overall. This gives a concrete path from execution evidence to weight updates rather than storing skills in memory or a skill library; Harness-returned data make up 38.8% of the 178.6M training tokens, and part of the traces come from a checkpoint trained in a previous round. Controlled comparison in the same Harness on the same tasks and seeds; the open 27B model reaches 39.5% on Composite-Unseen, 1.2 points above GPT-5.6-Sol (38.3%) in the same Harness.

The same interface extends beyond manipulation: on PhyBench, HexaAnything autonomously designs and carries out three simulated experiments (spring stiffness, gravitational acceleration, coupled-oscillator normal-mode frequencies), with GPT-6-Astra, Opus 5.5, and GPT-5.6-Sol all keeping mean relative error below 5% on every task; on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials, some faster than published references. A scientific experiment succeeds only when its quantitative conclusion is correct, which requires keeping measurement conditions, readings, and fits as checkable code; on the real robot a programming agent writes, validates, and freezes tools between episodes, and humans and agents call them through the same URAI interface. PhyBench uses ten trials per model and task, and the real-robot results rest on three trials per task; the paper states that published references come from other hardware and operating limits and that three trials per task are few.

Perspective

The result speaks to embodied-agent researchers and engineering teams working in controlled simulation and constrained real-robot settings: it shows that when the action model is treated as a callable tool and the Harness owns state bookkeeping, verification, and recovery, long-horizon manipulation and quantitative experiments can improve without retraining the action model. What a reader can reuse directly is the interface design: write task requirements as evaluable predicates, keep the verifier separate from the model that proposes actions, and let traces that pass independent evaluation become training data. The paper also lays out a staged roadmap (Harness bootstrap, model-Harness co-evolution, physical deployment) and notes that on the real robot tools can be written, validated, and frozen by a programming agent between episodes, with humans and agents sharing one tool interface.

The paper describes its evidence as preliminary: the Harness effect is measured with one action model (XR-1) on one benchmark, HexaAnything's overall margin over Codex driven by the same model is 1.6 points, and Codex is ahead on Atomic-Seen; long-horizon evidence comes from RoboCasa365's built-in composite tasks, with longer tasks designed but not yet evaluated; RoboDojo tool evolution rests on few seeds (five for Fold cloth), and in that task the operator chose the pixels for each call; PhyBench runs in simulation with ten trials per model and task, and real-robot results rest on three trials per task with published references obtained on other hardware. Open problems listed include preventing the model and verifier from co-adapting, learning physical causality from sparse trajectories, keeping lifelong memories private, and combining expert corrections, simulator traces, and real-robot failures without letting the agent optimize toward a narrow or self-generated evaluator. In addition, this evidence bundle is a full-text parse in which figures appear as textual descriptions, so specific curves and per-version numbers still need to be checked against the original.

Sources