Skip to main content
Back to timeline
arXivSource publication:

RLE-Bench tests coding agents on 51 robot-learning tasks: GPT-6 Astra and Claude Fable 5.1 lead, while mechanical design defeats every model

Synopsis

The authors introduce RLE-Bench, which organizes 51 robot-learning tasks into four workflows—interactive control, policy learning, perception and estimation, and mechanical design—aggregates task-specific metrics into an RLE Index, and evaluates 11 model–harness combinations, finding that GPT-6 Astra and Claude Fable 5.1 lead by a substantial margin, that policy development shows the smallest gaps, and that no model consistently makes sound mechanical design decisions.

AI-generated editorial illustration: RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

Interpretation

RLE-Bench splits robotics development into four workflows, nine task families, and 51 tasks, where each task is a tuple of instruction, public development environment, resource budget, executable artifact contract, verifier-controlled evaluation environment, and verifier; submitted artifacts range from agent context and harness code plus manual to model checkpoints, pose-estimator code, end-to-end controller code, and MJCF designs with controller code. Existing robotics benchmarks mainly evaluate individual artifacts such as policies or controllers, and CaP-X evaluates coding agents but only for writing code as policies; RLE-Bench brings controller synthesis, robot policy training, perception and estimation, and mechanical design into one evaluation. The paper provides a task-family-to-workflow table and states that tasks come from two sources, repurposed open-source benchmarks and hand-designed tasks; evaluation runs under hidden scenes, dynamics, embodiments, random seeds, or held-out tasks using benchmark-owned instrumentation.

Across 11 model–harness combinations, GPT-6 Astra and Claude Fable 5.1 lead by a substantial margin, with their advantage most pronounced in Interactive Control and Perception and Estimation, both of which require interpreting multimodal observations and acting on repeated environment feedback; in Policy Development the gap between models is substantially smaller, and in Mechanical Design gaps are also small with no model consistently making sound design decisions. The result turns “can coding agents do robotics engineering” from a single leaderboard into workflow-specific capability profiles, and points to visual grounding and reasoning about physical consequences as two distinct axes of differentiation. Conclusions come from the RLE Index (an equally weighted macro-average of the four workflow scores) and per-workflow comparisons for 11 combinations over 51 tasks; the paper also reports native physical diagnostics such as motion tracking error, survival rate, bin clearance, and parts per minute.

Development-time evaluation feedback delivers real, front-loaded improvement: in 29 of 36 sessions the submitted version outperforms the agent's first development-time evaluation; but the improvement is bounded by what the feedback can measure—on rehearsable open-design subtasks development success predicts the official score within 5 points, while on robustness subtasks whose test-time perturbations are never shown to the agent the official score falls 19 points (LIBERO) and 41 points (RoboTwin) below the development score. This turns “does evaluation feedback help” from a qualitative judgment into replayable development-curve evidence and marks the boundary of that benefit: agents' own perturbation proxies do not close the robustness gap. The paper monitors and logs use of the development-time evaluation service, reporting a 29-of-36 session-level comparison and the score gaps on the two subtask types; the authors conclude that building robust policies remains difficult.

External robotics scaffolding generally helps but yields diminishing returns as models get stronger: L1 provides basic robot and sensor APIs, L2 adds SAM3 and Contact-GraspNet, and L3 adds privileged object and fixture positions; richer scaffolding generally improves scores and reduces costs for most models, whereas GPT-6 Astra performs strongly with L1 and gains little or occasionally loses performance with additional support. This gives a capability-conditional conclusion about adding external modules to agents rather than treating it as a universally effective default. The conclusion comes from a controlled comparison of three scaffolding levels within the same task family, covering GPT-5.6, the Opus series, Gemini-3.7-Flash, and GPT-6-Astra among others.

Perspective

The benchmark targets evaluation and comparison of robot-learning engineering capability: researchers, agent developers, and teams that need to judge how far coding agents can go on physical tasks can use it for reproducible comparisons in simulation and use the workflow profiles and cost data to choose models and scaffolding configurations. The authors position it as a qualifying exam for developing agents that can build, test, and improve reliable robot-learning systems.

The authors list scope limits: simulation performance does not establish real-world reliability or safety, and the scale and diversity of simulated scenes are insufficient to assess behavior across all deployment conditions; the four workflows capture only part of robot-learning engineers' work; and the planned public release introduces a risk of future training data contamination. In addition, the mechanical design result that all nine models receive zero worst-arm credit for static stability margin suggests reasoning steps between visible objectives and coupled physical consequences remain uncovered by the evaluation, and the 19-point and 41-point gaps between development and official scores on robustness subtasks leave open the question of how agents could build effective perturbation proxies.

Sources