Six-domain evaluation of GPT-6 Astra as an embodied policy: navigation leads, hybrid control lifts manipulation success, but direct in-hand control and dense-reference locomotion remain unreliable
Synopsis
This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
Interpretation
A unified evaluation across six domains that separates Astra-generated robot commands (Direct) from cooperation with learned policies or whole-body controllers (Hybrid), retaining each study's native metrics and protocols. Earlier evaluations of Astra concentrated on manipulation with uneven performance; this work brings navigation, locomotion, and humanoid loco-manipulation into one framework and explicitly separates fixed-configuration evaluations from development trials. Covers six domains, multiple benchmarks and fixed subsets, reporting sample sizes, comparison conditions, and shared initial states; the authors state that fixed-configuration evaluations and development trials are reported separately.
Hybrid control outperforms direct control on several manipulation benchmarks: on RoboDojo, Hybrid succeeds on 24/50 (48%) versus 13/50 (26%) for Direct, with mean Score 62.60 versus 37.81; across ten dexterous tasks the mean scores are 61.6 for Hybrid, 44.2 for the policy, and 16.6 for Direct; on RoboCasa365, Hybrid completes 29/75 (38.7%), Direct 25/75 (33.3%), and the standalone policy 17/75 (22.7%). The results show selective intervention: on RoboDojo, 85.6% of executed control steps follow the policy while only 14.4% are generated or corrected by Astra; in dexterous manipulation, Astra's interventions account for only 11.98% of executed steps. Benchmarks report paired or shared-initial-state comparisons, repeated runs per task, and partial-completion scores; RoboCasa reports paired success/failure counts and an exploratory McNemar test result.
Navigation is Astra's strongest domain: in local comparisons it reaches 46/50 (92%) on RxR instruction following, 39/50 (78%) on R2R, 41/50 (82%) on HM3D object search, and 28/50 (56%) on MP3D, with higher success and SPL than the evaluated released navigation policies. Object-search SPL is only 22.43% (MP3D) and 43.69% (HM3D), showing that success rate conceals inefficient search; the authors therefore report path efficiency separately from success. Four subsets of 50 episodes each, using the same episode lists, scoring rules, and action budgets as released policies; camera settings and action adapters are documented in the appendix.
Direct in-hand control and dense-reference locomotion remain unreliable: the fraction of steps within rotation tolerance is 0.51% for Astra versus 76.90% for RL on the cylinder and 4.40% versus 63.50% on the cuboid; five sequential attempts on a single obstacle course all fail to reach the goal, while PASSAGE reaches it in 13.18 s. The authors attribute the failures to tension between preserving a grasp and reorganizing finger contacts, and to dense references requiring root velocity, limb posture, and their evolution to describe compatible motion rather than merely respecting joint limits. In-hand tasks use paired saved physical initial states, though observations and decision rates differ; the locomotion study is one adaptation trajectory rather than five independent trials, with separate analytic compact-reference diagnostics that do not involve Astra.
Perspective
The work addresses researchers and engineering teams who want to understand what responsibilities a general model can assume in robot control, and it applies to simulated benchmarks and fixed subsets under specified interfaces, controllers, and prompting configurations. For practitioners it offers two actionable paths: let the model lead on navigation and task-level decisions, and let learned policies or whole-body controllers supply low-level motion in contact-rich manipulation and locomotion, with the model handling target correction, contact-condition preparation, and completion checking. The resource counts and latency records also provide a reference for estimating inference budgets.
In in-hand control, Astra and RL differ in observations and decision rates; the locomotion study is one adaptation trajectory rather than independent repeats; and in the compact-reference diagnostics the two generators themselves differ, so the effect of dimensionality is not isolated. The gap between success rate and path efficiency in object search indicates that detours still need finer failure analysis. In addition, some results come from development trials and researcher changes to prompts, interfaces, or controller settings, so while model weights are fixed, system configuration varies, and stability across scenes still needs more repetition and transfer validation.
