RoboQuest benchmark: the strongest frontier multimodal agent succeeds in only 23% of ten kitchen manipulation tasks requiring active exploration
Synopsis
The authors introduce RoboQuest, a benchmark for goal-directed embodied exploration with ten mobile manipulation tasks spanning search, manipulation-based inspection, and interactive testing; five frontier multimodal agents run zero-shot through a common visuomotor interface and the best succeeds in only 23.2% of episodes, while a VLA policy fine-tuned on the released 5,000 demonstrations almost never succeeds; isolated skill tests show agents can perform most required actions when hidden information is supplied, and failure analysis attributes only 18–21% of failures to execution, with most arising from stopping exploration too early and deciding before observing required evidence.
Figure 1: A visual depiction of the salient features of RoboQuest that includes tasks demanding information-seeking physical interactions with active perception.
arXivInterpretation
RoboQuest removes task-critical information from initial observations so it cannot be resolved by passive viewing and must be gathered through physical interaction, organizing ten tasks into search, inspect, and test families where each task specifies only an externally verifiable physical goal and leaves the investigation strategy to the agent. Existing manipulation benchmarks (RLBench, LIBERO, CALVIN, ManiSkill3, RoboTwin 2.0, BEHAVIOR-1K, RoboCasa365, RMBench) evaluate under complete or immediately accessible observability, and active-perception work is largely limited to viewpoint adjustment or single-step un-occlusion; RoboQuest expands exploration into directed search, active inspection, and interactive testing, where exploratory interactions can alter the environment irreversibly. Ten tasks run in RoboCasa365 kitchens simulated in MuJoCo with a Franka Panda arm on a mobile base at 20 Hz, observed by two scene cameras and a wrist camera; 50 instances per task span 18 kitchen layouts and 5 kitchen styles, and all agents play the same instances.
Five frontier multimodal agents achieve low zero-shot success: GPT-6 Astra leads at 23.2%, Opus 5.5, GPT-6.1 Sol, and Fable 5.1 follow at 11–14%, and Gemini 3.8 Flash reaches 2.0%; without the easiest task, Puzzle Box, GPT-6 Astra reaches 15.1%. This is a systematic evaluation of multiple frontier models under a common visuomotor interface for goal-directed embodied exploration, reporting success, progress, cost, and termination modes together. 500 episodes per model, 2,500 total; total evaluation cost $21,352, from $997 for GPT-6.1 Sol to $9,610 for Fable 5.1; GPT-6 Astra has the highest success rate on 7 of 10 tasks, with Opus 5.5 leading only on Wobbly Stand.
Isolated skill tests and failure analysis agree that the bottleneck is exploration and decision-making rather than execution: the three models succeed on 72–81% of isolated skills overall, failure analysis attributes 18–21% of failures to execution, and missing evidence (43/43/46%) plus wrong decisions (23/31/25%) account for 66–74%. The work separates execution from exploration and attributes each failure to the earliest unrepaired breakdown using deterministic per-unit rules, localizing where failures originate. Isolated tests use 20 scenes per skill; failure analysis covers 931/1143/1182 units from 384/431/439 failed episodes for GPT-6 Astra/Opus 5.5/GPT-6.1 Sol; in controlled occlusion experiments, hiding objects drops success by 8–10 points and the additional failures fall entirely under missing evidence.
Exploration often stops too early and disturbances are rarely prevented or repaired: in about half of failures a needed object was never placed, 27–29% of all failures left it hidden in a compartment never opened, while hiding places themselves were rarely missed (only 1–2% of failures); side effects account for 7–13%, about two-thirds involving objects falling to the floor, and only 4 of 38 recovery attempts succeeded. These patterns identify specific weaknesses in evidence acquisition, evidence retention, and consequence monitoring for both reasoning agents and standard VLAs, beyond raw manipulation ability. Attribution uses simulator state logged at every 20 Hz tick and the camera views delivered to the model, with ground-truth segmentation masks for visibility; rules and per-unit labels are released with the benchmark; in Puzzle Box, adding a cover cuts Opus 5.5 from 94.4% to 36.8% while GPT-6 Astra is barely affected.
Perspective
The benchmark targets mobile manipulation agents running in RoboCasa365 kitchens simulated in MuJoCo, suited to evaluating settings where task-critical information is absent from initial observations and must be gathered through physical interaction with autonomous decisions about when to commit; the ten tasks cover search, inspection, and testing uncertainty, with 50 instances per task and held-out kitchen styles and task configurations, making it suitable for measuring zero-shot exploration on unseen scenes and configurations. The released 5,000 demonstrations (500 per task, 366 hours at 20 Hz) with two-level per-frame language annotations support training and diagnosis, and the isolated skill tests and per-unit failure attribution rules support locating the relative contributions of exploration and execution.
Readers should still watch: conclusions rest on MuJoCo simulation and the five tested frontier models plus one fine-tuned VLA, so generalization to real robots and other architectures remains to be tested; for side effects, the text states it cannot be determined from recorded data whether damage was an unforeseen consequence or a collision; some tables and figures (e.g., Fig. 3, Fig. 5, Fig. 7) are not fully rendered in the parsed text, so specific numeric details require the original; the fine-tuned policy records 0.0% success on most tasks with most episodes failing by timeout, and its training and inference details are mainly in Appendix C.
