Mine Odyssey benchmarks spatial agents on 180 tasks across 30 real-world Minecraft reconstructions, with GPT-6 Astra at 85.6% success and the strongest open-weight model at 23.9%
Synopsis
The authors introduce Mine Odyssey, a benchmark built from 30 BuildTheEarth Minecraft reconstructions of real-world locations (20 outdoor, 10 indoor, spanning 20 countries and regions on five continents), with 1,059 annotated waypoints and 180 natural-language tasks requiring ordered visits to multiple destinations, paired with a programmable action interface, independent arrival verification, and sandboxed isolation; across eight frontier models, GPT-6 Astra reaches 85.6% success, Claude Opus 5.5 reaches 73.9%, and the strongest open-weight model, DeepSeek-V4.1-Flash, reaches 23.9%, with indoor success generally lower than outdoor.
Figure 1: Overview of Mine Odyssey environments and task construction. (a) 30 Minecraft reconstructions and geographic coverage. (b) Task construction with White House and Ueno Park examples: reference locations are mapped to waypoints and ordered into navigation tasks. S denotes the start; numbered markers and dashed lines indicate visit order.
arXivInterpretation
Introduces Mine Odyssey, turning Minecraft reconstructions of real-world places into automatically verifiable multi-goal ordered navigation tasks: 30 locations, 180 tasks, each with 2–12 stops, specified as natural-language instructions giving the visit order. Existing benchmarks often cover narrower settings such as household environments or urban navigation, with limited range of spatial layouts, scales, and traversal requirements; this work places outdoor districts, rural settlements, mountainous terrain, palaces, stadiums, and historic passenger ships in one testbed spanning 20 countries and regions on five continents. Maps come from the BuildTheEarth community project; waypoints are built in three manual stages (geographic reference, matching in the reconstruction, entering the world to create the waypoint at an observed reachable standing position), with a Baritone preliminary reachability check followed by manual traversal; arrival is verified by an independent evaluator sampling player position at 1 Hz against annotated coordinates, crediting only the next required waypoint.
Provides a programmable agent framework and anti-hacking sandbox: AgentBridge exposes native APIs for movement, camera rotation, and item interaction, while mcapi and xdo are driven through Bash so agents can compose commands, scripts, loops, and conditional logic, and use observe_after_sec to inspect and interrupt execution. Unlike evaluations with fixed action primitives, navigation execution is opened to agent-written programs with a persistent workspace for retaining information across decisions; coordinate-disclosure paths are blocked, teleportation and Baritone automatic navigation are disabled, and a Bubblewrap sandbox excludes world saves, waypoint files, and evaluator records. Defaults are first-person view, frozen noon and clear weather, Adventure mode, Peaceful difficulty, with minimap panel and navigation hints disabled; each task is limited to 500 steps and six hours with at most three claim_done calls; the sandbox supports parallel evaluation and GPU-accelerated rendering in containers.
Success rates differ substantially across eight frontier models: GPT-6 Astra 85.6%, Claude Opus 5.5 73.9%, Claude Opus 5 (max) 50.6%, GPT-5.6 Sol 43.9%, Gemini 3.8 Flash 26.7%, DeepSeek-V4.1-Flash 23.9%, Grok 4.6 17.2%, and GLM-5.3-Flash 15.6%; five of the eight complete fewer than half the tasks. The results quantify the current level of agentic spatial intelligence as ordered-navigation success across environments and long horizons, and show an indoor–outdoor gap: Astra's lead over Opus 5.5 widens from 8.8 percentage points outdoors to 18.2 points indoors. Each model is evaluated on all 180 tasks, reporting SR, Checkpoint Coverage, SPL, average turns, and path length; the authors note that indoor–outdoor comparisons concern task subsets with different layouts and route composition and do not isolate the effect of being indoors.
Analysis exposes failure modes and efficiency structure: in more than half of failed attempts agents do not reach even the first required destination; Opus 5.5 reaches 86.1% CC and 50.35% SPL, only 2.9 and 3.35 points below Astra, despite an 11.7-point SR gap, showing high intermediate-stop coverage need not translate into full task completion. Separates partial progress from full completion and adds case evidence on prolonged local stays, route recovery, and incomplete local exploration, complementing what success-rate rankings alone cannot show. Ablations show HUD raises SR by 6.1 points for DeepSeek and 4.0 for GLM, route guidance slightly raises CC but lowers DeepSeek's SR by 2.8 points, midnight lowers SR by 1.7 points for both, and rain lowers it by 3.9 for DeepSeek and 2.2 for GLM; in local-stay statistics Gemini spends 51.7% of its turns in place.
Perspective
The benchmark targets research on agents that must visit multiple destinations in order within unfamiliar 3D environments, and is suited to evaluating map use, route planning, local traversal interactions such as doors and stairs, and plan revision during execution. It serves controlled, automatically verifiable, parallel evaluation in Minecraft reconstructions, and the authors state the task suite and evaluation framework will be publicly released. For readers comparing long-horizon navigation across diverse real-world layouts, it is a directly usable testbed; for those concerned with physical-world deployment, its value lies in decomposing spatial intelligence into observable success, intermediate-stop coverage, and path-efficiency metrics.
The specific numbers in the ablation table are not fully presented in the body text, so readers needing condition-by-condition comparison should consult the original table. The indoor–outdoor success difference comes from task subsets with different layouts and route composition, and the authors explicitly state this does not isolate the effect of being indoors. SPL reference paths are computed by A* on a 0.5-block horizontal lattice with approximate collision geometry, excluding gap jumps, swimming, dynamic stance changes, and button-operated transport, while agent path lengths are estimated from position samples at approximately 1 Hz that may miss movement between samples, so path-efficiency metrics are approximate. Cost and output-token comparisons use reported usage rather than direct measurement of inference compute and average over all tasks without holding successful-task membership fixed.
