Skip to main content
Back to timeline
arXivSource publication:

NavGPT-3 links language-model reasoning to VLA execution through an OS-like runtime, matching human RxR-CE followers at 90.43 SR

Synopsis

NavGPT-3 connects a language-model Planner and the 8B navigation policy NavGPT VLA through a harness with an OS-like runtime: reasoning, execution, and monitoring run as threads with their own context, tools, and permissions, and the runtime assigns motion authority by event priority so the robot can be interrupted and switched between threads; the complete system reaches 81.51 SR on R2R-CE and matches human followers on RxR-CE at 90.43 SR and 78.47 nDTW, lowering the system's minimum reaction time from 3–19 s per language-model decision to 0.5–1 s per action-policy step.

Source-provided article image: NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
Figure 1 ·

Figure 1: Overview of NavGPT-3 : (a) runtime abstraction with the planner, VLA, and spatial tools; (b) synchronous tool use versus asynchronous thread coordination; and (c) benchmark performance. NavGPT-3 reaches human performance on RxR-CE .

arXiv

Interpretation

The harness organizes context, spatial tools, persistent route state, and execution returns into one inspectable, revisable embodied interface: the Planner can call observe_map to read an observed-only top-down map with visited places numbered in visiting order, use navigate_to_node to return to any visited place, use annotate_node to label places, revise locally after a delegated VLA rollout using returned route keyframes and execution status, and explicitly call terminate_episode to end the task. Compared with hierarchical systems that pair a planning model with a learned controller, or systems that retain trajectory evidence only to answer questions, this work uses returned evidence to revise an instruction-following route itself, and a VLA stop ends only one delegation rather than the episode. On a fixed 100-episode R2R-CE subset, basic tools alone give GPT-6 Astra 72 SR at 325.0k tokens and Claude Opus 5 66 SR at 721.0k tokens; adding Map and Backtrack raises these to 83 SR and 73 SR while cutting tokens by 41% and 33%.

The OS-like runtime coordinates reasoning, execution, and monitoring through threads and motion authority: events are handled by priority, protective events rank highest, episode limits take precedence over routine review, and stale events cannot restore an obsolete operation; once a permission is revoked, commands carrying the old version number are ineligible, and resuming motion requires fresh sensing and a new permission. Existing hierarchical systems cover graph memory, backtracking, verification, and recovery, but do not treat the granting, revocation, and version checking of motion authority as a runtime mechanism for managing asynchronous threads. In a matched comparison on a Unitree Go2, overlapping reasoning with priority monitoring raises SR from 73.3 under sequential queued handling to 83.3 and cuts p95 halt latency from 1840 ms to 180 ms; missed stops fall from 7/30 to 0/30, accepted stale commands are 1/30, and resume SR is 80.0.

NavGPT VLA allocates visual tokens by scene change through codec allocation: the first observation of each view is an I-frame with full score, later P-frames are scored by the mean absolute pixel difference of consecutive thumbnails, and the weight combines change, recency, and view while the allocator and token total stay unchanged. Relative to allocating a fixed budget by position alone (recency and view direction), this rule moves tokens from low-change images to images that change, and training randomizes budgets and allocation limits so non-trivial histories enter the regime where the rule can act. At 8B across three cumulative training scopes, codec raises R2R-CE SR by 1.7, 5.7, and 2.1 points and RxR-CE SR by 1.6, 2.6, and 2.6 points; at 4B it raises RxR-CE SR by 3.7, 0.5, and 1.4 points, leaves R2R-CE SR within one point on the first two scopes, and improves both SR and SPL once cross-embodiment data is added.

Data scale contributes more than model scale: growing the mixture from 1.70M VLN records to the complete approximately 19.28M raises R2R-CE SR from 62.6 to 72.54 at 4B and from 65.3 to 74.51 at 8B, and RxR-CE SR from 68.7 to 76.77 and from 70.5 to 78.19, while the 8B model leads the 4B model by 2.0 to 4.3 points at every scope. The mixture is organized along four axes—capability, embodiment and observation interface, conditions and edge cases, and visual-language understanding—unifying 83 sources into one eight-waypoint action space, with LocateAnything grounding contributing 2.31M effective records, 12.0% of the complete mixture. The final step, adding visual grounding and growing the mixture from 11.57M to 19.28M records, still raises R2R-CE SR by 3.3 and 2.9 points at 4B and 8B; the 4B model with cross-embodiment data (69.2 SR) surpasses the 8B model trained on VLN only (65.3), and the complete 4B model (72.54) surpasses the 8B model without the final grounding step (71.6).

Perspective

The interface targets embodied navigation systems that must execute long-horizon instructions in continuous environments: on the simulation side it covers R2R-CE, RxR-CE, and object-goal navigation and embodied question answering on MP3D and HM3D v2, and on the robot side it covers a Unitree Go2 with monocular RGB input, remote or onboard VLA inference, and a LiDAR hazard monitor. The authors state they will release all models, code, and evaluation records, so follow-up work can swap in stronger reasoning models or action policies on the same harness. The runtime design lets the Planner review a fixed copy of the current state while VLA execution continues, interrupting only for revision or termination, a structure suited to replacing periodic checks with event-triggered review.

The ablation studies use a fixed 100-episode R2R-CE subset; the authors note it covers 10 of 11 scenes, has shorter instructions with a narrower length distribution, and state that this establishes coverage and similarity of the listed attributes rather than representativeness or execution-protocol equivalence, with full-split results remaining the basis for benchmark comparisons. Comparisons across Planners describe each complete Planner–scaffold combination rather than ranking the underlying models in isolation. The object-goal and EQA comparisons use 500-episode subsets, and NavGPT-3 has no full-split run there. Each step of the data and model scaling curves changes data volume and source composition together, so the curves measure the cumulative recipe rather than the value of any single source. The robot evaluation uses 30 trials over five routes, three event timings, and two repetitions, and the authors note that detection false negatives and false positives are annotated independently of the detector's output. In addition, several equations, table values, and figure captions are incomplete in the provided text, so specific coefficients, learning rates, and per-item numbers should be checked against the original figures and tables.

Sources