Skip to main content
Back to timeline
arXivSource publication:

NavHarness carries maps, search records and house notes into the next conversation, lifting GOAT-Bench s-SR by 18.6 to 34.8 points

Synopsis

NavHarness is a training-free embodied harness that makes memory processing part of the navigation loop: each task opens a fresh multi-round agentic session that inherits prior experience through maps, task records and house notes, while outcome verification and run-end consolidation decide what later sessions inherit; on GOAT-Bench it improves s-SR over context-only independent sessions by 18.6 points with GPT-6 Astra, 22.6 with Opus 5, 34.8 with GPT-4o and 30.3 with Qwen3.8-27B, and with SLAM-estimated poses Astra reaches 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE.

AI-generated editorial illustration: NavHarness: Towards Lifelong Embodied Navigation

Interpretation

The paper separates how experience is used from whether it is retained: structured recovery handovers outperform length-matched ordinary summaries, which trail by 8.3 points in s-SR. Prior long-horizon navigation work largely focused on retaining scene knowledge or training retrieval policies; here the transfer of evidence across a session boundary is itself the object of study, with controlled comparisons. Evaluated on all 2,669 GOAT-Bench Val-Unseen subtasks with Opus 5 held fixed and matched task budgets, including empty-handover, mismatched-task-handover and matched-length-summary controls.

Across four reasoning backbones, NavHarness consistently improves task success and path efficiency over same-backbone independent sessions, and reaches state-of-the-art GOAT-Bench and IR2R-CE results using SLAM-estimated poses rather than simulator ground-truth poses. Stronger prior results often rely on simulator poses or auxiliary perception modules; these results are reported without simulator poses, a global navmesh, or navigation-specific learned policies. Full GOAT-Bench split of 2,669 subtasks in 360 episodes across 36 scenes with three seeds; all 36 IR2R-CE tours and 1,824 tasks; Astra reaches 83.7 s-SR and 62.3 SPL on GOAT-Bench and 85.9 s-SR and 76.1 SPL on IR2R-CE.

The two parts of cross-task memory contribute separately: hiding earlier task records costs 9.7 s-SR points and increases steps by 25%, clearing spatial memory costs 12.6 points and increases steps by 32%, and removing both costs 14.7 points. These interventions measure maps and markers separately from textual task records, showing the two cannot substitute for each other. Opus 5 held fixed on the full 2,669-subtask evaluation with three seeds and paired 95% confidence intervals using scene-clustered bootstrap resampling.

In cross-house continuous deployment, run-end consolidation distills the journal into house notes, raising pooled s-SR from 72.8% to 80.5% and SPL from 36.5 to 44.3, with case studies showing agents using earlier experience to interpret ambiguous goals, investigate unresolved questions and resume failed searches. The control disables only run-end consolidation while retaining maps and task records, so it measures what organized long-term memory adds beyond persistent spatial and task-level records. A simulated multi-day deployment over all 36 houses with ten tours per house using Opus 5, with paired gains of 7.7 and 7.8 points and confidence intervals resampling by house, plus qualitative analysis of 48 cases.

Perspective

The work targets settings where a robot pursues successive navigation goals within one house, with goals given by object category, language description or image, and without resetting the robot between tasks. It applies to executor models capable of multi-round multimodal reasoning; the four backbones used are Qwen3.8-27B, GPT-4o, Opus 5 and GPT-6 Astra, and the authors note that the smaller Qwen3.5-4B and Qwen3.5-9B struggled with multi-round multimodal navigation. Spatial state consists of ORB-SLAM3 poses, named markers and small place-photo sets, and the map-query and update tools can be adapted to other pose estimators or richer representations. Evaluation runs in the static scanned environments of GOAT-Bench and IR2R-CE, and the continuous deployment is a simulated multi-day sequence of ten tours per house in which the robot is placed at prescribed starts at tour boundaries.

The paper itself flags open questions. On memory reliability, an incorrect completion accepted by the judge can propagate into house notes, correcting one record does not necessarily update its copies or validate other spatial claims, and the four-view checker still accepts 5.2% of false completion claims and rejects 4.1% of true ones. The spatial representation cannot express object motion, temporal change or the state of manipulable objects, its metric frame is never re-anchored to ground truth, and tracker loop-closure corrections are not propagated into previously written occupancy cells. Evaluation is limited to static simulated environments, and physical deployment would also require robust communication, low-latency obstacle avoidance and independently measured trajectories. As experience accumulates, retrieval must remain selective while handling stale or conflicting accounts. In addition, this evidence bundle is a full-text parse, but several table values are empty after parsing, so specific component scores are taken from the prose and appendix text.

Sources