AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time
Synopsis
This work introduces AdaGeoVLN, a streaming vision-language navigation framework that couples VGGT geometry-foundation-model representations at depths 11, 17, and 23 to the first three Qwen3.5-4B decoder layers and retains historical VGGT global-attention KV states under a fixed budget according to instruction relevance, geometric confidence, and transition novelty, achieving 55.7%/51.4% and 54.1%/44.7% SR/SPL on R2R-CE and RxR-CE Val-Unseen with a single RGB stream and no additional navigation-specific external data, with ablations showing that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations and that bounded navigation-aware retention preserves navigation performance while reducing GFM-KV memory.
Fig. 2 : Hierarchical GFM–VLM fusion across representation depth. VGGT depths 11, 17, and 23 are coupled to Qwen3.5-4B decoder layers L 0 L_{0} , L 1 L_{1} , and L 2 L_{2} . Each branch applies RMSNorm, 2 × 2 2\times 2 grouping, projection, token-wise gating, and a learned scale before residual addition at image-token positions.
arXivInterpretation
Geometry selection across representation depth: VGGT representations at depths 11, 17, and 23 are coupled to Qwen3.5-4B decoder layers L0, L1, and L2 instead of repeatedly injecting a terminal feature. Relative to the common practice of exposing the policy only to the geometry encoder's terminal representation, or repeatedly injecting that same terminal feature, different geometric depths enter successive policy stages without assigning predefined semantic roles to individual layers. On R2R-CE Val-Unseen, hierarchical fusion improves SR/SPL by 9.2/9.1 percentage points over the geometry-free policy, 6.8/7.3 points over single-deep fusion, and 13.6/14.5 points over the Deep×3 control that matches the number and positions of fusion operations.
Navigation-conditioned geometric memory: under a fixed capacity, VGGT global-attention KV states are retained by TopK selection based on instruction relevance, geometric confidence, and transition importance. Relative to temporal retention strategies that treat observation age as a proxy for utility, navigation utility (matching instruction subgoals, depth and point-map confidence, viewpoint transition novelty) determines which historical geometric evidence remains available for future geometric inference. At a similar measured KV footprint, Nav. 900K surpasses Hybrid Inc. (8+24) by 0.6/1.1 percentage points in SR/SPL while using 3.33% less mean GFM-KV memory and 9.35% less mean allocated GPU memory; relative to Hybrid Inc. (8+48) it improves SR/SPL by 0.4/1.4 points while reducing KV and allocated GPU memory by 39.73% and 20.22%.
Leave-one-signal-out ablation of the retention score: at the fixed 900K budget, removing instruction relevance, geometric confidence, or transition importance each reduces both SR and SPL. The ablation compares signal contributions at approximately constant KV memory (variation of only 0.20 MB), showing the complete score outperforms every tested two-signal variant rather than assessing weight optimality. Removing confidence causes the largest SR decrease (2.1 percentage points), followed by instruction relevance (1.5 points) and transition importance (1.4 points).
Simulation benchmarks and physical demonstration: evaluation on R2R-CE and RxR-CE Val-Unseen, plus deployment of the resulting policy on a Unitree G1 humanoid. Relative to the representative methods listed in the tables, AdaGeoVLN reports strong SR/SPL under a single RGB stream with no additional navigation-specific external data, and provides qualitative evidence from simulation and indoor physical execution. On R2R-CE, SR is 55.7% and SPL 51.4%, exceeding JanusVLN* by 2.9 and 2.2 percentage points; on RxR-CE, SR is 54.1%, SPL 44.7%, nDTW 61.8, and NE 5.64 m, corresponding to gains of 2.7, 0.4, and 2.7 points in SR, SPL, and nDTW over JanusVLN* with a 0.82 m NE reduction; the physical part is a qualitative demonstration of two indoor trials.
Perspective
The results target streaming vision-language navigation agents executing instructions in continuous environments, under a setting with a single RGB stream, no panoramic observations, no odometry, no depth input, and no additional navigation-specific external training data; evaluation covers the Val-Unseen splits of R2R-CE and RxR-CE, ablations are mainly on R2R-CE Val-Unseen, and the physical demonstration is an indoor trial on a Unitree G1 with a ZED X Mini camera. It provides a starting point for further study of jointly designing geometric representation depth and historical evidence retention, and for transferring the approach to other geometry foundation models and policy backbones.
The retention coefficients are fixed as equal weights without tuning, and the authors note that the signal ablations do not establish optimal weighting, leaving tuning or learning these coefficients as future work; budget sensitivity shows SR/SPL rising with capacity from 600K to 900K, so behavior at larger budgets or under other allocations remains an open question; the physical part is a qualitative demonstration of two indoor trials, and its behavior across different environments and longer trajectories is yet to be observed; cross-method comparisons should also be read alongside the sensing assumptions and external supervision differences listed in the tables.
