Skip to main content
Back to timeline
arXivSource publication:

PanoVLN lifts R2R-CE success rate to 77.3%, 11.9 points above the previous best, using panoramic vision

Synopsis

PanoVLN performs vision-and-language navigation in continuous environments with a 4B vision-language model and RGB-only panoramic input, combining longer action-sequence supervision, confidence-guided execution, decision-centric training data, and fused panoramic geometric features to reach 77.3% and 78.0% success rate on R2R-CE and RxR-CE Val-Unseen, exceeding the previous best by 11.9 and 8.7 percentage points, and enabling real-world quadruped navigation with fewer pauses.

AI-generated editorial illustration: PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

Interpretation

The paper finds that simply replacing perspective images with equirectangular panoramas (ERPs) under the same training and inference setup yields only limited gains and can even reduce performance, and diagnoses that action prediction, training supervision, and visual representation all need adaptation. Most prior VLN work takes perspective images as input; this work treats whether panoramas are inherently better as a question to test rather than an assumption, and reports the input-only replacement as a controlled observation. The paper reports that with the same VLM, training trajectories, action supervision, and execution procedure, input replacement alone gives limited gains or a drop, motivating the subsequent adaptations.

Longer action-sequence supervision combined with confidence-guided execution (CGE) lets the model predict and execute longer action segments from a single panorama, with CGE using action uncertainty to decide how many predicted actions to execute before reobserving and replanning. Prior policies execute a fixed-length prefix of their prediction at inference; this work makes execution length dynamically determined by prediction uncertainty and separates prediction horizon from execution length as two distinct choices. The horizon study shows ERP policies trail perspective policies at short horizons but overtake them as targets lengthen, with perspective peaking at a shorter horizon while ERP favors longer ones; in the execution study CGE outperforms random prefixes and all fixed strategies, including one-action execution.

The paper constructs a decision-centric dataset of 98K trajectories across 800 HM3D scenes, with routes containing frequent branching points and instructions that clearly specify the chosen path, verified by motion consistency, choice grounding, and stop grounding checks. Existing VLN training data contains relatively few branching points, giving limited supervision for selecting the intended path from panoramic observations; this work densifies sampling around route decisions and termination during data construction. At matched trajectory counts against ScaleVLN and a ScaleVLN-rewrite variant using the same instruction pipeline, improved instructions alone bring gains and the full pipeline yields further gains across training scales; adding the data raises SR by a further 3.4/3.9 points and SPL by 2.7/2.0 points.

Geometric features extracted by PanoVGGT from the same RGB panorama are aligned with the VLM's semantic features by ERP region and fused residually, capturing scene content and spatial layout without adding visual tokens. VLM visual features mainly capture semantics; this work adds panoramic geometry without depth input and without increasing token count, linking landmarks to neighboring passages and directions. With the same fusion design, visual-token budget, and CGE, frozen UniK3D, DA2, DAP, and PanoVGGT encoders are compared; alternative encoders have mixed effects while PanoVGGT improves SR on both benchmarks.

Perspective

This work targets embodied agents that follow natural-language instructions through indoor and outdoor continuous environments, especially robot platforms equipped with panoramic cameras and relying on synchronous calls to a remote policy. It shows that converting wider visibility into performance requires coordinated changes in supervision horizon, execution length, and visual representation, so it is directly relevant to navigation systems using panoramic or wide-field sensors and to real-world deployments that need fewer policy calls and pauses. Simulation evaluation is limited to the Val-Unseen splits of R2R-CE and RxR-CE, Matterport3D scenes, and the Habitat platform; real-world evaluation uses a Unitree Go2, an Insta360 X5, and a remote RTX 3090, comparing five methods on 20 shared instruction-route pairs per setting without scene-specific fine-tuning.

Several numbers in the loaded text appear as placeholders, such as success rates and improvement margins in the abstract and experiment subsections, horizon values, the CGE budget and minimum execution steps, residual scale, learning rates, and network overhead; some can only be inferred indirectly from other passages or appendix tables, so reproducing exact hyperparameters and per-metric numbers requires checking the original. The real-world portion uses 20 shared instruction-route pairs per setting, a limited sample, so stability across more scenes and longer routes remains an open question. The geometry encoder is used frozen and the residual scale for semantic-geometric fusion is fixed, leaving behavior under other sensor configurations or outdoor lighting conditions to be observed.

Sources