Skip to main content
Back to timeline
arXivSource publication:

GeoScaffold internalizes depth, connectivity, and traversability into a Video-LLM navigation policy via training-time reconstruction, consistently outperforming leading vision-only navigators on continuous VLN benchmarks

Related research and updates

Synopsis

The work proposes GeoScaffold, a geometric supervision framework that first learns a compact depth tokenizer on depth maps from training trajectories and freezes it, then fine-tunes the policy with a handful of learnable geometry query tokens whose hidden states reconstruct navigation-critical geometry such as depth, connectivity, and traversability, thereby internalizing geometry into the policy itself; after training the tokenizer, target generators, and reconstruction heads are discarded, leaving the backbone and action interface unchanged, and experiments show it consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.

Source-provided article image: GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation
Figure 1 ·

Figure 1: Comparison of end-to-end VLN paradigms. (a) Streaming Video-LLMs act directly from RGB, but lack 3D perception. (b) Geometry-enhanced policies add a spatial encoder, but it stays resident at inference and geometry is never internalized. (c) GeoScaffold internalizes geometry via training-only query-token supervision, leaving the deployed RGB-and-text action path unchanged.

arXiv

Interpretation

It proposes a framework that pays the geometric supervision cost once at training time: a compact depth tokenizer is trained and frozen, then a handful of learnable geometry query tokens have their hidden states trained to reconstruct navigation-critical geometry. Relative to existing geometry-aware extensions that rely on depth sensors, 3D encoders, or per-step perception tool calls, this method does not charge a persistent inference-time price but internalizes geometry into the policy. The abstract explicitly describes the training pipeline and supervision targets (depth, connectivity, traversability) and states that these components are no longer needed at inference.

The geometry query states become compact geometric latents used for action decoding, and through shared weights geometry is also internalized into the backbone's own representations. Geometry is moved from an external module into the policy's internal representations, so action decoding directly benefits from geometric latents. The abstract states that the supervision turns the query states into compact geometric latents and, through shared weights, internalizes geometry into the backbone's representations.

After training, the tokenizer, target generators, and reconstruction heads are discarded, leaving the backbone and action interface unchanged, forming a scaffold-like one-time structure. Compared with extensions that retain extra geometric modules, this method keeps the original interface at deployment, which is convenient for lightweight edge deployment. The abstract explicitly states that these components are discarded after training and that the backbone and action interface remain unchanged.

On continuous VLN benchmarks, GeoScaffold consistently outperforms leading vision-only navigators. It achieves better performance than vision-only baselines without requiring geometric inputs at inference time. The abstract summarizes the experimental conclusion with 'Extensive experiments' and 'consistently outperforms', without giving specific numbers, benchmark names, or sample sizes.

Perspective

The framework targets continuous VLN settings that use streaming Video-LLM policies mapping egocentric RGB observations and instructions directly to low-level actions, and its benefit comes from depth maps available in training trajectories. It applies to embodied navigation systems that want improved spatial awareness without introducing inference-time depth sensors, 3D encoders, or per-step perception tool calls, especially for lightweight edge deployment. Because the tokenizer, target generators, and reconstruction heads are discarded after training, the deployment side only needs to keep the original backbone and action interface.

The abstract summarizes results with 'Extensive experiments' and 'consistently outperforms' without listing specific benchmark names, evaluation metrics, improvement magnitudes, or sample sizes, so the magnitude and stability boundary of the advantage cannot be judged from the abstract. The compactness of the depth tokenizer, the number of geometry query tokens, the exact form of the reconstruction targets, and behavior under different training-trajectory depth quality all need confirmation in the main text. In addition, the abstract does not state applicability when reliable depth supervision is unavailable or in cross-domain scenarios; these are open questions worth watching.

Sources