Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

GeoScaffold internalizes depth, connectivity, and traversability into a Video-LLM navigation policy via training-time reconstruction, consistently outperforming leading vision-only navigators on continuous VLN benchmarks

The work proposes GeoScaffold, a geometric supervision framework that first learns a compact depth tokenizer on depth maps from training trajectories and freezes it, then fine-tunes the policy with a handful of learnable geometry query tokens whose hidden states reconstruct navigation-critical geometry such as depth, connectivity, and traversability, thereby internalizing geometry into the policy itself; after training the tokenizer, target generators, and reconstruction heads are discarded, leaving the backbone and action interface unchanged, and experiments show it consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.