GeoCoTDrive grounds planning-critical 2D regions first, then retrieves localized 3D priors, improving safety-critical planning metrics on nuScenes, NAVSIM and Bench2Drive
Synopsis
The work proposes GeoCoTDrive, an explicit geometric chain-of-thought framework for vision-language-action driving: it first autoregressively grounds 2D regions tied to ego decisions, then samples localized 3D geometry tokens from a frozen geometric foundation model inside those regions and interleaves them into the autoregressive context for trajectory generation; it also formulates planning-relevant grounding and builds the 146K-annotation PlanningGrounding dataset, improving safety-critical planning metrics on nuScenes, NAVSIM and Bench2Drive.
Figure 1: Comparison of geometry integration paradigms for VLA . (a) Structure-perception fusion injects agent and map tokens. (b) Geometric fusion incorporates global geometric tokens. (c) GeoCoTDrive introduces an explicit geometric chain-of-thought pipeline: VLM first grounds planning-relevant regions, then retrieves localized 3D geometric priors, and finally interleaves the grounding text with geometry tokens to condition trajectory generation.
arXivInterpretation
GeoCoTDrive models VLA planning as an explicit geometric chain of thought that thinks with 2D first and drives with dedicated 3D priors: the model autoregressively generates planning-relevant 2D regions, samples localized geometry tokens from a frozen geometric foundation model's feature map according to those regions, and interleaves them into the autoregressive context before continuing to decode the trajectory. Prior spatial-aware VLAs mostly introduce geometry as global scene-level representations through implicit feature fusion, such as VGGDrive injecting cross-view VGGT features, SpaceDrive converting UniDepth 3D coordinates into positional encodings, and LaST-VLA distilling VGGT priors into continuous latent CoT states; here geometry retrieval is explicitly derived from 2D decision-critical regions the model itself generates. The method section gives a complete autoregressive formulation chain: region generation, region-based sampling with a regular grid and bilinear interpolation inside each box, an MLP alignment head projecting features into language space, concatenation and interleaving of geometry tokens, and a teacher-forced objective supervising only grounding and trajectory responses, with the geometric foundation model kept frozen.
The paper formulates planning-relevant grounding, a region-level grounding task, and constructs PlanningGrounding, a VQA-style dataset with 146K planning-relevant annotations across five categories (critical-object, road-boundary, conflict-area, occluded-unknown-area, dense-object-area), plus 0.6K manually refined test samples. Conventional object-level grounding is category-driven and built on human-annotated objects with explicit semantic categories, which may overlook planning-critical cues such as conflict areas, occluded regions and dense agent clusters; PlanningGrounding shifts the target from semantically defined objects to decision-critical regions conditioned on the driving command. Each sample combines a front-view image, ego status and a high-level command; a planning-oriented prompt guides Gemini-3.1 to identify regions and assign planning-related categories, followed by automatic format validation and human verification. Raw data comes from nuScenes (28K), NAVSIM (102K) and Bench2Drive (16K), with road-boundary at 37.33% and critical-object at 34.16% dominating the category distribution.
On nuScenes open-loop planning, GeoCoTDrive reaches 0.32 m average L2, 0.11% average collision rate and 2.02% average intersection rate, competitive among VLA-based methods with favorable safety-related performance; on NAVSIM it reaches PDMS 87.6 with a Qwen2.5-VL-3B backbone versus 85.9 for the same-backbone baseline; on Bench2Drive closed-loop it reaches DS 78.68 and SR 55.42, above ORION and SpaceDrive. On NAVSIM the gain over the Qwen2.5-VL baseline is described as a modest PDMS improvement alongside better safety-related sub-metrics, and the paper notes the method remains competitive while using a smaller VLM backbone. Results follow the official protocols of three benchmarks: nuScenes reports L2, collision rate and intersection rate; NAVSIM reports PDMS with NC, DAC, TTC, Comf. and EP sub-scores; Bench2Drive reports DS and SR. Training uses 8 A800 GPUs with two stages of 3 epochs each.
Ablations show that both how geometry is injected and which grounding labels supervise it matter: directly interleaving global geometry tokens yields limited gains, whereas retrieving localized priors from planning-relevant regions gives more consistent improvements, and replacing perception labels with PlanningGrounding lowers nuScenes average collision rate from 0.12% to 0.07% and average intersection rate from 2.20% to 1.44%. Perturbing grounding results shows higher grounding mIoU consistently improves safety-related metrics; although OmniDrive reaches grounding mIoU comparable to an intermediate GeoCoTDrive setting, its downstream safety metrics remain worse, which the authors read as evidence that gains come from explicitly coupling 2D grounding evidence with localized 3D priors. Ablations cover geometry integration strategy, grounding annotation source, geometric foundation model (both VGGT and DA3-LARGE beat the baseline, with DA3-LARGE slightly better overall and VGGT marginally higher on EP), sampling grid size (16 and 25 both reach PDMS 87.3), and the two-stage training (either stage alone underperforms the combination).
Perspective
The result applies to end-to-end planning settings that take a single front-view image, ego status and a high-level driving command as input, under the open-loop and closed-loop protocols of nuScenes, NAVSIM and Bench2Drive. Two kinds of readers benefit directly: researchers who want to add localized geometric conditioning to an existing VLA planner without touching the geometry backbone, since the geometric model stays frozen and only the LLM and a lightweight alignment head are trained; and researchers who need region-level planning supervision, for which PlanningGrounding provides command-conditioned annotations in five categories. The paper states that future work will extend GeoCoTDrive to multi-view and temporal settings and improve grounding robustness.
The stated scope is single front-view input without explicit multi-view or temporal modeling, which limits coverage of side and rear traffic, prolonged occlusions and interaction histories; localized geometry retrieval also depends on grounding accuracy, and incorrect or incomplete regions may yield less relevant geometric cues and impair trajectory prediction. A careful reader would still watch whether the quantitative link between grounding quality and downstream safety metrics holds across benchmarks; whether the grid-size saturation (16 and 25 both reaching PDMS 87.3) persists at larger scale or higher input resolution; and whether the respective contributions of driving-knowledge adaptation and localized geometry conditioning remain stable across base models. Some table values are not fully rendered in the text, so exact reproduction should be checked against the appendix and the released code.
