ViRA aligns planner features to frozen visual foundation models, lifting Rap∗ from 72.8 to 83.7 EPDMS on NAVSIM v2 navtest and yielding a 92.3-EPDMS ViRA-Diffusion
Synopsis
The authors introduce ViRA, a planner-agnostic visual representation alignment framework that supervises the planner encoder with frozen visual foundation model (VFM) intermediate features during training and discards the alignment modules at inference, leaving planner architecture and inference cost unchanged; across NAVSIM v2 and HUGSIM, ViRA consistently improves driving performance for regression-based, diffusion-based, and scoring-based planners, with gains depending on the VFM target and on whether the planner uses auxiliary perception supervision.
Figure 1: Better visual representations improve planning, with gains depending on both VFMs and planners. (a) ViRA aligns planner representations with a frozen VFM during training, leaving inference unchanged. (b) Segmentation and visual attention suggest improved representations. (c) Better visual representations improve planning, but the gains depend on the VFM target and planner.
arXivInterpretation
VFM-guided visual representation alignment consistently improves driving performance across planners and extends to zero-shot closed-loop evaluation. Prior work mostly studies a single planner–VFM pairing; ViRA provides a common interface to compare regression-based TransFuser, diffusion-based DiffusionDrive, and scoring-based Rap∗ under controlled alignment. On NAVSIM v2 navtest, EPDMS improves by 5.4, 7.4, and 10.9 points respectively; on navhard, Rap∗ rises from 32.8 to 49.2 (+16.4); zero-shot HUGSIM overall HD-Score improves by 2.7, 6.4, and 3.0 points.
The gains require a pre-trained target rather than alignment-loss regularization alone. A randomly initialized target is substituted under identical alignment objectives to isolate the contribution of VFM pre-training. Random-target alignment lowers TransFuser's BEV segmentation mIoU from 36.4 to 30.7 and EPDMS by 1.6 points, and lowers Rap∗ EPDMS by 1.7 points, whereas the pre-trained DVGT target improves both planners by 5.4 and 10.9 points.
The choice of VFM target matters, and targets can be complementary to a planner's existing representations. Holding the planner fixed and varying only the alignment target quantifies target-selection effects and tests semantic-versus-geometric complementarity. For perception-free Rap∗, EPDMS spans 5.5 points across five targets (79.4–84.9), with DINOv3 giving the largest gain; even with a DINOv3-L backbone, aligning to the geometry-focused VGGT further raises EPDMS from 83.8 to 86.3 (+2.5).
Auxiliary perception supervision reduces sensitivity to VFM target selection and mainly compensates for less effective targets. A controlled comparison of TransFuser with and without auxiliary perception supervision across the same five VFM targets. With auxiliary supervision the EPDMS spread is 0.5 points (88.8–89.3); removing it widens the spread to 2.7 points (86.4–89.1), with the maximum nearly unchanged (89.3 to 89.1) while the minimum drops (88.8 to 86.4).
Perspective
The work targets end-to-end planners trained and evaluated in simulation: NAVSIM v2 navtest covers 12,146 real-world scenarios, navhard uses 244 challenging real-world scenarios plus 4,164 synthetic follow-up scenarios rendered with 3D Gaussian Splatting, and HUGSIM provides zero-shot closed-loop evaluation. ViRA uses frozen VFM intermediate representations as representation-level supervision during training and discards the VFM and alignment heads at inference, so the deployed architecture and inference cost are unchanged—DVGT-guided alignment raises EPDMS from 72.8 to 83.7 while retaining the baseline 20 ms latency and 69M parameters, close to the strongest backbone replacement, DINOv3-L, which reaches 83.8 EPDMS but needs 65 ms and 351M parameters. This setting suits researchers and engineering teams wanting to add pre-trained visual knowledge to existing planners at low cost; the authors note that extracting and storing VFM representations may add training-time cost, which feature caching, lightweight distillation, or selective alignment could reduce.
Several open questions remain: the conclusions rest on three representative planners and simulation benchmarks, and the authors state that multi-modal sensor fusion, vision-language-action models, and real-vehicle deployment are not covered, so generalization to a wider set of planners and real-world data needs further validation; the training-time cost of extracting and storing VFM representations is not quantified, leaving the practical benefit of feature caching or selective alignment to be assessed; failure cases show that even with more complete BEV semantic maps, planned trajectories may still deviate from human experts in challenging scenarios, indicating that representation quality is not sufficient and that trajectory decoding, temporal reasoning, interaction modeling, and decision-making under uncertainty remain separate ingredients; and VFMs may inherit geographical, weather, cultural, and traffic-distribution biases from pre-training data, whose effect after transfer to planners requires cross-environment evaluation and transparent failure analysis.
