Skip to main content
Back to timeline
arXivSource publication:

MapLightning replaces dense BEV grids with 1D map tokens, gaining +10.1 mAP over MapTRv2 on nuScenes at 1.73x faster inference

Related research and updates

Synopsis

MapLightning replaces the dense BEV grid with a compact set of 1D learnable map tokens as the intermediate representation for online vectorized HD map construction, where a transformer mapper concatenates map and image tokens, applies full self-attention, discards the image tokens, and retains the updated map tokens for decoding; this design uses up to 16.7x fewer intermediate tokens and achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse 2, with a lightweight variant surpassing MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse 2 while delivering 1.73x faster inference (40+ FPS) and 53% less memory, and further showing improvements on uncertainty-aware map construction and downstream trajectory prediction.

Source-provided article image: MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens
Figure 3 ·

Figure 3: Image-to-image attention in the self-attention mapper. We select one image token as the query (magenta cross) and visualize its attention scores over the other image tokens, which serve as keys (red: high, blue: low), at the first and last ( L = 4 L=4 ) mapper layers. From left to right: (1) a query on a lane divider attends along that divider and to the parallel divider across the lane; (2) a query on a lane divider attends along the divider’s full extent toward the horizon; (3) a query on a pedestrian crossing attends across the crossing’s full width; and (4) a query on a road boundary attends along the curb. From Layer 1 to Layer 4, responses on these structures become stronger and more complete, while diffuse background responses (e.g., the sky in column 4) fade, illustrating progressive refinement of the image tokens. Such image-to-image interaction is absent in cross-attention, where image tokens serve only as fixed keys and values. Best viewed zoomed in.

arXiv

Interpretation

A compact set of 1D learnable map tokens replaces the dense BEV grid as the intermediate representation, and map tokens are concatenated with image tokens for full self-attention, after which the image tokens are discarded and only the updated map tokens are retained for decoding. Prior online vectorized HD map construction methods typically rely on dense BEV grids as the intermediate representation; this work switches to compact 1D map tokens and chooses self-attention over vanilla cross-attention to enable joint interactions and contextual aggregation among image and map tokens. The abstract reports that this representation uses up to 16.7x fewer intermediate tokens than dense BEV-based methods and states that it achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse 2.

The lightweight design allows the map decoder to use full rather than deformable cross-attention for better global context. In BEV-based methods the decoder is typically constrained to deformable cross-attention; this work states that its lightweight representation permits full cross-attention. The abstract lists this as the second advantage of the design, an argument at the design level without a separate ablation number.

The network does not use camera projection parameters, making it robust to camera-extrinsic perturbations. Unlike BEV-based methods that depend on camera projection parameters, this design removes that dependency at the representation level. The abstract states this property by contrast with BEV-based methods and does not give specific numbers for perturbation experiments.

The lightweight variant surpasses MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse 2 while delivering 1.73x faster inference (40+ FPS) with 53% less memory, and further shows improvements on uncertainty-aware map construction and downstream trajectory prediction. Relative to the existing MapTRv2 baseline, this work reports simultaneous gains in accuracy, speed, and memory, and extends to uncertainty-aware construction and trajectory prediction. The abstract provides concrete numbers across two datasets, including the mAP deltas on nuScenes and Argoverse 2, 1.73x faster inference, 40+ FPS, and 53% less memory, and states that code and models will be released.

Perspective

The work targets online vectorized HD map construction within autonomous-driving perception, applicable to settings with multi-camera image input that require real-time output; the abstract states state-of-the-art accuracy and efficiency on nuScenes and Argoverse 2 and reports the lightweight variant's mAP gains over MapTRv2, 1.73x faster inference, 40+ FPS, and 53% less memory. At the representation level it does not use camera projection parameters, making it robust to camera-extrinsic perturbations, which is relevant to deployment settings with unstable camera calibration. The abstract also mentions improvements on uncertainty-aware map construction and downstream trajectory prediction, suggesting the representation can be reused for map-related downstream tasks. Code and models will be released, enabling reproduction and transfer to proprietary data.

The currently readable scope is the abstract only, which does not include network architecture details, training configuration, ablations, or the exact setup of the camera-extrinsic perturbation experiments, so it is not possible to attribute each advantage to a specific design choice. The 16.7x token reduction, the +10.1 and +16.2 mAP gains, the 1.73x speedup, 40+ FPS, and 53% memory saving are taken as stated in the abstract, and their measurement conventions and comparison conditions need to be checked in the original text. The magnitude of improvements on uncertainty-aware map construction and downstream trajectory prediction is not given numerically in the abstract. Code and models are stated as forthcoming in the abstract, so their actual availability and reproduction conditions remain to be confirmed.

Sources