Skip to main content
Back to timeline
arXivSource publication:

SLLCP calibrates frozen vision-language models into safety prediction sets, raising collision-trajectory flagging from 4.6% and 39.1% to 89.6% and 88.4% over 15k CARLA trajectories

Related research and updates

Synopsis

The authors propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen vision-language models that transforms their unreliable predictions into probabilistically calibrated safety prediction sets, with label-conditional finite-sample distribution-free coverage under exchangeability; over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, whereas the base VLMs flagged only 4.6% and 39.1%, respectively.

Source-provided article image: Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving
Fig. 1 ·

Fig. 1: Uncertainty in VLM-based trajectory safety predictions ( safe or unsafe ) varies across driving scenes. SLLCP upweights calibration examples (bottom) near each test input (top) in the projected embedding space and calibrates each safety label separately to address class imbalance in the collected data (e.g., due to unsafe cases being rarer). Trajectory footprints are shown in blue, with yellow waypoints spaced at 0.5 0.5\, s intervals.

arXiv

Interpretation

SLLCP is proposed as a post-hoc conformal prediction layer over frozen vision-language models that converts their approximate predictions about driving-trajectory safety into probabilistically calibrated safety prediction sets. Classical approaches are limited by the quality of their forecasting model, and VLMs, though promising for reasoning about the consequences of high-level actions, produce approximate predictions unsuitable for safety-critical autonomous driving; SLLCP calibrates on top of model outputs rather than retraining them. The abstract describes the layer as post-hoc, operating over frozen VLMs, and producing probabilistically calibrated safety prediction sets; implementation details are not given.

A localized procedure is introduced that upweights relevant past experience when calculating uncertainty thresholds, reflecting that the ability to estimate safety can depend on the observed driving scene. Scene dependence is made explicit in threshold computation rather than using a single threshold across scenes. The abstract describes a localized procedure that upweights relevant past experience; the weighting scheme and scene partition criteria are not specified.

Label-conditional finite-sample distribution-free coverage is provided under exchangeability. The method is given a theoretical coverage property rather than only empirical performance. The abstract states label-conditional finite-sample distribution-free coverage under exchangeability; no theorem numbers or proof details are given.

Over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, while the base VLMs flagged only 4.6% and 39.1%, respectively. Missed collision-causing trajectories are substantially reduced relative to the base VLMs, and the effect is observed across two different backbones. The abstract reports 15k CARLA trajectories, unseen scenarios, and comparison numbers for two backbones; false-positive rates, threshold settings, and statistical uncertainty are not reported.

Perspective

The work targets safety monitoring of planned trajectories in autonomous driving, suited to settings where a vision-language model serves as the safety prediction source and calibrated prediction sets are wanted without retraining the model; its localized procedure addresses the case where safety-estimation ability varies with the observed driving scene, and the coverage guarantee holds under exchangeability. The abstract shows validation across two different backbones (Qwen and Cosmos), indicating backbone substitutability for this post-hoc layer; the evaluation setting is CARLA trajectories and the collision-causing flagging task.

The abstract does not report false-positive rates, prediction-set sizes, threshold settings, how scenes are partitioned, or the composition of the 15k trajectories, so the trade-off between missed and false alarms cannot be judged; the coverage guarantee relies on exchangeability, and the relationship between unseen-scenario evaluation and exchangeability is not developed in the abstract; the two backbones have quite different baselines (4.6% and 39.1%), and the abstract does not explain the source of that gap. These are open questions at the abstract level.

Sources