Public articles linked to the same research event.
arXiv The authors propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen vision-language models that transforms their unreliable predictions into probabilistically calibrated safety prediction sets, with label-conditional finite-sample distribution-free coverage under exchangeability; over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, whereas the base VLMs flagged only 4.6% and 39.1%, respectively.
The authors propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen vision-language models that transforms their unreliable predictions into probabilistically calibrated safety prediction sets, with label-conditional finite-sample distribution-free coverage under exchangeability; over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, whereas the base VLMs flagged only 4.6% and 39.1%, respectively.
The authors propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen vision-language models that transforms their unreliable predictions into probabilistically calibrated safety prediction sets, with label-conditional finite-sample distribution-free coverage under exchangeability; over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, whereas the base VLMs flagged only 4.6% and 39.1%, respectively.
The authors propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen vision-language models that transforms their unreliable predictions into probabilistically calibrated safety prediction sets, with label-conditional finite-sample distribution-free coverage under exchangeability; over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, whereas the base VLMs flagged only 4.6% and 39.1%, respectively.