Skip to main content
Back to timeline
arXivSource publication:

ShieldVLA gates VLA fine-tuning with Hamilton-Jacobi reachability, cutting cumulative safety cost by 57% and lifting success rate by 0.13

Synopsis

ShieldVLA is a safety-aligned fine-tuning framework for vision-language-action (VLA) models that learns a model-free Hamilton-Jacobi reachability value function from visual observations as a safety critic, uses that critic to gate policy optimization into reward maximization inside feasible regions and recovery near unsafe states, and replaces manual per-step cost labels with rubric-based VLM safety scores; across five navigation and manipulation benchmarks (Dubins-VL, TurtleBot-Nav, Safety-CHORES Nav/Fetch, Franka-Reach) and multiple VLA backbones it reduces cumulative safety cost by 57% on average and improves task success rate by 0.13 over SafeVLA.

AI-generated editorial illustration: ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

Interpretation

A feasibility-gated fine-tuning framework for VLA models: a model-free Hamilton-Jacobi reachability critic estimates whether the current state still permits safe recovery, and safety intervention is applied only on states the critic flags as infeasible, replacing the global reward-cost trade-off of Lagrangian methods. Prior work such as SafeVLA enforces safety as a soft penalty on expected cumulative cost and lacks a per-trajectory feasibility signal; ShieldVLA reframes safety as a state-dependent feasibility problem and applies HJ reachability to fine-tuning foundation models rather than only as a runtime filter. The paper gives a proposition and proof that the safety Bellman operator is a contraction in the supremum norm (Appendix A), and a gating-versus-penalty ablation on TurtleBot-Nav with the same critic: gating reaches SR 0.54 / CSC 6.90 versus SR 0.33 / CSC 7.80 for the Lagrangian penalty.

Rubric-based VLM safety supervision: semantic safety feedback is converted through structured rubrics and calibration into per-frame safety margins, removing the need for manually annotated per-step cost labels. Free-form VLM judgments are noisy and inconsistent for stable policy optimization; the work transplants LLM/VLM-as-a-judge and rubric-based reward modeling to safety supervision, calibrating with a per-episode binary safety label via class-balanced logistic regression and adding an offline Refinement-through-Differentiation loop for tail failure modes. On TurtleBot-Nav the 2B scorer shows mode collapse (AUROC 0.500), while switching to the 8B scorer plus RTD-added rubrics raises it to 0.608; on Safety-CHORES Nav the 8B scorer reaches per-frame AUROC 0.548 and Spearman +0.272 against log(1+cost), versus +0.094 for 2B.

Improved safety-performance trade-offs are demonstrated across five navigation and manipulation benchmarks and multiple VLA backbones, and hold up under test-time visual perturbations. Evaluation spans Dubins-VL, TurtleBot-Nav (OmniVLA), Safety-CHORES Nav and Fetch (SPOC-VLA), and Franka-Reach (OpenVLA-OFT), and additionally reports three families of out-of-distribution visual perturbation (+Color, +Light, +All) rather than in-distribution metrics alone. In the main table ShieldVLA pairs the lowest collision rates with the highest success rates in nearly all environments, e.g. TurtleBot-Nav SR 0.54 / CSC 6.90 versus SafeVLA's 0.34 / 14.6; standard deviations across 3 training seeds are small relative to the gap over SafeVLA point estimates; under perturbation ShieldVLA holds the lowest CSC in every (environment, condition) cell of Table 6.

Ablations attribute the gains to feasibility gating and the continuous calibrated cost signal rather than to the critic alone. Replacing the calibrated VLM rubric margin with a binary collision indicator while keeping the critic and gate identical collapses task performance on three of four environments, showing that a continuous severity signal is what lets the gate fire selectively rather than indiscriminately. The binary-cost ablation drops TurtleBot-Nav SR from 0.54 to 0.31 and Franka-Reach SR from 0.45 to 0.25; the gating-versus-penalty comparison on TurtleBot-Nav is superior on both axes.

Perspective

The results target research and engineering settings where a pretrained VLA policy is fine-tuned in simulation: Dubins-VL, TurtleBot-Nav, Safety-CHORES Nav/Fetch, and Franka-Reach cover wheeled navigation, mobile manipulation, and tabletop manipulation, with backbones including OmniVLA, SPOC-VLA, and OpenVLA-OFT. The method assumes access to a per-episode binary safety label for calibration and assumes sufficient near-boundary exploration in the replay data; VLM scoring is used only during training and no VLM calls are issued at test time. The authors name physical-robot validation (with domain-adapted recalibration of the rubric) and scaling the scorer to break the mode-collapse ceiling as the most immediate next steps.

The authors note that the classical HJ contraction does not transfer cleanly to the function-approximation, finite-data setting, so all safety claims are empirical; using a calibrated VLM rubric in place of the ground-truth signed distance adds a second source of error, so the learned safe set is the one induced by the rubric signal, and the gap to simulator dynamics is bounded only on Dubins-VL by the HJ (GT cost) row. The scorer has a capacity floor, with sub-8B models showing mode collapse on tail frames, and in environments with very rare unsafe events the gate threshold becomes effectively a hand-tuned hyperparameter. The head-to-head comparison is against a single published baseline (SafeVLA) because older safe-RL methods lack a published VLA instantiation. In addition, several numbers in the loaded text (such as the average cost reduction and success-rate gain in the abstract and conclusion, the mean and standard deviation cells of Table 3, and some cells of Tables 4 and 5) are missing from the markdown, so this summary relies on the 57% and +0.13 figures given in the external story plus the specific numbers that are readable in the body.

Sources