Skip to main content
Back to timeline
arXivSource publication:

Safety-routing evaluations distort under distribution shift: selection cost can match routing's entire benefit, and an attacker can cut judged recognition by 19.6 points

Synopsis

The work sizes the selection cost of picking the best single model as comparator on the evaluation data in safety-routing benchmarks: 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories on HELM Safety, comparable to the whole deficit attributed to routing, rising seven- to ninefold on AgentDojo when suites are held out; scored honestly under shift, routing buys little, and an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points.

Source-provided article image: False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift
Fig. 1 ·

Fig. 1: The baseline, not the router, decides the verdict. Phase 1, evaluation. One fitted router is scored against two fixed-model comparators on identical folds (E24, union label). The fold protocol shows where each comparator is chosen, the honest pin on the four training folds, the in-sample pin on the held-out fold it is then scored on. Against the in-sample pin the router never beats it. Against the honest pin it ties in the median and is ahead on average. The gap is the selection cost, 0.043 to 0.113 0.0430.113 , whose sign is guaranteed (Observation 1) and whose size depends on the judge (E24j). Phase 2, deployment. The router sees only the request text and decides once, before any tool output exists. What routing can buy, each on its own surface, is small. The fitted router serves the honest pin’s own model on every held-out request in 91 true % 91\text{true}\mathrm{\%} of two-model pool cells, a perfect pre-dispatch router on AgentDojo stays within two points of harm (E43), and the best honest cascade is one cheap model (E34). After the router commits, an injection arrives in a tool result and defence rests on the model’s recognition. On held-out reruns a targeted template lowers gpt-5.4 ’s judged recognition by 19.6 19.6 points, confirmed by an independent label (E54b, E55). Four action-level policy settings record zero judged successes on one shared set of episodes.

arXiv

Interpretation

The paper sizes, on both harm and accuracy, the cost of selecting the comparator model on the evaluation data in routing benchmarks, and bounds it by optimism plus a shift-dependent regret. Prior work proves the direction of this bias; this work gives its magnitude and shows it is larger under the held-out splits measured. On HELM Safety the selection cost is reported as 0.003-0.030 under random splits and 0.045-0.113 under held-out categories, with its direction holding under either published judge alone; it rises seven- to ninefold on AgentDojo when suites are held out.

Scored honestly under shift, routing buys little on these benchmarks: in most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. Routing benefit is re-measured against a baseline chosen without the test labels rather than against the best single model picked on the evaluation data. The conclusion comes from scoring results on HELM Safety and AgentDojo pool cells and the nearly saturated corpus.

A model's expressed recognition of a late injection is steerable: on held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. Recognition-based defences are scored against an attacker who chooses what the model sees, rather than measuring recognition on fixed inputs alone. Held-out reruns give the 19.6-point drop with independent-label confirmation; in an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model.

Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. The existence of the selection cost is tested both by a registered interval test and by a later permutation null, and zero-observed-harm corpora are identified as a common setting for interval-test misses. Seven corpora, two test families (registered interval test and permutation null), and a count of zero-observed-harm corpora.

Perspective

The work addresses designers and users of evaluations for safety routing and recognition-based defences: under distribution shift, use a baseline chosen without the test labels, and score recognition-based defences on harm against an attacker who can choose what the model sees. Its quantitative results apply to the measured HELM Safety, AgentDojo, and seven safety corpora, and to the reported judge and label conditions; the offline counterfactual composition into a controller applies to the fallback-model configurations examined.

This is an abstract-level reading without figures or full experimental detail, so the breakdown of selection cost across judges, the per-corpus test outcomes for the seven corpora, and how the fallback model determines the direction of harm in the offline counterfactual composition remain open questions to confirm in the original. In addition, the 19.6-point recognition drop is measured under held-out reruns and a specific attacker setting, and its behaviour under other models and deployment conditions is an open question.

Sources