安全路由评测在分布偏移下失真:选择成本可达路由全部收益,且攻击者可压低模型识别率19.6分
核心概要
该工作量化了安全路由基准中“在评测数据上挑选最佳单模型作为对照”所带来的选择成本:在 HELM Safety 上随机划分时为 0.003–0.030、留出类别时为 0.045–0.113,与归因于路由的全部差距相当,在 AgentDojo 留出套件时上升七到九倍;按偏移诚实评分后路由收益很小,并发现攻击者若知道所面对模型可将 GPT-5.4 的判定识别率压低 19.6 分。
Fig. 1: The baseline, not the router, decides the verdict. Phase 1, evaluation. One fitted router is scored against two fixed-model comparators on identical folds (E24, union label). The fold protocol shows where each comparator is chosen, the honest pin on the four training folds, the in-sample pin on the held-out fold it is then scored on. Against the in-sample pin the router never beats it. Against the honest pin it ties in the median and is ahead on average. The gap is the selection cost, 0.043 to 0.113 0.0430.113 , whose sign is guaranteed (Observation 1) and whose size depends on the judge (E24j). Phase 2, deployment. The router sees only the request text and decides once, before any tool output exists. What routing can buy, each on its own surface, is small. The fitted router serves the honest pin’s own model on every held-out request in 91 true % 91\text{true}\mathrm{\%} of two-model pool cells, a perfect pre-dispatch router on AgentDojo stays within two points of harm (E43), and the best honest cascade is one cheap model (E34). After the router commits, an injection arrives in a tool result and defence rests on the model’s recognition. On held-out reruns a targeted template lowers gpt-5.4 ’s judged recognition by 19.6 19.6 points, confirmed by an independent label (E54b, E55). Four action-level policy settings record zero judged successes on one shared set of episodes.
arXiv深度剖析
论文把路由基准中“对照模型在评测数据上选出”这一做法的代价量化到危害与准确率两个尺度上,并给出乐观偏差加偏移相关遗憾的界。 此前工作只证明该偏差的方向,本文给出其大小,并显示在留出划分下更大。 在 HELM Safety 上报告选择成本为随机划分 0.003–0.030、留出类别 0.045–0.113,且方向在任一已发布评判器单独使用时均成立;AgentDojo 留出套件时上升七到九倍。
在按偏移诚实评分后,路由在这些基准上收益有限:多数池单元中嵌套路由最终服务的是诚实基线的模型,近乎饱和的 AgentDojo 上完美预派发路由最多值两个点的危害。 把路由收益放到不含测试标签的基线之下重新衡量,而非与在评测数据上选出的最佳单模型比较。 结论来自 HELM Safety 与 AgentDojo 的池单元与近乎饱和语料上的评分结果。
模型对晚期注入的“表达出的识别”可被操纵:在留出重跑中,知道所面对模型的攻击者把 GPT-5.4 的判定识别率压低 19.6 分,并由独立标签确认。 把识别类防御放到“攻击者选择模型所见内容”的条件下评分,而非仅在固定输入下测量识别。 留出重跑给出 19.6 分的下降,并有独立标签确认;离线反事实组合进控制器时,同一攻击依回退模型不同而抬高或压低估计危害。
跨七个按预先固定规则选出的安全语料,三个通过注册的区间检验,四个优于事后置换零假设,四个区间未通过中有三个是某些模型观测危害为零的语料。 用预先注册的检验规则与事后置换零假设两种方式检验选择成本的存在性,并指出零观测危害语料是区间检验失效的常见情形。 七个语料、注册区间检验与置换零假设两类检验,以及零观测危害语料的计数。
启示与展望
该工作面向安全路由与识别类防御的评测设计者与使用者:在存在分布偏移的设定下,应使用不含测试标签选出的基线,并以危害为尺度对能选择模型所见内容的攻击者评分识别类防御。其量化结果适用于所测的 HELM Safety、AgentDojo 与七个安全语料,以及所报告的评判器与标签条件;离线反事实组合进控制器的结果适用于所考察的回退模型配置。
本文为摘要级阅读,未包含图表与完整实验细节,因此选择成本在不同评判器下的具体分解、七个语料各自的检验结果、以及离线反事实组合中回退模型如何决定危害方向,仍需查阅原文确认。此外,识别率下降 19.6 分是在留出重跑与特定攻击者设定下测得,其在其他模型与部署条件下的表现是开放问题。
