TellTail 仅凭查询即可从黑盒检索系统中识别出 53 个嵌入模型中的部署模型
相关研究与后续进展核心概要
该工作提出 TellTail,一种仅通过查询来识别黑盒检索系统背后嵌入模型的指纹攻击:当系统暴露完整排序结果时可完美识别,暴露无序 top-3 结果时成功率为 94.3%,仅暴露语言模型生成答案时成功率为 92.5%,并在 OpenWebUI 端到端部署中仍能识别 53 个检索器中的 44 个。
Fig. 1: Overview of the fingerprinting procedure. Depending on the selected TellTail variant, fingerprint construction follows one of two paths. For generic-query variants, TellTail-Random and TellTail-Topic (green arrows), we first submit the benign queries to the target system, use the observed outputs to construct a proxy corpus for each candidate model, and then run the same queries against the proxy corpus to derive per-candidate sets of retrieved passages which act as fingerprints. The model-specific query variant, TellTail-OPT (red arrows), optimizes model-specific query suffixes toward selected targets to construct these fingerprints. The resulting queries are then submitted to the black-box retrieval system, and the exposed behavior is compared with the candidate fingerprints to infer the deployed retriever. Depending on the available interface, the observed behavior may consist of retrieved passages or evidence from post-retrieval processing (e.g., LLM response).
arXiv深度剖析
TellTail 表明,尽管嵌入模型在重叠数据上训练、目标相似且表征被认为趋于收敛,它们仍可被引导产生模型特有的检索结果,从而仅凭查询即可识别部署的检索器。 此前针对机器学习系统的指纹研究覆盖了分类器、LLM 版本和文生图模型,但识别已部署检索系统背后的嵌入模型尚未被探索。 在 53 个检索器、三种访问级别下评估,攻击者仅能访问其中 19 个候选模型;完整排序下完美识别,无序 top-3 下 94.3%,仅生成答案下 92.5%。
TellTail 提供两种互补策略:通用查询指纹通过比较检索重叠来推断模型,模型特定优化查询则诱导目标检索器产生特定行为而对其他模型迁移性差。 该工作将通常被视为攻击泛化缺陷的弱迁移性重新定位为指纹识别的有用特性,并采用 token 屏蔽来抑制跨模型迁移。 token 屏蔽将非目标模型的平均 appeared@ 迁移率从 56.2% 降至 3.0%,同时目标模型上的成功率仅从 99.4% 降至 91.4%。
在真实 RAG 部署 OpenWebUI 中,TellTail-OPT 仍能识别部署的检索器,尽管存在文档分块、查询生成和重排序。 该案例研究将受控的响应级评估扩展到具有应用侧组件的实际 RAG 应用,并分离出各组件的影响。 端到端 ASR 为 83.0%(53 个中正确分类 44 个),低于理想化 RAG 的 92.5%;查询生成是主要退化来源,而重排序影响甚微。
启示与展望
该结果适用于检索主要由稠密嵌入驱动的系统,攻击者拥有候选检索器集合、对目标系统具有标准黑盒查询访问权限,并了解目标领域。它使提供方审计和攻击者侦察成为可能,并激励设计在保持检索质量的同时抑制模型特定信号的防御。
优化查询有时看起来不自然,且评估聚焦于稠密检索;向混合检索(如 BM25)扩展以及设计抑制指纹的防御仍是开放方向。端到端案例研究中的查询生成是主要退化来源,因此当应用重写查询时,指纹成功率可能下降。
