TellTail identifies the deployed embedding model among 53 retrievers from queries alone in black-box systems
Related research and updatesSynopsis
The work introduces TellTail, a query-only fingerprinting attack that identifies the embedding model behind a black-box retrieval system: it perfectly identifies the deployed retriever from full rankings, achieves 94.3% success from unordered top-3 results, and 92.5% from language-model-generated answers alone, and still identifies 44 of 53 retrievers in an end-to-end OpenWebUI deployment.
Fig. 1: Overview of the fingerprinting procedure. Depending on the selected TellTail variant, fingerprint construction follows one of two paths. For generic-query variants, TellTail-Random and TellTail-Topic (green arrows), we first submit the benign queries to the target system, use the observed outputs to construct a proxy corpus for each candidate model, and then run the same queries against the proxy corpus to derive per-candidate sets of retrieved passages which act as fingerprints. The model-specific query variant, TellTail-OPT (red arrows), optimizes model-specific query suffixes toward selected targets to construct these fingerprints. The resulting queries are then submitted to the black-box retrieval system, and the exposed behavior is compared with the candidate fingerprints to infer the deployed retriever. Depending on the available interface, the observed behavior may consist of retrieved passages or evidence from post-retrieval processing (e.g., LLM response).
arXivInterpretation
TellTail shows that embedding models, despite being trained on overlapping data with similar objectives and argued to converge in representation, can be steered to emit model-specific retrieval results, enabling identification of the deployed retriever from queries alone. Prior fingerprinting work covered classifiers, LLM versions, and text-to-image models, but identifying the embedding model behind a deployed retrieval system remained unexplored. Evaluated on 53 retrievers under three access levels with only 19 candidates accessible to the attacker; perfect identification from full rankings, 94.3% from unordered top-3, and 92.5% from generated answers alone.
TellTail offers two complementary strategies: generic-query fingerprinting compares retrieval overlap to infer the model, while model-specific optimized queries induce a chosen retrieval behavior on the target retriever but transfer poorly to others. The work reframes poor transferability, usually a bug limiting attack generalization, as a useful feature for fingerprinting, and uses token blocking to suppress cross-model transfer. Token blocking reduces mean appeared@ transfer on non-target models from 56.2% to 3.0% while only reducing success on the target model from 99.4% to 91.4%.
In a real RAG deployment, OpenWebUI, TellTail-OPT still identifies the deployed retriever despite document chunking, query generation, and reranking. The case study extends the controlled response-level evaluation to a practical RAG application with application-side components and isolates their effects. End-to-end ASR of 83.0% (44 of 53 correctly classified), down from 92.5% in the idealized RAG; query generation is the primary source of degradation, while reranking has little effect.
Perspective
The result applies to systems whose retrieval is primarily driven by dense embeddings, where the attacker has a candidate retriever set, standard black-box query access to the target system, and knowledge of the target domain. It enables provider auditing and attacker reconnaissance, and motivates defenses that suppress model-specific signals while preserving retrieval quality.
Optimized queries can look unnatural, and the evaluation focuses on dense retrieval; extending fingerprinting to hybrid retrieval such as BM25 and designing defenses that hinder fingerprinting remain open directions. Query generation in the end-to-end case study is the primary source of degradation, so fingerprinting success may drop when applications rewrite queries.
