LADE extracts cross-model safety signals from first-token dark knowledge, lowering jailbreak success while preserving the safety-utility balance
Synopsis
LADE extracts model-agnostic latent safety signals from dark knowledge in the first-token output probability distribution by contrasting harmful and benign queries, maps the selected tokens across tokenizers, and filters harmful queries via kNN distance before generation; across Llama-2-7B-Chat, Llama-3-8B-Instruct, Qwen2-7B-Instruct, Qwen3-8B, Gemma-7B-it, and Mistral-7B-Instruct-v0.3, it attains the lowest or near-lowest average compliance under multiple jailbreak attacks and the best safety-utility average on six of eight models.
Figure 1 : The key idea of LADE . (Left) Latent safety signals ( red points), identified by contrasting first-token probabilities between harmful and benign queries, are extracted from a reference LLM and transfer across target LLMs, while model-specific surface tokens ( gray points) remain scattered. (Right) Refusal tokens capture only a small subset of safety signals and fail to separate harmful from benign queries (top), whereas adding extracted latent safety signals yields clear separation (bottom).
arXivInterpretation
The authors find that the first-token output probability distribution of safety-aligned LLMs contains latent safety signals, defined as tokens whose probabilities differ most between harmful and benign queries, and that these signals align consistently across safety-aligned models, forming a model-agnostic discriminative direction. Prior decoding-stage defenses rely on internal hidden states or surface refusal tokens such as "Sorry" and "As"; this work locates the discriminative information in the dark knowledge of the first-token distribution and provides empirical evidence of cross-model transfer. Within the Gemma family, average classification accuracy rises from 0.6836 on the base Gemma-7B to 0.9396 on the instruction-tuned Gemma-7B-it, 0.9474 on Gemma-2-9B-it, and 0.9647 on Gemma-3-4B-it; cross-model reference-target transfer accuracy is comparable to using the same model as both reference and target.
LADE comprises three components: extracting latent safety signals from dark knowledge, tokenizer mapping, and kNN-based discrimination; it operates entirely on output probabilities without accessing internal hidden states and makes a single filtering decision before generation begins. Unlike SafeDecoding, RDS, and SafeInfer, which adjust token probabilities or hidden states at every decoding step, LADE performs one pre-generation filtering decision and therefore requires no architecture-specific auxiliary module. The tokenizer-mapping ablation shows that enabling both subword split handling and duplicate-mapping ratio estimation yields the highest and most uniform cross-model transfer accuracy; ratio estimation alone improves the average but leaves noticeable degradation, indicating that subword split handling is the dominant factor for cross-tokenizer transfer.
Across six to eight open-source LLMs and multiple jailbreak attacks, LADE achieves the lowest or near-lowest average compliance and outperforms existing defense baselines on the safety-utility trade-off. Self-Reminder and SafeDecoding collapse on weakly aligned models such as Gemma-7B-it and Mistral-7B-Instruct-v0.3 or induce substantial over-refusal, whereas LADE maintains low compliance and low refusal counts on these models. In Table 1, LADE attains the lowest average compliance on Qwen2-7B-Instruct, Gemma-7B-it, and Mistral-7B-Instruct-v0.3; in Table 2, LADE achieves the best average on six of eight models, for example 1.60 on Llama-2-7B-Chat versus 5.00 for RDS.
Compared with dedicated guard models, LADE achieves the highest average accuracy of 0.969 on Llama-3-8B-Instruct and offers advantages in latency and memory. Guard models are standalone classifiers trained on safety annotations, whereas LADE reuses the target LLM's first-token forward pass and requires no additional model or safety-specific training. In Table 3, LADE's average accuracy of 0.969 exceeds WildGuard-7B at 0.947, Llama-Guard-4-12B at 0.852, and Prompt-Guard-2-86M at 0.575; in Table 4, LADE uses the least memory (14.96 GiB) and is several times faster than Llama-Guard-4-12B and WildGuard-7B.
Perspective
LADE applies to deployments of safety-aligned LLMs where first-token output probabilities are accessible, such as open-weight models or interfaces that expose the probabilities of the few most likely first tokens; its offline stage uses Hex-Phi as the harmful extraction and reference set and XSTest as the benign extraction set, with Llama-3-8B-Instruct as the reference model, and the online stage makes one pre-generation filtering decision per query. On weakly aligned models, discrimination accuracy decreases, so LADE is positioned as a complement to, not a substitute for, safety alignment. For closed APIs, Appendix I shows that keeping only the few most likely first tokens reproduces nearly all filtering decisions, indicating the route remains viable under restricted probability access.
LADE depends on the target model having safety alignment, and discrimination accuracy is lower on weakly aligned models; on Mistral-7B-Instruct-v0.3 it induces a higher refusal count on Alpaca than the no-defense baseline, indicating that the harmful reference manifold partially overlaps with some conversational benign queries; against a white-box adaptive attack that knows every component of LADE, accuracy drops from 1.000 to 0.480, which, while still above the compared guard models, is explicitly described by the authors as covering a single controlled white-box attack and not establishing robustness to all adaptive attacks. In addition, for indirect or context-dependent harmful queries, such as those involving false medical claims or specific predictions about future events, the first-token distribution does not exhibit clear discriminative patterns, leaving such borderline cases an open question.
