Skip to main content
Back to timeline
arXivSource publication:

Appending a classification instruction after the user turn lifts out-of-distribution malicious-input detection by up to about 4 AUC points

Related research and updates

Synopsis

Using strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B), this work compares post-user suffixes for activation probes and finds that a classification suffix consistently improves out-of-distribution detection on a single-position probe (up to about 4 AUC points), that the gain comes from the classification format rather than the named criterion, that a content-free suffix matches the real malicious/benign one, and that the benefit carries to production multi-position pooling probes (attention, multi-max, MLP) though the best suffix there is readout-dependent.

Source-provided article image: Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
Figure 2 ·

Figure 2: A classification suffix lifts OOD detection; format, not criterion, carries the ranking gain. Δ \Delta shared-benign AUC vs. no suffix (mean over 9 LODO datasets, ± \pm SEM). A neutral suffix ( reflect ) barely moves ranking; asking the model to classify lifts it (Llama, Gemma), and content-free A/B matches named intent . Qwen is near-ceiling. Precision (the criterion’s low-FPR gain): App. A.1 .

arXiv

Interpretation

Appending a classification instruction after the user turn improves activation-probe out-of-distribution detection on unseen attack types, by up to about 4 AUC points on a single-position probe. The LLM-as-judge-style suffix for probes had been treated as a cheap trick without systematic testing of wording or in-the-wild generalization; this work supplies a controlled suffix ladder under strict LODO evaluation. Leave-one-dataset-out evaluation across 13 safety benchmarks (jailbreak, injection, benign chat) and three open-weight model families, comparing single-position probes with and without a suffix.

The gain comes from the classification format itself, not the named criterion: a content-free suffix matches the real malicious/benign one, with the criterion adding precision only at strict thresholds. Separates the format effect of asking the model to represent the input as a class from the semantics of the criterion, showing the source of suffix effectiveness is the classification-style phrasing rather than a specific safety definition. Within the same controlled suffix ladder, compares classification suffixes, content-free-label suffixes, off-topic suffixes, and merely-attentive suffixes.

The benefit is not an artifact of the single-position read and carries to production multi-position pooling probes (attention, multi-max, MLP), though the best-performing suffix there is readout-dependent. Extends the conclusion from single-position probes to multiple pooling readouts used in deployment, while showing suffix choice is tied to the readout rather than universally optimal. Repeats the suffix comparison on attention, multi-max, and MLP pooling probes and reports that the best suffix varies with the readout.

Served through a KV-cache fork, the suffix is a cheap drop-in for any activation-probe monitor, though not an automatic win: which suffix helps, and by how much, depends on the model and the readout. Provides a deployment path and cost profile while making the conditionality of the benefit explicit, avoiding treating the suffix as a universal gain. Serves the suffix via a KV-cache fork and compares suffix gains across different models and readouts.

Perspective

The results target LLM-agent deployments that use activation probes as runtime monitors, for input-side interception of prompt injection, jailbreaks, and unsafe requests, with evaluation covering 13 safety benchmarks and three open-weight model families: Llama-3.1-8B, Qwen3.5-9B, and Gemma-4-12B. For engineering teams seeking to strengthen existing probes at very low serving cost, a classification suffix is a drop-in change worth trying and can be served through a KV-cache fork; however, suffix choice and the size of the gain need to be re-confirmed for the specific model and readout.

The size of the suffix gain and the best choice depend on the model and readout, so re-validation is still needed when transferring to a new model or readout; the precision difference between content-free-label suffixes and the real criterion at strict thresholds suggests threshold setting affects the value of criterion semantics. In addition, this reading is based on the abstract and does not include figures or full experimental details, so per-benchmark numbers, the exact suffix wording list, and pooling-probe implementation configurations still need to be confirmed from the original text.

Sources