Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Appending a classification instruction after the user turn lifts out-of-distribution malicious-input detection by up to about 4 AUC points

Using strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B), this work compares post-user suffixes for activation probes and finds that a classification suffix consistently improves out-of-distribution detection on a single-position probe (up to about 4 AUC points), that the gain comes from the classification format rather than the named criterion, that a content-free suffix matches the real malicious/benign one, and that the benefit carries to production multi-position pooling probes (attention, multi-max, MLP) though the best suffix there is readout-dependent.