Public articles linked to the same research event.
arXiv Using strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B), this work compares post-user suffixes for activation probes and finds that a classification suffix consistently improves out-of-distribution detection on a single-position probe (up to about 4 AUC points), that the gain comes from the classification format rather than the named criterion, that a content-free suffix matches the real malicious/benign one, and that the benefit carries to production multi-position pooling probes (attention, multi-max, MLP) though the best suffix there is readout-dependent.
Using strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B), this work compares post-user suffixes for activation probes and finds that a classification suffix consistently improves out-of-distribution detection on a single-position probe (up to about 4 AUC points), that the gain comes from the classification format rather than the named criterion, that a content-free suffix matches the real malicious/benign one, and that the benefit carries to production multi-position pooling probes (attention, multi-max, MLP) though the best suffix there is readout-dependent.
Using strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B), this work compares post-user suffixes for activation probes and finds that a classification suffix consistently improves out-of-distribution detection on a single-position probe (up to about 4 AUC points), that the gain comes from the classification format rather than the named criterion, that a content-free suffix matches the real malicious/benign one, and that the benefit carries to production multi-position pooling probes (attention, multi-max, MLP) though the best suffix there is readout-dependent.
Using strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B), this work compares post-user suffixes for activation probes and finds that a classification suffix consistently improves out-of-distribution detection on a single-position probe (up to about 4 AUC points), that the gain comes from the classification format rather than the named criterion, that a content-free suffix matches the real malicious/benign one, and that the benefit carries to production multi-position pooling probes (attention, multi-max, MLP) though the best suffix there is readout-dependent.