Skip to main content
Back to timeline
arXivSource publication:

Matched evaluation finds DPO leads overall on safety control and text monitoring, while representation probes stay competitive at far lower marginal compute

Synopsis

Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.

AI-generated editorial illustration: When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

Interpretation

For safety control, DPO provides the strongest jailbreak robustness under matched supervision, Flow ranks second, and CAA and probe-based steering stay close to the base model; yet DPO degrades most after benign fine-tuning, with AIM attack success rising from 0.027 to 0.436. Prior steering work was mostly compared against other steering techniques, and behavioral alignment versus representation steering was rarely compared side by side under the same models, data, and protocol; this work runs a matched comparison on the same 5,406 PKU-SafeRLHF pairs across Qwen2.5-1.5B/14B and Llama-3.1-8B. Macro-average ASR tables cover three models, two attacks (AIM and refusal suppression), and pre/post benign fine-tuning; benign fine-tuning uses 20,000 Alpaca-Cleaned examples and steering vectors are applied without re-extraction.

Representation steering is competitive mainly in low-data, high-quality contrastive settings: Flow can match or beat DPO on small filtered contrastive subsets, while CAA and probe-based steering show little benefit from adding examples. Separating data quantity from data quality shows DPO improves with scale on both original preference data and filtered contrastive pairs, whereas steering gains come more from supervision quality than quantity. The data-scaling study on Qwen2.5-1.5B-Instruct spans 100 to 63,094 pairs, with the 100- and 200-pair settings averaged over three seeds and a reported mean within-setting ASR standard deviation of 0.0054.

For safety monitoring, specialized text monitors detect best overall (Qwen3Guard AUROC 0.996, AUPRC 0.995), while the strongest representation probe (mean pooling) reaches 0.982 AUROC and 0.978 AUPRC; in streaming, Qwen3Guard reaches 0.948 recall with a median first-alarm position of 0.034, whereas the rolling probe reaches 0.913 recall with the lowest sequence-level FPR of 0.017. Representation probes and text monitors had rarely been compared under a shared safety setting; this work reports full-response detection, streaming early detection, and computational cost on the same natively generated trajectories from Qwen2.5-32B-Instruct. Full-response detection uses 918 trajectories (384 harmful, 534 safe) and streaming uses 384 harmful and 200 safe trajectories, with thresholds calibrated on a prompt-level split to target 1% and 5% FPR and then fixed.

Probe-guided interventions recover much of the safety DPO loses after benign fine-tuning: at the 5% calibration threshold, blocking and corrective regeneration cut the post-update 14B DPO model's mean ASR from 0.500 to 0.042, with over-refusal rising only from 0.108 to 0.124 (blocking) and 0.120 (regeneration); text monitors triggering blocking also reduce ASR. Whether monitoring signals can improve downstream control was underexplored; this work couples monitoring and control and compares blocking, corrective regeneration, and adaptive Flow steering. Main-text results average AIM and refusal suppression on Qwen2.5-14B-Instruct, with appendices covering other model sizes, attacks, thresholds, and text-monitor baselines; the authors note some control experiments use a single training seed and jailbreak evaluation uses non-adaptive attacks.

Perspective

The results target safety engineering and deployment settings that must trade off control against monitoring: when high-quality safety data are limited, Flow-style representation steering is a practical option; when the generating model's internal activations can be reused, representation probes offer monitoring at low marginal cost; and when an aligned model loses safety after benign fine-tuning, post-generation blocking or corrective regeneration triggered by probes or text monitors can help. The evaluation scope is public models including Qwen2.5-1.5B/14B/32B and Llama-3.1-8B, public benchmarks including PKU-SafeRLHF, StrongREJECT, HarmBench, and XSTest, and two attack types, AIM and refusal suppression.

The authors note two scope limits: some control experiments use a single training seed, so results do not capture run-to-run variability, and jailbreak evaluation uses non-adaptive attacks, so robustness to attackers who tailor prompts to the deployed safeguard is not assessed. On monitoring, the calibration split contains only 51 safe trajectories, so achievable false-positive rates are discrete and realized test FPR may differ from its target; detection position is normalized relative to response length and does not establish whether an alarm precedes harmful content. Appendix replay experiments show probe advantages narrow under replay rather than native activation reuse and require an additional forward pass; the hidden-reasoning access experiment uses replayed activations, which the authors treat as evidence for the value of reasoning access rather than a general probe advantage. In addition, this evidence bundle is full text, but some table values are not fully rendered in the text, so a few specific numbers cannot be checked item by item here.

Sources