Skip to main content
Back to timeline
arXivSource publication:

SENTINEL cuts jailbreak success to about 5% via input-output intention matching while lowering over-refusal

Related research and updates

Synopsis

The authors propose SENTINEL, a fine-tuning-free, generation-time jailbreak defense that reframes mitigation as intent extraction: it matches semantically aligned input-output context windows to extract intention-revealing subsequences, scores them via refusal-direction projections, and halts generation when needed; on HarmBench across multiple LLMs it reduces attack success rates to close to 5% (near 3% for strongly aligned models such as Llama-3-8B) while keeping low over-refusal on OR-Bench, remains robust to white-box adaptive attacks, and mechanistically re-distributes jailbreak features from alignment blind spots to aligned regions.

Source-provided article image: Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching
Figure 1 ·

Figure 1: Intention Extraction Modeling. We solve an optimization problem over the extraction probabilities 𝐩 , 𝐪 \mathbf{p,q} to align semantically close context windows while maintaining intra-set semantic diversity. The optimization finds that context windows with features { f μ 2 , f μ 3 , f μ 4 } \{f_{\mu_{2}},f_{\mu_{3}},f_{\mu_{4}}\} are highly similar to { f ν 2 , f ν 3 , f ν 4 } \{f_{\nu_{2}},f_{\nu_{3}},f_{\nu_{4}}\} . We therefore extract an intention-related input–output context subsequence, e.g., “make a bomb from home (adv token)”, removing most adversarial tokens and making defense easier.

arXiv

Interpretation

Reframes jailbreak defense as an intent extraction task that exploits the semantic consistency between input and output in instruction-tuned models to localize harmful intent. Prior defenses largely rely on input perturbation or harmful-output suppression and rarely model where malicious intent resides; this work makes intent extraction the central defense step. Formalizes the problem with alignment and informativeness constraints and validates on HarmBench across multiple models; random scoring yields IPS near 0.5 while context matching often approaches 1.

Designs a plug-and-play defense module that requires no parameter or training-procedure changes and runs during autoregressive generation, halting generation when harmfulness is detected. Unlike latent-space defenses that require fine-tuning (e.g., Circuit Breaker, LAT) or a trained small model (IBProtector), it deploys without modifying model parameters. Evaluated on Llama2-7B, Llama3-8B, Vicuna-7b-v1.5, and Mistral-7B-v0.2; reported total invocation under 1.5 seconds even for the longest prompts (e.g., AutoDAN), with intent extraction empirically under 0.1 seconds.

Provides a mechanistic interpretation: extracted intention subsequences re-distribute jailbreak input features from alignment blind spots back to aligned regions, reactivating alignment behavior. Connects the effectiveness of an input-space defense to latent-space visualization evidence, linking input-space methods with latent alignment structure. Feature-distribution visualizations show jailbreak features agglomerating in a high-density island far from benign/harmful contours, while extracted subsequences return to aligned regions; jailbreak subsequences fall mostly in the harmful region and safety-boundary samples mostly in the benign region.

Remains resistant under a white-box adaptive attack where the adversary optimizes trainable embedding suffixes to both evade intention matching and preserve harmful behavior. Constructs attacks targeting the intention-extraction step itself, testing robustness under adversarial conditions. Under two adaptive settings ASR and SR stay low: compliance-only lowers IPS to 0.14-0.32 but outputs are often off-topic; compliance plus on-topicness raises IPS to 0.95-0.96 but exposes intent, with ASR around 3.00-6.67.

Perspective

The results target generation-time defense for pretrained checkpoints (llama2-7b, llama3-8b, vicuna-7b-v1.5, mistral-7b-v2) without additional fine-tuning, suited to deployers who need to lower jailbreak success while controlling over-refusal without changing model parameters. The method operates primarily in token space and depends on refusal-direction projections and a pre-computed threshold (set two standard deviations above the benign mean projection, so about 97.5% of benign prompts are unaffected), so its applicability presupposes an available refusal direction and benign calibration data for the target model. For engineering teams seeking interpretable, low-overhead defense and for researchers studying intent extraction and alignment mechanisms, this framework offers a modular path that plugs directly into autoregressive generation.

Several open questions remain for a careful reader: intent extraction operates in token space, which the authors note may be less expressive than continuous latent representations, suggesting future work could define intention matching directly in representation space; defense effectiveness depends on refusal directions and threshold calibration, whose stability across models or distribution shifts warrants further observation; in the adaptive attack, the compliance-only setting lowers IPS but often produces off-topic outputs, while compliance plus on-topicness preserves behavior but exposes intent, leaving a finer attack surface between the two as an open direction; additionally, some tables in the main text (such as the main results and ablation tables) do not provide complete numerical values in the text, so precise comparisons of baseline ASR and FPR should consult the original tables and appendices.

Sources