Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

SENTINEL cuts jailbreak success to about 5% via input-output intention matching while lowering over-refusal

The authors propose SENTINEL, a fine-tuning-free, generation-time jailbreak defense that reframes mitigation as intent extraction: it matches semantically aligned input-output context windows to extract intention-revealing subsequences, scores them via refusal-direction projections, and halts generation when needed; on HarmBench across multiple LLMs it reduces attack success rates to close to 5% (near 3% for strongly aligned models such as Llama-3-8B) while keeping low over-refusal on OR-Bench, remains robust to white-box adaptive attacks, and mechanistically re-distributes jailbreak features from alignment blind spots to aligned regions.