OPUR unifies jailbreak objectives via a harmfulness-reweighted distribution, lifting AgentHarm attack success on Llama-3.1-8B to 75.00% under stochastic decoding
Synopsis
The work reformulates LLM-agent jailbreaking probabilistically: it shows that the gradient of the log expected harmfulness with respect to the input equals the expected log-likelihood gradient under a harmfulness-reweighted output distribution, unifying expected-harmfulness maximization with target-likelihood optimization; building on this, OPUR constructs surrogate targets from the model's own stochastic rollouts to guide adversarial suffix optimization, achieving higher targeted attack success rates on AgentHarm, InjecAgent, and WebShop.
Figure 1: Overview of our probabilistic view of jailbreak optimization. Given an adversarial input, the model induces a response distribution pdis(y) = pθ(y | x, s), while the harmfulness function induces a preference pvic(y) ∝ψ(y). Considering either view alone may favor responses that are likely but insufficiently harmful, or harmful but unlikely. Their normalized product defines the reweighted distribution π, which emphasizes responses that are simultaneously likely under the model and harmful under the attack objective. Since exact sampling from π may be unavailable in practice, its high-probability regions can be represented by constructed surrogate targets and opti- mized through their likelihood. As optimization proceeds, model probability mass is progressively shifted toward harmful outputs, providing a unified interpretation of expected harmfulness maxi- mization and target-likelihood optimization.
arXiv · Page 2Interpretation
The paper establishes a gradient-level equivalence between the expected-harmfulness objective and the target-likelihood objective: the gradient of the log expected harmfulness equals the expected log-likelihood gradient under a harmfulness-reweighted response distribution. Previously, expected-harmfulness attacks (e.g., REINFORCE-GCG) and target-likelihood attacks (e.g., GCG, UDora) were treated as separate objectives; the identity shows they induce the same update direction at a fixed input, placing behavioral risk and likelihood optimization within one response-distribution framework. The result is stated as a theorem with a proof in Appendix A, under the assumption that differentiation and expectation can be interchanged; the paper also notes the reweighted distribution shares the same product form as existing probabilistic adversarial attacks.
It proposes OPUR, which approximates the harmfulness-reweighted distribution using the model's own stochastic rollouts, builds surrogate contexts via temperature-controlled probabilistic position sampling and probe-based screening, and aggregates cross-entropy gradients across rollouts to update the adversarial suffix. To address deterministic top-position selection overfitting a single rollout under stochastic decoding, OPUR probabilistically explores multiple non-overlapping position sets and retains the lowest-loss contexts, balancing candidate diversity with optimization reliability. The algorithm is described in four steps (stochastic rollout generation, probabilistic position sampling, context screening, suffix update) with position-scoring and surrogate-loss formulas; each rollout samples 4 injection-position sets and keeps 2 candidates, with code and configurations released.
Across three agent benchmarks, OPUR attains higher targeted attack success rates under stochastic decoding: on AgentHarm, Llama-3.1-8B-Instruct reaches 75.00% average ASR, above the strongest UDora baseline; on InjecAgent it reaches 49% (Llama) and 40% (Ministral); on WebShop it reaches 26.67% (Llama) and 36.67% (Ministral). Relative to GCG, REINFORCE-GCG, and the Sequential/Joint UDora variants, OPUR leads on most categories; on Ministral's AgentHarm, where performance is near saturation, OPUR matches the strongest UDora variants (99.43%). Results use ASR as the metric across malicious-instruction and malicious-environment scenarios on two open-source instruction-tuned models; ablations show that increasing the optimization rollout budget generally improves ASR, though gains are not strictly monotonic.
Perspective
The results target red-teaming of agents whose victim is an open-source instruction-tuned model: experiments run on AgentHarm (malicious instructions), InjecAgent and WebShop (malicious environments), all with ReAct prompting and public benchmarks, without targeting real users, accounts, or deployed services. For researchers wishing to reproduce or extend the framework, the paper releases code and experiment configurations and specifies the attack formulation, datasets, baselines, model settings, optimization parameters, and decoding configurations, supporting independent verification and defense comparisons on comparable benchmarks. The method applies to models where gradients and token probabilities are accessible, and its surrogate contexts are used only for optimization, not to modify the agent's actual generated responses.
Several ASR values appear as placeholders in the text and need to be confirmed against the tables; some tables (e.g., the AgentHarm and InjecAgent ablations) are not fully rendered in the text, so trends can only be inferred from descriptive statements. The theorems rely on assumptions such as interchangeability of differentiation and expectation and the harmfulness function not explicitly depending on the suffix, and how well these hold in practice deserves attention. Moreover, gains from larger rollout budgets are not strictly monotonic, and the optimal budget across benchmarks and models remains to be characterized; on models already near saturation, the method's separation from baselines is limited.
