Skip to main content
Back to timeline
arXivSource publication:

PEV, which adds Gaussian noise to prompt embeddings, reaches a 100% jailbreak rate on JailbreakBench across six open-weight LLMs

Synopsis

The work introduces the Perturbed Embedding Vector (PEV) jailbreak: adding independent Gaussian noise to a prompt's embedding and resampling it repeatedly, with no gradient computation and no weight modification, achieves a 100% jailbreak rate on all 100 JailbreakBench prompts across six open-weight LLMs, with the first successful jailbreak typically arriving within one minute and at lower compute cost than the compared prior methods.

Source-provided article image: Jailbreaking Open-Weight LLMs via Random Embedding Perturbations
Figure 1 ·

Figure 1: Our Perturbed Embedding Vectors ( PEV ) attack. Gaussian noise z ~ i ∼ σ ⋅ 𝒩 ⁡ ( 0 , I ) \tilde{z}_{i}\sim\sigma\cdot\mathcal{N}(0,I) is added to each token embedding E ⁡ ( t i ) E(t_{i}) before being fed to the transformer block. At σ = 0.003 \sigma=0.003 the same “safe” model (GLM-4-9B) that refused the unperturbed harmful prompt now produces a response in another language, an unsafe completion, or degenerate output across runs.

arXiv

Interpretation

PEV reaches a 100% jailbreak rate on all 100 JailbreakBench prompts for all six open-weight models, while none of the compared baselines reaches a full rate; reported values include RefusalCones between 89 and 98 per model and means of 75.2 for -GCG, 81.0 for SoftPrompt, 61.0 for LatentFusion, and 91.3 for NeuroStrike. Prior jailbreaks typically require gradient computations, per-prompt optimization, or altering internal model weights; PEV reduces the attack to a single forward pass over a noise-injected embedding, repeated by resampling noise for the same prompt. Six models of different sizes (Qwen2.5-1.5B, SmolLM3-3B, Phi-4-mini, Mistral-7B, Llama-3.1-8B, GLM-4-9B), 100 prompts each with 20 runs per prompt, outputs classified by Claude Opus 4.6 using the JailbreakBench judge instructions, with manual verification of every response labeled unsafe.

PEV obtains the first successful jailbreak within one minute on every tested model without preprocessing, and the authors report the average compute cost to the first successful attack as up to an order of magnitude less than previous attacks. RefusalCones and NeuroStrike have per-prompt runtimes comparable to PEV after calibration, but their dominant cost is preprocessing, making them an order of magnitude slower overall; -GCG was excluded from the timing comparison because it takes much longer. All runs were performed on a single NVIDIA A100-SXM4-80GB via Google Colab, with wall-clock time recorded per batch of 20 runs per prompt and total time reported after completing 100 prompts times 20 runs.

Each model has a peak noise magnitude that maximizes the jailbreak rate: unsafe response rates rise to a peak and then decrease as the perturbation grows, and the ratio of the approximate noise norm to the average prompt embedding norm spans a wide range from 6% to 240%. This indicates that prompt semantics remain sufficiently preserved under perturbations that are large relative to the original embedding for the model to fulfill the harmful request, which the authors flag as deserving future investigation. The peak was identified by running PEV at various noise magnitudes on a small sample of unsafe prompts and measuring the percentage of unsafe responses; the chosen values and per-prompt run statistics, including maximum and mean runs, are reported.

PEV typically needs few resamples: most models jailbreak within 20 runs, with GLM-4-9B and Phi-4-mini cumulative curves converging rapidly to 100%, while Qwen2.5-1.5B holds out longer with a handful of stubborn prompts and a maximum of 160 runs. Because PEV is stochastic, repeated independent noise injections accumulate success on the same prompt, in contrast to the compared deterministic methods, for which the authors observe that more runs do not increase the jailbreak rate. The distribution of iterations to first unsafe response is recorded per prompt, with outputs separated into safe, unsafe, and degenerate classes; degenerate responses increase with perturbation while the safe response rate remains non-trivial.

Perspective

The result applies to open-weight instruction-tuned models where the embedding layer is accessible; the authors state the technique is currently restricted to open-weight models and suggest perturbation could be a cheap way to generate unsafe training data for closed-weight models to help design stronger guardrails and training data. The method requires only one-time write access to the embedding layer and then executes a forward pass in the original model, modifying neither weights nor hidden layers and leaving the user-visible prompt unchanged, so it transfers across prompts. The authors also frame perturbation as a tool for studying LLM dynamical behavior, suggesting that useful prompts may lie in a small, possibly measure-zero, region of all possible inputs and that perturbation explores beyond it, with possible connections to dynamical systems theory.

Several open questions remain: the peak noise magnitude must be scanned per model, and its relation to embedding dimension and prompt length is not given analytically; the authors report noise-to-embedding norm ratios spanning 6% to 240% and flag as future work why such large relative perturbations still preserve prompt semantics; degenerate responses increase with perturbation while the safe response rate stays non-trivial, which also invites deeper analysis; and jailbreak determination relies on Claude Opus 4.6 with JailbreakBench judge instructions plus manual verification, leaving open whether success rates agree under different judges.

Sources