BlindBias jailbreaks black-box endpoints from sampled text alone, taking the top mean score in 20 of 24 comparisons
Synopsis
The work introduces BlindBias, a framework for decoding-time jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation: combining sample-based distribution reconstruction, risk-gated residual control, and speculative multi-token execution, it achieves the highest mean score in 20 of 24 comparisons across four target endpoints and three benchmarks, and shows that selective intervention trades attack effectiveness for less frequent controller invocation.
Interpretation
BlindBias moves decoding-time residual control from settings that require weights or numerical probabilities to interfaces that return only sampled text, achieving the highest mean score in 20 of 24 comparisons across four target endpoints (GLM-5, Gemini-3.5-Flash, Qwen3-32B, Kimi-K2.5) and three benchmarks (AdvBench, HarmBench, SORRY-Bench). Prior decoding-time attacks rely on target weights or numerical token probabilities, while black-box jailbreaks have mainly operated at the prompt level through prompt search, template evolution, multi-turn interaction, or encoded instructions; this work adapts the BiasNet residual controller to sampled-text interfaces while retaining fine-grained control. The main table reports per-cell scores for four targets, three benchmarks, and two metrics (Harm and Info) against four official-implementation baselines: PAIR, GPTFuzz, LogiBreak, and FlipAttack; the advantage is not universal, as FlipAttack leads both metrics on GLM-5 SORRY-Bench.
Sample-based reconstruction combines sparse counts from finite sampling with a prior to recover a usable control signal: the uniform prior lowers PPL from 11.85 to 7.07 and raises Harm/Info from 1.50/1.26 to 3.03/2.16. The work shows that under sample-only access, reconstruction accuracy and steering utility are related but not interchangeable: a global unigram prior fits predictions better yet yields lower attack scores than the uniform prior, while numerical log probabilities remain the strongest reference (PPL 3.17, Harm 4.06). The ablation runs on a 40-record training cache and 100 held-out AdvBench prompts with Qwen3-32B as the target, reporting PPL, unseen-event NLL, and downstream attack scores; the numerical-probability reference is explicitly marked as outside the sample-only setting.
Risk gating concentrates reconstruction and intervention at a small subset of positions: soft gating reduces average API calls from 4,000 to 283 (92.9%) versus ungated control while Harm falls from 3.96 to 3.58, and versus hard gating it lowers active positions from 9.5% to 5.25% with 36.8% fewer API calls. The work uses a risk score that evolves with the response prefix to decide when to reconstruct and modify the distribution, rather than intervening only in a fixed opening window, so control can reactivate later in generation. The gating comparison on Gemini-3.5-Flash reports Harm, Info, API calls, and active-position frequency; the text also notes that active-position frequency alone does not establish endpoint query savings, which also depend on sampling, retries, and response length.
Speculative multi-token execution and contextual-prior transfer form two extensions: the former amortizes target calls over spans where the gate suppresses the residual via draft-and-verify, while the latter shows a same-family proxy prior performs best on all three metrics and cross-family proxies still improve Harm over a shuffled control. The speculative path borrows the draft-and-verify structure of speculative decoding but has a local risk model verify whether each prefix permits bypassing the controller; prior transfer tests whether a proxy prior can supply a control signal when the target vocabulary is unavailable. The prior-transfer experiment fixes the Qwen3-32B target and controller and compares Qwen3-1.7B, SmolLM2-1.7B, Gemma-3-1B, and a shuffled prior; the text states that verification enforces the gate rule on accepted prefixes and does not imply token-for-token equivalence with repeated single-token calls.
Perspective
The framework targets text-only continuation interfaces that permit repeated stochastic sampling and assistant-prefix continuation, where the attacker cannot access target weights, hidden states, or numerical token probabilities; the main setting uses a local tokenizer to define the action vocabulary and trains a separate controller for each target. Within this setting, the results apply to the four evaluated endpoints and three English harmful-request benchmarks, and can serve as a starting point for assessing decoding-time risk on such interfaces, for reducing endpoint-level query cost, and for extending reliable control beyond tokenizer-based action spaces.
Reconstruction quality and steering utility do not align: a global unigram prior fits predictions better yet scores lower on attack than the uniform prior, and numerical log probabilities remain a stronger reference, so the ceiling of the sampled signal is still to be characterized. Active-position frequency is not the same as endpoint-level query savings, which also depend on sampling, retries, and response length. Verification in the speculative path enforces the gate rule on accepted prefixes only, and the text explicitly does not claim token-for-token equivalence with repeated single-token calls. The vocabulary-free string action space cannot emit target token strings absent from calibration and retains only the first proxy token, making it an exploratory stress test. In addition, training uses reference prefixes while inference uses generated prefixes, a difference in prefix distributions that the text flags as not removed.
