Same bytes, two kinds of authority: splitting forged chat-template markers into ordinary subwords cuts injection success by 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B
Synopsis
Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.
Interpretation
Under a byte-identical contrast that differs only in whether forged markers keep their reserved ids, the reserved representation carries most of the attack: the identity gap is 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B and 8.1 percentage points on Qwen3-8B direct harm, significant in every run of these seven configurations. Prior template-level attacks such as ChatInject vary the template's visible form, moving text and token ids together, while tokenizer-level work re-segments strings that never carried a reserved id. The Matched control separates the number of extra tokens from whether the reserved ids survive, giving a direct fixed-byte measurement of what the reserved representation is worth. 400 cases per attack type in InjecAgent, four open-weight families plus Qwen3-32B as a scale check, eight configurations, three to five runs each, exact paired McNemar tests requiring significance in every run; Matched stays within 1.7 percentage points of Reserved on every configuration, so the extra tokens alone cost the attacker almost nothing.
The authority sits in the single vector at the marker position rather than in a particular id: on Llama-3.1 replacing the marker-position input vector with that of the nearest ordinary token restores the attack in full (98.4 versus 98.2), while the mean of the marker's subword vectors reaches only 58.4; substituting another reserved control token's vector keeps the attack at or near full strength on both models. Earlier work showed that special-token injection works or that re-segmentation shifts behaviour, but did not localise which vector carries the effect; the fixed-position, fixed-length vector swap separates the id from the vector. Input-embedding swaps on all 510 direct-harm cases for Qwen3-8B and Llama-3.1, keeping the token sequence, marker positions and every other row fixed; the accompanying embedding-proximity experiment shows restoring the reserved id is worth 18 to 47 percentage points while moving a split closer in embedding space is worth at most 17 and nothing on Llama-3.1.
Instruction tuning strengthens the preference for reserved markers: in all three base and instruction-tuned pairs tested (Qwen3-1.7B, Qwen3-8B, Seed-OSS-36B), the instruction-tuned checkpoint's logit identity gap leans further toward the reserved marker, while the cost of the extra tokens does not shift. Prior work attributed template attacks to template structure or stylistic cues without tying the preference to a training stage; this study reads a logit difference from a single forward pass on an identical prompt and forced continuation, avoiding a generative readout that is not comparable across base checkpoints. Each pair is read over the same 200 cases for both checkpoints with paired shifts tested by the Wilcoxon signed-rank test; no base checkpoint prefers the reserved marker, with Qwen3-1.7B-Base indifferent and the Qwen3-8B and Seed-OSS-36B bases disfavouring it.
The existing mitigation misses the tool channel: among the 400 most-downloaded chat models on Hugging Face, 33 of 67 distinct tokenizer configurations, covering 255 checkpoints, declare tool-protocol tokens such as <tool_response> as special:false, so split_special_tokens leaves them intact, and a forged block built only from those tokens yields identity gaps of 9.4 to 19.9 percentage points on every model that declares them. Tokenizer-side intervention had been treated as a general switch over special tokens without auditing which channels it actually covers; this work provides a configuration census and measures the effect on the uncovered channel separately. Each of the 400 checkpoints is loaded and its added tokens encoded with and without the flag, then classified by protocol role and checked for reachability; the tool-channel result holds on Qwen3-8B, Qwen3-32B, GLM-4.5 and Seed-OSS-36B and is larger than the covered system channel on Qwen3-8B.
Perspective
The result is aimed at deployers of self-hosted open-weight models who control the serving stack and the tokenizer call and can label which spans of a prompt are untrusted. In that setting the fixed-byte contrast is a measurement tool: it estimates how much of the same injection the encoding choice blocks, and it applies unchanged to models trained to resist injection, where it can test whether a defence removes this authority or only the surface cues that trigger it. The authors also provide a training-free, reversible source-aware encoder that removes the special-token matching step on untrusted spans while keeping the pre-tokeniser and BPE merges, and that changes no token on benign tool traffic.
The authors scope the work to self-hosted open-weight models, leaving hosted APIs that accept only strings outside it; InjecAgent's success criterion is the tool call parsed from the next turn, with execution-level success measured on AgentDojo; and the vector swaps of Section 4.4 intervene on the model's input rather than on bytes an attacker can send. On Qwen3-8B the text of the markers dominates rather than the reserved representation, and the data-stealing gap is not significant there, which the authors read as a prior null result reflecting a readout at its ceiling rather than an absent effect. An adaptive attacker who abandons the reserved marker for ordinary tokens near it in embedding space recovers most of what the defence removes, so removing the reserved id removes the default attack's advantage but not the authority itself. In addition, some table values in the loaded text appear as blanks; no specific number is inferred here beyond those stated in the prose.
