GlitchPatch repairs glitch tokens in frozen language models via local retokenization, reaching an 85.10% mean fix rate and 0.00% regression across ten models
Related research and updatesSynopsis
The work proposes GlitchPatch, an external framework that repairs glitch tokens without accessing model internals by optimizing input tokenization: offline, Behavioral Path Optimization (BPO) searches for the behaviorally optimal equivalent token sequence for each glitch token and compiles validated replacements into a rule table; online, only the IDs of matched glitch tokens in the canonical token sequence are substituted. Across ten models spanning six tokenizer families, it achieves an 85.10% mean fix rate, 14.37 percentage points above the strongest baseline, reduces the mean glitch rate from 14.88% to 2.27%, attains 0.00% regression in full-vocabulary evaluation, and keeps the mean online latency increase below 0.1%.
Figure 2: BPE tokenization of “Hello!” with Qwen3.8-27B.
arXivInterpretation
It presents the first exploration of repairing glitch tokens from outside the model, and an empirical study of BPE merge-rule deletion reveals two findings: deleting a glitch token's merge rule fixes a portion of failures but disrupts normal longer tokens sharing the same intermediate merge node, raising the overall glitch rate; and different decomposition granularities yield non-monotonic fix rates across models while collateral damage to normal tokens grows monotonically with depth. Prior repair methods such as GlitchProber, GlitchEdit, and GlitchCleaner must operate inside the model (activation hooks, embedding weight edits, or a gated LoRA branch); this work reframes repair as an input-tokenization optimization and shows that global rule deletion cannot jointly achieve repair gains and protection of normal tokens. On Llama-3.1-8B-Instruct and Qwen3.8-27B, every vocabulary entry is tested on the three GlitchQuiz tasks (repetition, spelling, length); merge rules are deleted globally at three backtracking depths, all entries are re-encoded, and detection is re-run, reporting fix rate, regression rate, and glitch rate by depth.
It proposes GlitchPatch: offline, BPO ranks candidates on an equivalent-path graph with a proxy score, fully evaluates the top paths, and confirms rules through independent calibration; online, only a single-pass rule match and ID substitution are performed, with no change to model parameters or internal states, and unmatched inputs preserved by construction. Repair becomes a reversible input-level patch whose rule table contains only glitch-token IDs and their replacement sequences, so normal-token encodings remain completely unchanged; a proxy score plus full-path verification controls cost when candidate paths are numerous and exhaustive evaluation is impractical. The paper formalizes the decoding-equivalence constraint, candidate graph construction, utility function, and thresholds, and ablates each component: restricting search to BPE merge history lowers the fix rate to 73.82%, removing the pair term to 81.73%, removing full-path evaluation to 85.34%, while lowering the utility threshold raises it to 90.23% at the cost of weaker behavioral stability.
Across ten models and six tokenizer families, GlitchPatch reaches an 85.10% mean fix rate, higher than every baseline on every model, reduces the mean glitch rate from 14.88% to 2.27%, and shows 0.00% regression in full-vocabulary evaluation; practically, mean offline build is 0.89 hours, online latency increases by about 0.065%, token-level loss on natural texts containing glitch tokens drops by a mean 13.47%, and the General Score across four benchmarks changes by a mean 0.008 percentage points. Against baselines that require model-internal access (GlitchCleaner 70.73%, GlitchEdit 58.39%, GlitchProber 24.63%), the method attains a higher fix rate under the stricter input-only restriction, with shorter build time and no extra parameters or extra model calls. Full-vocabulary per-token detection removes selection bias; in contextual regression, short-answer accuracy on adjacent normal-token content changes by at most 0.2 percentage points; on out-of-distribution downstream tasks (entity QA, verbatim copy, code-identifier completion) over a uniformly sampled set of 500 glitch tokens per model, the overall mean accuracy is 70.0% versus 59.5% for GlitchCleaner, 53.5% for GlitchEdit, and 48.3% for GlitchProber.
It characterizes repairability boundaries: of Qwen3.8-27B's 29,451 glitch tokens, 5.00% are structurally unrepairable because no budget-feasible equivalent path exists in the vocabulary, and available path count correlates positively with fix rate (85.4% for 1-3 paths, 93.8% for 4-10, 98.7% for more than 50), with a final fix rate of 92.21% on the subset that has feasible paths and does not fail all selection probes initially. It attributes fix-rate differences to vocabulary structure and path availability rather than reporting only aggregate numbers, giving an actionable basis for judging which glitch tokens are worth repairing. All glitch tokens are grouped by available path count with reported fix rates, and retention and failure cases are enumerated with their proportions, including no feasible path, insufficient utility, calibration failure, and final-evaluation failure.
Perspective
The result targets providers and engineering teams that deploy frozen checkpoints and want a fast, reversible repair of glitch tokens without changing weights or the inference implementation, for models whose tokenizers satisfy the adapter contract (byte-level BPE or SentencePiece, including normalization, leading-space, byte-fallback, and special-token handling). The method identifies glitch tokens through per-token behavioral tests (repetition, spelling, length) and is constrained by decoding equivalence, so its repair scope is limited to glitch tokens for which a budget-feasible equivalent path exists in the vocabulary; the rule table is built per model and tokenizer, offline build takes 0.36 to 2.08 hours, and online adds only rule matching and sequence reconstruction.
Open questions remain: glitch identification relies on three behavioral tasks (repetition, spelling, length), so whether other forms of anomalous behavior are covered awaits testing; fix rates vary across vocabulary sizes and tokenizer families (80.50% to 91.04%), and about 5% of tokens are structurally unrepairable and need other routes; the rule table and hyperparameters were selected on a Qwen3.8-27B development set and then fixed for the remaining models, so robustness of cross-model transfer deserves continued tracking; language-understanding and contextual evaluations use GPT-5.6-Terra to generate target-containing passages, and how that generation distribution affects conclusions is an open question; entity QA relevance is judged by a model, and although 94% judge-human agreement is reported on a 50-sample subset, the scale of human verification remains limited.
