Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

GlitchPatch repairs glitch tokens in frozen language models via local retokenization, reaching an 85.10% mean fix rate and 0.00% regression across ten models

The work proposes GlitchPatch, an external framework that repairs glitch tokens without accessing model internals by optimizing input tokenization: offline, Behavioral Path Optimization (BPO) searches for the behaviorally optimal equivalent token sequence for each glitch token and compiles validated replacements into a rule table; online, only the IDs of matched glitch tokens in the canonical token sequence are substituted. Across ten models spanning six tokenizer families, it achieves an 85.10% mean fix rate, 14.37 percentage points above the strongest baseline, reduces the mean glitch rate from 14.88% to 2.27%, attains 0.00% regression in full-vocabulary evaluation, and keeps the mean online latency increase below 0.1%.