Replacing the trainable input embedding table with fixed 16-bit binary token codes yields 2.36 perplexity at 32 layers on roughly 16-17B tokens while removing 67.1M parameters
Synopsis
The study asks whether a decoder-only language model needs an independently trainable input vector per token: with V=65,536 and d_model=1024, it replaces the usual V×d_model input table with fixed 16-bit binary token codes plus a parameter-free tiled lift to model width, removing 67.1M trainable parameters (about 12.5% of the untied learned-input baseline), and across three training seeds 32-layer models trained on roughly 16-17B tokens obtain mean held-out perplexities of 2.44 for the learned-input baseline, 2.36 for canonical binary codes, and 2.39 for affine-recoded codes, with the authors stating these descriptive results do not establish statistical superiority or equivalence.
Interpretation
The paper proposes replacing the trainable input embedding table with fixed minimal binary token codes: for vocabulary size V, an injective fixed-length binary identifier requires K=⌈log2 V⌉ bits, lifted to model width by a parameter-free tiled scheme. Relative to the usual practice of learning an independent input vector per token, token identity on the input side becomes a fixed code, which in the V=65,536, d_model=1024 setting removes 67.1M trainable parameters, about 12.5% of the untied learned-input baseline. Evidence comes from the paper's stated configuration and parameter accounting plus controlled training across three seeds; the text does not give a step-by-step derivation of the parameter accounting.
Across three training seeds, 32-layer models trained on roughly 16-17B tokens obtain mean held-out perplexities of 2.44 for the learned-input baseline, 2.36 for canonical binary codes, and 2.39 for affine-recoded codes. This provides a descriptive comparison between fixed binary codes and a trainable input table at matched scale and training volume, and additionally covers a table-free implementation using one fixed invertible affine recoding over F_2^16. Evidence is the mean perplexity over three seeds; the authors explicitly state these descriptive results do not establish statistical superiority or equivalence, so they should be read as observations in this setting rather than confirmation.
Standardized evaluation of released base checkpoints with the LM Evaluation Harness adds commonsense, knowledge, and language-modeling benchmarks; the three paper checkpoints show broadly similar, task-dependent performance, while external SmolLM2 reference models are substantially stronger on many tasks. This extends the effect of replacing the input table from perplexity to several downstream benchmark families, and uses external reference models to position performance, indicating the conclusion is limited to the studied regime. Evidence is the checkpoint comparison under a standardized evaluation harness; the text does not list per-benchmark scores, the task list, or sample sizes.
The paper's conclusion is limited to the studied regime: learning nontrivial language modeling does not require a free trainable token-indexed input table; the Transformer still learns continuous representations, and the output vocabulary projection remains standard and trainable. This clarifies the boundary of the change—what is removed is the input-side token-indexed table, not the model's ability to learn continuous representations or the output-side projection. The conclusion is drawn by the authors from the training and evaluation results above and is self-limited to the studied regime, with no claim of general equivalence or superiority.
Perspective
The work addresses readers studying input-side design in decoder-only language models, in the setting of V=65,536, d_model=1024, 32 layers, roughly 16-17B training tokens, three training seeds, and standardized evaluation with the LM Evaluation Harness. It enables follow-up work within this scope to examine the feasibility of fixed binary token codes with a parameter-free lift and to compare the canonical binary-code and fixed invertible affine recoding over F_2^16 variants. The authors explicitly limit the conclusion to the studied regime, so its significance is that a free trainable token-indexed input table is not required in this setting, rather than offering a general replacement across scales and vocabularies.
The authors state the descriptive results do not establish statistical superiority or equivalence, so the differences among 2.44, 2.36, and 2.39 across three seeds should be treated as observations rather than confirmation. The text does not list per-benchmark scores, the task list, or sample sizes, nor does it give a step-by-step derivation of the parameter accounting; these are open pieces of information a reader would consider when assessing transferability. External SmolLM2 reference models are substantially stronger on many tasks, and the comparability of their training setup with the paper checkpoints is not developed in the text. The conclusion is limited to V=65,536, d_model=1024, 32 layers, and roughly 16-17B tokens, so extension to other vocabulary sizes, model widths, or training volumes remains an open question.
