Skip to main content
Back to timeline
arXivSource publication:

A 300M-parameter model pretrained only on a synthetic prior cuts bits per byte on six Wikipedia languages from 8 to about 1

Related research and updates

Synopsis

The authors present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on synthetic non-linguistic sequences generated by recurrent structural causal models, which with frozen weights learns in context to predict real text: on English, Chinese, Hindi, Arabic, Japanese, and Korean Wikipedia, bits per byte falls from the uniform eight to about one at one million bytes of context, and the same model learns to count, compare magnitudes, and add approximately on numerals, predicts deterministic sequences such as Rudin–Shapiro, Kolakoski, and the prime indicator, and compresses six non-text domains from source code to speech below gzip and PPMd.

AI-generated editorial illustration: Learning to Learn a Language

Interpretation

PFLM predicts natural language from a context prefix alone, having never seen a word of any real language, with bits per byte falling from eight to about one as context grows. Every language model to date, however synthetic its curriculum, has been trained on natural language; the authors state that to the best of their knowledge no prior work has shown a language model learning to predict natural language in context after pretraining only on samples from a constructed non-linguistic prior. Bits per byte is measured on six Wikipedia languages as context grows, with weights frozen throughout and adaptation occurring only through context.

Samples from the prior reproduce the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence at multiple scales. The prior is constructed rather than fitted, and its samples occupy the same statistical neighborhood as natural text without sharing any specific script; an ablation that fixes every node to update every step leaves scalar fingerprints unchanged but reduces the prior's long-range mutual information, integrated over distance, by 18%, indicating the multi-rate firing schedule carries dependencies across many steps. Comparison against four data-free comparators (uniform, Zipf, random bigram, random hidden Markov model) and six Wikipedia corpora of 500 sequences of 4096 bytes each, plus a fixed-update-rate ablation for the long-range mutual-information analysis.

The same model learns to count through carries, compare magnitudes, and add approximately on numerals, and predicts deterministic sequences such as Rudin–Shapiro, Kolakoski, and the prime indicator. These abilities appear in a model pretrained only on the synthetic prior and never trained on numerals or these sequences, and comparison and addition accuracy falls with the ratio between quantities, the distance effect seen in people comparing digits. Counting is scored by exact match of the greedy decode and grouped by digit count and carry condition; comparison and approximate addition are few-shot with queries binned by ratio; deterministic sequences are reported as bits per symbol over sequence length.

On six further domains—source code, molecules, proteins, DNA, symbolic music, and audio—PFLM achieves lower bits per byte than bigram, gzip, and PPMd. The same weights compress across domains without per-domain training, with the smallest margin over the best compressor on DNA and the largest on audio and music. Direct bits-per-byte comparison against three general-purpose compressors on the same domain data.

Perspective

The result applies to a frozen-weight, context-only setting: the model never takes a gradient step during evaluation and adapts to a corpus only through its context. Context length is the current limit, with a trained context of one million bytes (about two hundred thousand English words), and the authors note that at that length prediction on Wikipedia still falls short of what modern language models reach after training on trillions of tokens. They expect two lines of future work to be most rewarding: applying the prior to domains with little data, such as low-resource languages, undeciphered scripts, or animal communication, and joint pretraining on real text and prior samples to enhance the in-context learning of conventional language models.

Which components of the prior matter most for transfer to natural language remains an open question, with the fixed-update-rate ablation the only clue given in the text. Evaluation covers six languages and six non-text domains, but context length stops at one million bytes, so whether the curves keep falling at longer contexts and whether the gap to conventional language models narrows remains to be seen. The accuracy of the number tasks falls with the ratio between quantities, and the correspondence to human approximate number sense is analogical; the text gives no direct cross-species or cross-task comparison. In addition, this material is the full text, but some figure values appear in the prose as ranges, so exact numbers require the original figures and appendices.

Sources