Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

A 300M-parameter model pretrained only on a synthetic prior cuts bits per byte on six Wikipedia languages from 8 to about 1

The authors present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on synthetic non-linguistic sequences generated by recurrent structural causal models, which with frozen weights learns in context to predict real text: on English, Chinese, Hindi, Arabic, Japanese, and Korean Wikipedia, bits per byte falls from the uniform eight to about one at one million bytes of context, and the same model learns to count, compare magnitudes, and add approximately on numerals, predicts deterministic sequences such as Rudin–Shapiro, Kolakoski, and the prime indicator, and compresses six non-text domains from source code to speech below gzip and PPMd.