Skip to main content
Back to timeline
arXivSource publication:

Scaling synthetic pre-pretraining from 1B to 7B and 100B tokens keeps the token savings, but the authors trace them to long-range retrieval rather than a grammatical prior

Synopsis

The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.

AI-generated editorial illustration: Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

Interpretation

PPT's downstream gains persist under parameter and budget scaling: of 12 scale-mixture pairs, 9 gain at least 0.6 points on the downstream average, with a mean gain of 1.6 among them, and the gain is similar across scales (0.8 at 500M, 1.4 at 1B, 1.3 at 3B). Prior PPT work was tested only at or below 1B parameters, with pre-training budgets below 2B tokens on predominantly web text; this work extends to 500M-7B, up to 100B tokens, and four public data mixtures, reporting at least 21B pre-training tokens saved at the 3B scale (on Marin, PPT reaches at 63B tokens what PT-Only reaches at 84B). The comparison uses a PT-Only baseline from the same initialization plus a Control baseline running the identical 500 warm-up steps on held-out text; Control stays within 0.2-0.4 points of PT-Only at every scale, while -Shuffle Dyck outperforms Control in 10 of 12 pairs; the 3B Marin configuration is repeated with three random seeds and every downstream task category retains its sign across seeds.

There is no consistent evidence for the grammatical-prior explanation: PPT's downstream gains are not accompanied by stable gains in grammatical acceptability, with only 1 of 12 scale-mixture pairs showing a stable BLiMP gain against 10 of 12 for the downstream average, and no stable gain in the Morphology or Syntax subgroups. This diverges from prior work attributing PPT effectiveness to a grammatical prior (Hu et al. 2025; Mita et al. 2026), and the divergence appears already at the 500M scale covered by that work: the largest degradation occurs on C4, the very setting where prior work reported gains. BLiMP is reported overall and by semantics, morphology, and syntax subgroups; stability is judged by mean versus standard deviation across checkpoints in the second half of training (5K-10K steps), and the authors note that single-checkpoint estimates reverse sign in half of the 12 pairs with a 1.50-point average difference, motivating the multi-checkpoint average.

The gains come from long-range retrieval rather than formal grammar: tasks requiring retrieval of a specific earlier position work, while tasks that do not fail - -Shuffle Dyck (1.3), the non-formal NCA (1.2), and the formal MP-Struct Core (0.9) each improve 3 of 4 mixtures at 3B, whereas Set, which only asks whether a token has appeared before, degrades the downstream average by 7.2 points and improves none. The authors propose that PPT induces a long-range retrieval capability rather than the grammatical prior proposed by Hu et al. (2025); this converges with prior hints that dependency structure and reduced retrieval ambiguity matter, but diverges on what the model acquires from it. On verbatim retrieval, -Shuffle Dyck lowers NLL in all 12 scale-mixture pairs with 7 stable reductions, against 1 of 12 for BLiMP; the largest downstream gains are on LAMBADA, ReCoRD, and HellaSwag, all of which require information from the preceding passage.

PPT gains are insensitive to code and math share but depend on web text: raising the math share from 1.3% to 17.0% (lowering web share from 92.6% to 76.9%) yields an average gain of 2.0 points versus 1.9 for the full Marin mixture, while removing web text (leaving 82% code and 18% math) collapses the average gain to 0.2 points. Prior PPT was tested only on predominantly web-text corpora, leaving unexamined how it interacts with modern mixtures containing code and mathematics; this work shows the structural signal does not come from code and math but tracks the presence of web text. Five PT composition variants on Marin (full, 17% Math, DCLM only, FineWeb-Edu only, Marin without DCLM) show verbatim retrieval improving in all five with three stable, while BLiMP deltas remain small and none is stable.

Perspective

The result is aimed at decoder-only language model pre-training pipelines that are predominantly web text: at 500M-7B, up to 100B tokens, and four public data mixtures, PPT can be added as a low-cost warm-up, saving at least 21B pre-training tokens at the 3B scale, with the PPT stage accounting for about 4.74% of wall-clock time. It applies to teams seeking token efficiency without changing the main pre-training recipe, and to researchers designing synthetic warm-up tasks - which, on this account, should target long-range retrieval rather than natural language grammar.

OLMo3 is the only mixture that fails to benefit at more than one scale, and the authors leave a direct ablation of it and the responsible property to future work; 7B covers only Marin and -Shuffle Dyck and reaches roughly 11 tokens per parameter against 33 at 3B, so how much of the narrower 7B gain reflects capacity versus budget is unresolved; most configurations are single runs, and BLiMP differences often fall within checkpoint variance; why Set degrades performance so sharply, and how long-range retrieval transfers to specific downstream tasks, remain open questions.

Sources