Across 800 pretrained models, wild AI web text raises loss once data is plentiful, and a 31.1% AI share costs 1.6x the compute
Synopsis
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
Interpretation
The work quantifies the share and growth of "wild AI text" in real pretraining corpora: 27.5% of tokens passing FineWeb quality filters in June 2026 Common Crawl are labeled AI-generated by Pangram, rising to 31.1% in August 2026, against under 0.1% in June 2021, 10.1% in June 2024, and 16.1% in June 2025. Prior work on AI text studied either recursive training on a model's own output (model collapse) or curated synthetic rephrasings designed to help specific domains; this paper characterizes a third kind, text written by many models for human readers that arrives unlabeled and mixed into pretraining corpora at varying quality. Based on a monthly sample of 5,000 documents from January 2021 through August 2026 labeled by Pangram 3.3.2, and on the WildAI corpus of 96.04M documents and 83.31B tokens; the authors report Pangram 3.3.2's false-positive rate of 0.05% and false-negative rate of 1.99%, and measure 0.062% of 60,000 documents from 2021 crawls labeled AI as an upper bound on the false-positive rate.
Across 800 models (19.9M to 973M parameters), adding wild AI tokens to a fixed human corpus changes sign with the human data budget: for data-starved models (below about 10 human tokens per parameter) AI tokens lower loss on human text, but the benefit saturates and quickly reverses into harm; at the Chinchilla-optimal 20 tokens per parameter and above, AI tokens raise loss almost immediately while the same number of fresh human tokens keeps lowering it. Existing laws such as Chinchilla treat AI tokens as no different from human tokens and fail to predict this benefit-then-harm behavior that changes sign with the human budget; the authors propose a law with a saturating benefit term and a logarithmically growing harm term that reduces to Chinchilla when no AI text is added. All 800 models share one nanochat training recipe, human and AI documents are interleaved in a fixed hashed order, and every AI or fresh-human addition trains on its control's human documents in the same order and with the same seed; the paired error on held-out sizes (477M and 973M) is 0.83 on C4, against 1.41 for the next-best Shukor joint law and 4.32 for Chinchilla.
The work offers actionable pretraining-data recommendations: training on unfiltered web text at August 2026's 31.1% AI share requires 1.6x the compute to match its human subset at 20 tokens per parameter, rising to 3.0x by 2028 under the forecast share; repeating human text beats adding AI text; and validation loss should be reported separately on human and AI text. The authors further audit quality filters and find they favor AI text, with FineWeb's pipeline keeping AI documents 2.3x as often as human ones and DCLM 9.8x, and show that on a mixed validation set with a 22.3% AI share, 95.5% of harmful runs are reported as improvements. The filter audit uses a 10,000-document sample from 2026 Common Crawl; the repetition-versus-AI comparison uses six 19.9M and 35.8M models at 20 tokens per parameter, where repetition wins all 18 comparisons on three human-text sets, lowering loss by 1.7% to 3.8% while AI text raises it by 0.2% to 3.4%.
AI text remains valuable when the target is AI text: the optimal AI share on AI-generated text stays above 90% at every budget, and the fitted law values the first AI token at more than 20 human tokens; yet training on AI text changes how models write, with AI-typical phrases per 1,000 words at 268M and 20 tokens per parameter rising from 0.16 without AI text to 0.36 at an AI share of 1 and 0.54 at 4, and Pangram labeling 37.2% of stories AI against 18.6% for the control. This separates understanding AI-written input from writing like it, and shows mixed validation sets hide the harm because AI text is easier to predict: all 553 models with added AI text lower loss on AI-labeled text while 44% raise loss on human-labeled text. The style analysis generates continuations from twelve held-out 973M models for 500 WritingPrompts story prompts and WebText article openings and counts 26 AI-typical phrases; on CORE, AI text and fresh human text raise scores similarly (1.6 versus 1.9 points) while C4 loss moves in opposite directions (up 2.3% versus down 3.0%).
Perspective
The work targets English web pretraining where the goal is human text: it gives grounds for deciding how many AI tokens to add at a given human-token budget, when to filter AI text, and when to repeat human text, and it releases the WildAI corpus with AI, topic, and format labels, all 800 models, and code so that fits and figures can be regenerated on a CPU. For targets that are themselves AI text, such as agents reading other models' outputs, the direction reverses and the optimal AI share can exceed 90%.
The law is fit on models up to 268M parameters and tested up to 973M; larger models may memorize more of their training data or may instead learn AI and human text as related tasks that help each other, both of which the authors list as future work. The study covers only English web text labeled by Pangram and uses next-token loss as the main metric; at the authors' sizes AI text raises CORE scores about as much as fresh human text, but the authors note all models are too small to have meaningful performance. The AI-share forecast comes from extrapolating a random walk with drift, so the 50.7% estimate for the end of 2028 carries an interval, and the compute multipliers (1.6x, 3.0x) are implications of the law rather than measured large-model training runs.
