acceptodds
Under review as a conference paper at ICLR 2027

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Abstract

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models (19.9M to 973M parameters), varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Existing scaling laws such as Chinchilla (Hoffman et al., 2022), which treat AI tokens as no different from human tokens, fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token (measured in human token equivalents) to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. It implies that training on unfiltered 2026 web text requires 1.3x as much compute as training on its human subset at 20 tokens per parameter, a gap that grows with the human budget. We therefore recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately, since at a 2026 crawl's AI share, mixed validation sets hide the harm in 95.5% of harmful runs. AI text remains valuable when the target is AI text. We will release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.