Full text not available for this paper

Summary (Overview)

  • This paper investigates how wild AI-generated web text (text written by LLMs for human readers, not for training) affects language model pretraining, using 800 pretrained models ranging from 19.9M to 973M parameters.
  • The authors find that 27.5% of June 2026 web tokens passing FineWeb quality filters are AI-generated, rising to 31.1% by August 2026, and that existing quality filters (FineWeb, DCLM) actually favor AI text over human text.
  • AI text helps only data-starved models (below ~10 human tokens per parameter); for Chinchilla-optimal models (≥20 TPP), AI tokens raise loss almost immediately, while fresh human tokens continue to lower it.
  • The authors propose a new scaling law with separate saturating benefit and logarithmic harm terms that reduces to Chinchilla in the absence of AI text, achieving 41% lower error than the best existing law when predicting held-out models up to 3.6× larger.
  • Key recommendation: filter AI text when targeting human text; training on unfiltered web at 31.1% AI share requires 1.6× more compute than training on the human subset alone, rising to 3.0× by 2028 at forecast AI shares.

Introduction and Theoretical Foundation

Background and Motivation

  • Web text constitutes the majority of pretraining data and is increasingly AI-generated.
  • Models are trained far past the compute-optimal budget of ~20 tokens per parameter (e.g., Qwen3-32B uses ~1,100 tokens/parameter).
  • Projections indicate models will exhaust nearly all public human-written text by 2032 (Villalobos et al., 2024).
  • Every new web crawl forces a choice: keep AI documents, filter them, or find an optimal mixture.

Distinguishing Wild AI Text from Prior Work

The paper identifies a third category of AI text distinct from prior research:

CategoryDescriptionPrior Work
Model collapseRecursive training on a model's own outputShumailov et al. (2024); Gerstgrasser et al. (2024)
Synthetic dataCurated rephrasings designed to help specific domainsMaini et al. (2024); Kang et al. (2025)
Wild AI text (this paper)Written by many models for human readers, unlabeled, mixed with human textNo prior work

Theoretical Foundation: Chinchilla Scaling Law

The paper builds on the Chinchilla scaling law (Hoffmann et al., 2022):

L(N,D)=E+ANα+BDβ(1)L(N, D) = E + \frac{A}{N^\alpha} + \frac{B}{D^\beta} \tag{1}

where:

  • NN = model parameter count
  • DD = total training tokens
  • C≈6NDC \approx 6ND = training compute
  • E,A,B,α,βE, A, B, \alpha, \beta = fitted constants

Five Criteria for a New Scaling Law

The authors formulate five requirements:

  • C1: Token value must change with NN, TPPhTPP_h, and rr (AI/human token ratio)
  • C2: Must model both beneficial and harmful behavior
  • C3: Must have interpretable, separate benefit and harm terms
  • C4: First AI token's value must be finite, not assumed to be 1
  • C5: Must recover Chinchilla when no AI data is added

Methodology

Data Collection and Labeling

  • Corpus: Extended FineWeb (Penedo et al., 2024a) from July 2025 to June 2026, replicating the FineWeb filtering process on new Common Crawl data.
  • Dataset (WildAI): 96.04M documents, 83.31B tokens.
  • Two-stage labeling:
    1. EditLens Llama-3.2-3B labels >280B tokens
    2. Pangram 3.3.2 (false-positive rate 0.05%, false-negative rate 1.99%) on confident labels
  • Final splits: 58.91M human documents (42.19B tokens), 32.54M AI documents (35.32B tokens)
  • Topic/format labels: WebOrganizer (Wettig et al., 2025)

Experimental Setup

  • Models: 800 models, 19.9M to 973M parameters
  • Training budgets: 2.9 to 87.5 human tokens per parameter (TPPhTPP_h)
  • AI ratios: rr from 0 to 64 (AI tokens per human token)
  • Architecture: nanochat (Karpathy, 2025)
  • Evaluation sets:
    • Human text: C4 (north star), Paloma, pre-2022 FineWeb (FW22), FineWeb 2026 human split (FW26-H)
    • Mixed: FineWeb 2026 (FW26), containing 22.3% AI text
    • AI text: Cosmopedia (Cosmo), FineWeb 2026 AI split (FW26-AI)

The Proposed Scaling Law

The proposed law maintains the Chinchilla backbone with added benefit and harm terms:

L(N,DH,DA)=E+ANα+BDeffβ(1+H)(2)L(N, D_H, D_A) = E + \frac{A}{N^\alpha} + \frac{B}{D_{\text{eff}}^\beta}(1 + H) \tag{2}

Benefit term (saturating credit window, based on Qin et al. and Muennighoff et al.):

Deff=DH(1+ηg(r)),g(r)=R⋆(1−e−r/R⋆),R⋆=Ktρ(3)D_{\text{eff}} = D_H(1 + \eta g(r)), \quad g(r) = R^\star(1 - e^{-r/R^\star}), \quad R^\star = K t^\rho \tag{3}

where r=DA/DHr = D_A/D_H, t=DH/(20N)t = D_H/(20N), η\eta = free first-token parameter, KK and ρ\rho = fitted constants.

Harm term (logarithmically decelerating penalty):

H=γtunv[log⁡(1+r)−r1+r](4)H = \gamma t^u n^v \left[\log(1+r) - \frac{r}{1+r}\right] \tag{4}

where n=N/108n = N/10^8, γ\gamma, uu, vv = fitted constants.

The harm term has key properties:

  • Grows as r2/2r^2/2 for small AI additions
  • Grows as log⁡r−1\log r - 1 for large additions
  • Marginal harm peaks at r=1r = 1 (equal AI and human tokens)

Relative Scaling Law Evaluation

Following Held et al. (2025), the paper measures the change in loss against each run's human-only control:

Δ=log⁡L^(N,DH,DA)−log⁡L^(N,DH,0)\Delta = \log \hat{L}(N, D_H, D_A) - \log \hat{L}(N, D_H, 0)

Fitting: 726 models (19.9M–268M parameters) Held-out evaluation: 74 models (477M and 973M parameters, 1.8× and 3.6× the largest fitted size)


Empirical Validation / Results

Key Empirical Finding

"AI-generated text seems to almost always help [at low human budgets]. For Chinchilla-optimal models (20 or more TPP), AI text raises loss immediately, and more drastically at larger sizes."

At 5 TPP, AI decreases loss even at large AI budgets; at 20 TPP, AI additions increase loss with increasing AI budgets.

Scaling Law Benchmark Results

Table 1: Paired RMSE × 10³ on held-out sizes (lower is better)

LawkC4FW22FW26-HPaloma
Chinchilla (Hoffmann et al., 2022)54.324.655.409.15
Muennighoff et al. (2023)73.964.214.949.15
CD (Qin et al., 2026)83.683.634.179.15
Lovelace et al. (2026)93.553.744.148.38
Shukor joint (Shukor et al., 2025)131.411.502.257.80
Jain et al. (2024)78.008.038.049.34
Ours110.830.851.286.77

The proposed law achieves 0.83 × 10⁻³ on C4 versus 1.41 × 10⁻³ for the best existing law (Shukor et al., 2025) and 4.32 × 10⁻³ for Chinchilla.

Quality Filter Bias

  • FineWeb pipeline: keeps AI documents 2.3× more often than human (29.3% vs. 12.8%)
  • DCLM pipeline: keeps AI documents 9.8× more often (14.5% vs. 1.5%)

Filtering Results

  • Models above ~10–15 TPP (less at larger sizes) had lower C4 loss when AI text was filtered out, despite having less data.
  • The law gets the direction of filtering right in 18 of 19 pairs (paired error 3.28 × 10⁻³).
  • Chinchilla incorrectly predicts that removing AI text raises loss in all 13 pairs where it actually lowers it.

Repetition vs. AI Addition

At 20 TPP, repeating human tokens beats adding AI tokens:

  • After 1 added epoch: 2.3% lower loss with repetition
  • After 8 added epochs: 6.3–6.4% lower loss with repetition

Compute-Equivalent Gain (CEG) Forecasts

Using a random walk with drift forecast:

  • End of 2027: 42.3% AI tokens → 2.1× compute needed
  • End of 2028: 50.7% AI tokens → 3.0× compute needed
  • August 2026 (31.1% AI): 1.6× compute needed

Validation Set Masking

  • All 553 models with added AI text lower loss on AI-labeled FW26 text
  • 44% of models raise loss on human-labeled text
  • At 22.3% AI share in validation set: 95.5% of harmful runs appear as improvements
  • The sign flips at 5.1% AI share in the validation set

Theoretical and Practical Implications

Optimal AI Share by Target Text

Target TextOptimal AI Share at 5 TPPOptimal AI Share at 20 TPP
C4 (human)37%1.0%
Cosmopedia (AI)>90%>90%
FW26 (mixed)98%34%

Training on AI Text Changes Model Writing

  • AI-typical phrases per 1,000 words rise from 0.16 (no AI) to 0.36 (r=1) and 0.54 (r=4)
  • At r=2, Pangram labels twice as many stories as AI-generated (37.2% vs. 18.6%)

Practical Recommendations

  1. Filter AI text when the target is human text prediction
  2. Repeat human text before expanding with AI-generated web text
  3. Report validation loss on human and AI text separately — mixed validation sets hide harm
  4. AI text is valuable when the target is AI text (e.g., agentic workflows)

Conclusion

Main Takeaways

  • Wild AI text constitutes over a quarter of filtered web crawls and is rising.
  • AI text helps only data-starved models; it harms Chinchilla-optimal models.
  • The proposed scaling law (saturating benefit + logarithmic harm) reduces to Chinchilla without AI text and predicts models 3.6× larger with 41% lower error than the best existing law.
  • Training on unfiltered web at 31.1% AI share requires 1.6× compute; this rises to 3.0× by 2028.

Limitations and Future Work

  • Laws fit on models up to 268M, tested up to 973M; larger models may behave differently (more memorization, potentially more harm from AI data).
  • Only next-token loss measured; downstream task effects may differ.
  • English web text only, labeled by Pangram.
  • Future directions:
    1. Filter wild AI text (as MAI does) and add targeted synthetic data
    2. Identify domains where AI text helps
    3. Measure how existing AI text in pretraining corpora has affected current models
    4. Train models that understand AI-written input without writing like it (e.g., tagging AI text, masking its loss)

Released Resources

Related papers