Wild AI Text Becomes an Independent Data Axis
Highlights
The clearest signal this issue is that "wild AI text" has been established as an independent data axis. How Much Is an AI Token Worth for the first time separates naturally occurring AI content on the web, written for human readers rather than for training, from synthetic data and model collapse settings. Using 800 models (19.9M–973M) with controlled changes in AI/human token ratios, it provides scaling laws that can flip the sign of AI token value: separating saturating benefits from logarithmic harms, and degrading to Chinchilla when no AI text is present (How Much Is an AI Token Worth?). The core conclusion is that AI tokens only benefit data-hungry models; at the Chinchilla-optimal 20 TPP, the benefit vanishes and turns into harm, and existing repetition and mixing laws fail to predict this behavior. It directly addresses the field's concerns about "synthetic data benefit boundaries" and "repetition vs. new data under data constraints," and provides actionable prescriptions: filter AI text, prioritize repeating human text over expanding AI corpora, and report human/AI validation losses separately.
This issue sees a "unified model" style advance on the document extraction quality axis. ReScraper compresses the entire pipeline of "HTML scraping + rule-based cleaning" into a single 0.6B model, using four operation types to unify the conversion from raw HTML to pretraining text. Across three scales (400M/1.4B/2.8B) and 22 DCLM Core tasks, it demonstrates that an end-to-end unified model outperforms any "scraper+cleaner" cascade (including the multi-agent DataOrchestra) (ReScraper). Its result—repeating 7.5 times at 2.8B scale still beats the baseline—grounds the idea that "extraction quality can offset repetition costs" in data-constrained allocation decisions.
This issue provides strong negative evidence on mid-training data selection. Everything in Moderation tests, at 8B real scale, with 30 mixing ratios and a full SFT+RL pipeline, whether "coverage gaps can be repaired by post-training." The conclusion is that compensatory SFT improves absolute accuracy but cannot close any inter-domain gap; permutation tests show gains are systematically placed in a way that preserves gaps (Everything in Moderation). This contrasts with Issue 6's Stress-testing Alignment Midtraining, where conflicting data priors can be overridden by small downstream data: coverage gaps are irreversible, while conflicting data priors are coverable—making mid-training allocation decisions more irreversible than previously thought.
Community and Updates
A negative replication on classifier-based quality filtering (CQF) is worth noting: simply rearranging documents in a Wikipedia style can flip the filtering decision for about 7% of documents in FineWeb-Edu's CQF model, and this phenomenon consistently appears across 26 domains. NemoCurator-family models, at threshold 3, allow over 7% of rearranged low-score content through (Is a Document Educational or Just Wikipedia-Style?). This closes a loop with another finding from this issue's Wild AI paper: existing quality filters actually prefer AI text (the FineWeb pipeline retains AI documents 2.3 times more often than human documents; DCLM, 9.8 times). The fragility of CQF combined with its preference for AI text implies that "quality filtering" on current web corpora may systematically amplify AI contamination.
Open Questions
- To what scale does the Wild AI paper's sign-flipping law for AI token value hold? Under its 973M upper bound and base loss criterion, can the "filter AI text" prescription be reproduced at ≥1B scale with post-training upper bounds as the criterion, especially regarding AI text's impact on RL trainability?
- Over what range does ReScraper's "quality gains can offset repetition costs" hold? Is there a quantifiable substitution relationship between extraction quality improvements and repetition rounds that could be incorporated into data-constrained scaling laws?
- Does Everything in Moderation's "coverage gap irreversibility" depend on its single logical reasoning setting? If switched to knowledge-based or code-based mid-training domains, does the irreversibility of compensatory SFT still hold?
- Does the fragility of classifier-based quality filtering (where Wikipedia-style rearrangement flips about 7% of decisions) imply that CQF scores are unsuitable as continuous signals for allocation decisions? If CQF is format-sensitive, what more robust signals should quality filtering turn to?
Papers in this issue
Wild AI text now makes up over 31% of filtered web data, harming Chinchilla-optimal models while a new scaling law predicts its effects 41% better than prior work.
Editor's noteFor the first time, "wild AI text" is separated from synthetic data and model collapse settings, treated as a naturally occurring, unlabeled independent data source in pretraining corpora, with scaling laws that can flip the sign of AI token value: separating saturating benefits from logarithmic harms, and degrading to Chinchilla when no AI text is present. Controlled ablations across 800 models (19.9M–973M) show AI tokens only benefit data-hungry models; at the Chinchilla-optimal 20 TPP, the benefit vanishes and turns into harm, while existing repetition/mixing laws fail to predict this behavior. Relative to Issue 1's *Internal Data Repetition* and Issue 3's *Scaling Laws for Mixture Pretraining*, the increment is treating AI text as an independent data source and providing actionable prescriptions (filter AI text, prioritize repeating human text over expanding AI corpora, report human/AI validation losses separately). Publicly releases the Wild AI corpus (83B tokens), 800 models, and code. Limitations: max scale 973M, base loss as criterion.
ReScraper, a single 0.6B language model, replaces the entire heuristic scraping and cleaning pipeline, improving LLM pretraining data quality by 3.8–4.7% over all baselines.
Editor's noteCompresses the entire "HTML scraping + rule-based cleaning" pipeline into a single 0.6B model, using four operation types (extract/keep/edit/delete/rewrite) to unify the conversion from raw HTML to pretraining text, and demonstrates that an end-to-end unified model outperforms any "scraper+cleaner" cascade (including the multi-agent DataOrchestra). Controlled ablations at 400M/1.4B/2.8B scales and 22 DCLM Core tasks show each operation contributes gains individually, and at 2.8B scale, repeating 7.5 times still beats the baseline—quality gains can offset repetition costs, key evidence for data-constrained allocation decisions. Relative to Issue 6's WeVisDoc and Issue 1's EDGAR, the increment is jointly optimizing "extraction" and "cleaning" within a single model. Publicly releases data/models/code. Limitation: criterion remains base benchmark.
Mid-training domain data coverage has an inverted-U effect on downstream accuracy, and a fixed-budget SFT pass raises absolute accuracy but cannot close coverage-induced performance gaps.
Editor's noteFirst systematic test of the irreversibility of per-domain coverage in mid-training on subsequent SFT/RL: at 8B real scale, with 30 mixing ratios, 5 seeds, and a full SFT+RL pipeline, compensatory SFT improves 116/120 cells (mean +4.32pp) but cannot close any inter-domain gap (0/240 pairs closed at the 5pp threshold); permutation tests show gains are systematically placed to preserve gaps. Also provides per-domain internal optimal bands (10–40%) and saddle point evidence. Relative to Issue 6's *Stress-testing Alignment Midtraining* on conflicting data priors, the increment is explicitly studying whether "coverage gaps can be repaired by post-training." For allocation decision-makers, "compensatory SFT cannot close mid-training gaps" is a strong design constraint. Limitations: single logical reasoning setting, short RL leg.
Subdocument deduplication with a frequency- and length-aware copy-retention policy outperforms uniform keep-one and shard-sensitive suffix-array methods for LLM pretraining, achieving the best scores on FineWeb-Edu and code-heavy corpora.
Editor's noteAdvances sub-document deduplication retention strategies from fixed rules (keep-one/keep-k) to explicit frequency-length-aware retention functions, decoupling detection from retention, and derives an analytical form from implicit retention behavior of sharded local deduplication. Validated at 30B MoE real scale (294.9B tokens) against Keep-One controlled comparisons. Relative to Issue 1's *Internal Data Repetition* and Issue 3's *Scaling Laws for Mixture Pretraining*, the increment is treating the deduplication retention strategy itself as a tunable data allocation lever. Limitations: code and recipe not released, base benchmark as criterion.
Data mixtures provably accelerate scaling laws only when auxiliary data has heavier-tailed spectra and intermediate relative sample growth, with ridge regression achieving the optimal rate.
Editor's noteIn a high-dimensional regression setting, provides the first complete characterization of when data mixing can strictly improve scaling law exponents (not just prefactors): the necessary and sufficient condition for the mixed minimax rate to beat any single-dataset rate is that the auxiliary spectrum is heavier-tailed (δ>1) and auxiliary sample size grows in a specific interval, and proves ridge regression can achieve the mixed minimax rate. Relative to Issue 2's *Explaining Data Mixing Scaling Laws*, the increment is explicitly characterizing the condition interval for mixing to improve scaling exponents, with minimax and ridge matching analysis. Language model experiments (81.5M) qualitatively validate the γ₂ sweet spot. Limitations: theoretical setting is linear/ridge regression, LM experiments are small-scale, and post-training upper-bound criteria are not addressed.




