What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining
Summary (Overview)
-
Unified scaling law for repetition cost: The paper demonstrates that excess loss from data repetition follows a single power law in the variable , where is epochs, is total parameters (including embeddings), and is unique tokens—no additional size-dependent terms are needed.
-
Critical epoch count: Repeated tokens fall to half the value of fresh ones at , which grows with training budget per parameter but is nearly independent of model size.
-
Resolution of conflicting size trends: The direction of size-dependent repetition effects depends entirely on which quantity is held fixed (U, U/N, or D/N), explaining seemingly contradictory findings in prior literature.
-
Compute-optimal allocation at fixed data: With unique data fixed, optimal runs grow model size and epochs together (each roughly as ) until loss stops improving near the critical epoch count .
-
Beyond token counts: At identical counts, repetition order, allocation across samples, source entropy, and tokenization all significantly affect loss—consecutive replay can raise loss by up to 0.46 bits per byte.
Introduction and Theoretical Foundation
Background
Compute-optimal scaling laws (Kaplan et al., 2020) prescribe growing training sets with model size, but the finite stock of public human-generated text (Villalobos et al., 2024) forces pretraining to increasingly repeat data. With unique tokens, a model of parameters trained for epochs sees training tokens.
Key Questions
- How many epochs should a run take?
- How should this number change with model size?
- Does anything besides the epoch count matter?
Prior Conflicting Evidence
- Muennighoff et al. (2023): Up to ~4 epochs nearly as valuable as fresh data
- Hernandez et al. (2022): Repeating subsets degrades models
- Xue et al. (2023): Degradation worsens with model size
- Lovelace et al. (2026): Penalty term in , prescribing fewer epochs at large compute
- Qin et al. (2026): Token effectiveness falls with tokens per parameter
Theoretical Framework
The paper defines two baselines for evaluating repeated tokens:
Data-matched baseline: Same and , extra compute from additional epochs. The value of a repeated token (fresh-token equivalence):
where is the fresh budget whose one-epoch loss equals that of the repeated run. means repeats are as good as fresh data; means worthless.
Compute-matched baseline: Same and , repeated data replace fresh data. The excess loss:
Methodology
Experimental Setup
- Models: Decoder-only Transformers (23M to 2B parameters), trained with Muon optimizer for hidden matrices and Adam elsewhere
- Data: SlimPajama (cleaned/deduplicated RedPajama), GPT-2 byte-level BPE tokenization
- Evaluation: Bits per byte (bpb) on held-out test sets
- Sources: Seven subsets (CommonCrawl, GitHub, Books, etc.), trained separately and in natural mixtures
Experimental Designs
CommonCrawl Grid: Model sizes 23M–361M, training budgets 5–300 TPP, epoch counts 1–16. Each size has a measured one-epoch curve providing both baselines.
Fixed-U sweeps: Eight sizes (14M–2B) on same 1B unique tokens
Fixed-U/N sweeps: Seven sizes with
Fixed-D/N sweeps: Five sizes at with partial repetition
Key Measurement Approach
Unlike prior work that inverts fitted scaling laws, the authors read directly from measured one-epoch curves via interpolation. This avoids two problems:
- Effective-data laws cannot represent runs that end above one-epoch loss (about 25% of repeated runs do)
- Inverting fitted laws amplifies small loss residuals when curves are flat
Empirical Validation / Results
1. Excess Loss Scaling Law
The central empirical finding across the entire grid:
with bpb and (). The exponent on an additional size term is negligible ().
2. Value of Repeated Tokens
- Second epoch: Worth 0.87–1.02 of fresh tokens at every budget
- Value decay: Falls fastest at small budgets per parameter (from ~0.9 to 0 at over 2–8 epochs)
- At large budgets (): Falls only from ~0.9 to 0.65
3. Critical Epoch Count
- About 4–5 epochs up to TPP of 20
- Rises to ~10 at TPP of 300
- Nearly independent of model size
4. Compute-Optimal Allocation at Fixed U
With fixed at , gives:
Fitted exponents: and .
Key results:
- TPP held at 16–24 (below fresh-data optimum of ~25)
- Loss minimum at 164M params/5.3 epochs (B), 297M/5.6 epochs (B)
- Beyond the minimum, more compute raises loss
5. Size Trends by Design
Table 1: Direction of size trends matches how z scales with model size
| Design | Baseline | z with size | Predicted | Observed |
|---|---|---|---|---|
| Fixed U (§5.1) | data-matched | minimum at fewer epochs | 15.3 to 3.6 epochs, 127M to 2B | |
| Fixed U/N (§5.2) | data-matched | minimum nearly fixed | 3.7–4.4 epochs on CommonCrawl | |
| Fixed D/N, f, (§5.3) | compute-matched | 22–39% less at 1.2B | 33% and 21% less | |
| Fixed U, T5 (Xue et al., 2023) | compute-matched | larger models degrade more | larger models degrade more | |
| R = 4, D ≈ 20N (Muennighoff et al., 2023) | compute-matched | ≈ 0.6 | bpb | nearly as good as fresh data |
| Small subset, many repeats (Hernandez et al., 2022) | compute-matched | large | large excess loss | marked degradation |
6. Beyond Token Counts
Partial repetition: ()—cheaper than full-corpus repetition suggests
Order effects: Consecutive replay raises loss by up to 0.46 bpb; shuffling changes it by <0.01 bpb
Source entropy: Lower-entropy sources degrade faster—GitHub ends 0.16 bpb above minimum at 16 epochs; Books doesn't rise (correlation with unigram entropy: to )
Re-tokenization (BPE-dropout): Helps only at ; at reduces loss rise from 0.218 to 0.054 bpb, but hurts at (raises loss by 0.021 bpb at )
Theoretical and Practical Implications
Practical Guidance
-
Epoch counts must be transferred with unique data: Scaling unique data with model size keeps optimal epochs stable (~4); fixing U drops the minimum from 15 to 4 epochs between 127M and 2B
-
Budget determines repetition tolerance: Larger TPP tolerates more epochs; grows as
-
Allocation at fixed data: Grow model size and epochs together (roughly each) until loss stops improving near ; additional compute beyond that point is wasted
-
Layout matters: Space repeats across epochs, spread over samples, repeat lower-entropy sources less, re-tokenize only under heavy reuse
Methodological Implications
- Report baselines and fixed quantities: Comparing repetition studies requires matching baseline (data-matched vs. compute-matched) and the quantity held fixed (U, U/N, or D/N)
- Measured curves over fitted laws: Effective-data formulations cannot represent the largest repetition costs and amplify residuals when inverted
- Separate annealing: Compare choices at separately annealed horizons since intermediate checkpoints can misrank
Conclusion
The paper establishes that the cost of data repetition in pretraining follows a simple geometric structure governed by the variable :
- Cost: bpb
- Critical value threshold:
- Size trends: Determined entirely by how z scales with size under each experimental design
Beyond counts, the value of repetition depends on ordering (spaced > consecutive), allocation (spread > concentrated), source entropy (higher entropy tolerates more), and tokenization (re-tokenization helps only under heavy reuse).
Limitations: Models far smaller than frontier systems; no weight decay or dropout in recipes (which could shift constants); held-out loss rather than downstream capabilities measured.
Future directions: Larger models with iso-FLOP runs locating loss minima directly, complex mixtures with per-source repetition rates, and downstream evaluation.
Related papers
- The Best Optimizer Depends on Batch Size
Optimizer rankings reverse across batch sizes even after extensive tuning, driven by a bias-variance tradeoff in preconditioning that makes single-exponent scaling rules fundamentally limited.
- Distributionally Robust Mixture-of-Experts Training
DRMoET, a distributionally robust MoE training objective, improves downstream task accuracy by up to 2.1% over baselines by optimizing worst-case routing outcomes rather than just load balancing.
- Hyperparameter Scaling Laws Across MoE Sparsity
Unified hyperparameter scaling laws for MoE models treat activation ratio as a multiplicative power-law factor, enabling reliable learning rate and batch size transfer across sparsity levels.