What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

Summary (Overview)

  • Unified scaling law for repetition cost: The paper demonstrates that excess loss from data repetition follows a single power law in the variable z=(R−1)Ntotal/Uz = (R-1)N_{\text{total}}/U, where RR is epochs, NtotalN_{\text{total}} is total parameters (including embeddings), and UU is unique tokens—no additional size-dependent terms are needed.

  • Critical epoch count: Repeated tokens fall to half the value of fresh ones at Rc≈2.7(D/N)0.24R_c \approx 2.7(D/N)^{0.24}, which grows with training budget per parameter but is nearly independent of model size.

  • Resolution of conflicting size trends: The direction of size-dependent repetition effects depends entirely on which quantity is held fixed (U, U/N, or D/N), explaining seemingly contradictory findings in prior literature.

  • Compute-optimal allocation at fixed data: With unique data fixed, optimal runs grow model size and epochs together (each roughly as C\sqrt{C}) until loss stops improving near the critical epoch count RcR_c.

  • Beyond token counts: At identical counts, repetition order, allocation across samples, source entropy, and tokenization all significantly affect loss—consecutive replay can raise loss by up to 0.46 bits per byte.

Introduction and Theoretical Foundation

Background

Compute-optimal scaling laws (Kaplan et al., 2020) prescribe growing training sets with model size, but the finite stock of public human-generated text (Villalobos et al., 2024) forces pretraining to increasingly repeat data. With UU unique tokens, a model of NN parameters trained for RR epochs sees D=RUD = RU training tokens.

Key Questions

  1. How many epochs should a run take?
  2. How should this number change with model size?
  3. Does anything besides the epoch count matter?

Prior Conflicting Evidence

  • Muennighoff et al. (2023): Up to ~4 epochs nearly as valuable as fresh data
  • Hernandez et al. (2022): Repeating subsets degrades models
  • Xue et al. (2023): Degradation worsens with model size
  • Lovelace et al. (2026): Penalty term in (R−1)N/U(R-1)N/U, prescribing fewer epochs at large compute
  • Qin et al. (2026): Token effectiveness falls with tokens per parameter

Theoretical Framework

The paper defines two baselines for evaluating repeated tokens:

Data-matched baseline: Same NN and UU, extra compute from additional epochs. The value of a repeated token (fresh-token equivalence):

η=Ueq−UD−U,(1)\eta = \frac{U_{\text{eq}} - U}{D - U}, \tag{1}

where UeqU_{\text{eq}} is the fresh budget whose one-epoch loss equals that of the repeated run. η=1\eta = 1 means repeats are as good as fresh data; η=0\eta = 0 means worthless.

Compute-matched baseline: Same NN and DD, repeated data replace fresh data. The excess loss:

ΔL=L(N,U,R)−L(N,D,1)(2)\Delta L = L(N, U, R) - L(N, D, 1) \tag{2}

Methodology

Experimental Setup

  • Models: Decoder-only Transformers (23M to 2B parameters), trained with Muon optimizer for hidden matrices and Adam elsewhere
  • Data: SlimPajama (cleaned/deduplicated RedPajama), GPT-2 byte-level BPE tokenization
  • Evaluation: Bits per byte (bpb) on held-out test sets
  • Sources: Seven subsets (CommonCrawl, GitHub, Books, etc.), trained separately and in natural mixtures

Experimental Designs

CommonCrawl Grid: Model sizes 23M–361M, training budgets 5–300 TPP, epoch counts 1–16. Each size has a measured one-epoch curve providing both baselines.

Fixed-U sweeps: Eight sizes (14M–2B) on same 1B unique tokens

Fixed-U/N sweeps: Seven sizes with U=2.5NU = 2.5N

Fixed-D/N sweeps: Five sizes at D=20ND = 20N with partial repetition

Key Measurement Approach

Unlike prior work that inverts fitted scaling laws, the authors read UeqU_{\text{eq}} directly from measured one-epoch curves via interpolation. This avoids two problems:

  1. Effective-data laws cannot represent runs that end above one-epoch loss (about 25% of repeated runs do)
  2. Inverting fitted laws amplifies small loss residuals when curves are flat

Empirical Validation / Results

1. Excess Loss Scaling Law

The central empirical finding across the entire grid:

ΔL=Azp,z=(R−1)NtotalU=(R−1)RD/N⋅NtotalN,(3)\Delta L = A z^{p}, \qquad z = (R - 1) \frac{N_{\text{total}}}{U} = \frac{(R-1)R}{D/N} \cdot \frac{N_{\text{total}}}{N}, \tag{3}

with A=0.0203A = 0.0203 bpb and p=1.07p = 1.07 (R2=0.96R^2 = 0.96). The exponent on an additional size term Ntotal−gN_{\text{total}}^{-g} is negligible (g=−0.02g = -0.02).

2. Value of Repeated Tokens

  • Second epoch: Worth 0.87–1.02 of fresh tokens at every budget
  • Value decay: Falls fastest at small budgets per parameter (from ~0.9 to 0 at D/N=5D/N = 5 over 2–8 epochs)
  • At large budgets (D/N=300D/N = 300): Falls only from ~0.9 to 0.65

3. Critical Epoch Count

Rc≈2.7(D/N)0.24(4)R_c \approx 2.7(D/N)^{0.24} \tag{4}
  • About 4–5 epochs up to TPP of 20
  • Rises to ~10 at TPP of 300
  • Nearly independent of model size

4. Compute-Optimal Allocation at Fixed U

With D/ND/N fixed at τ\tau, C=6τN2C = 6\tau N^2 gives:

Nopt=C/(6τ),Ropt=τNopt/U=τC/6/U.(5)N_{\text{opt}} = \sqrt{C/(6\tau)}, \quad R_{\text{opt}} = \tau N_{\text{opt}}/U = \sqrt{\tau C/6}\big/U. \tag{5}

Fitted exponents: Nopt∝C0.52−0.53N_{\text{opt}} \propto C^{0.52-0.53} and Ropt∝C0.47−0.48R_{\text{opt}} \propto C^{0.47-0.48}.

Key results:

  • TPP held at 16–24 (below fresh-data optimum of ~25)
  • Loss minimum at 164M params/5.3 epochs (U=0.5U = 0.5B), 297M/5.6 epochs (U=1U = 1B)
  • Beyond the minimum, more compute raises loss

5. Size Trends by Design

Table 1: Direction of size trends matches how z scales with model size

DesignBaselinez with sizePredictedObserved
Fixed U (§5.1)data-matched∝Ntotal\propto N_{\text{total}}minimum at fewer epochs15.3 to 3.6 epochs, 127M to 2B
Fixed U/N (§5.2)data-matched∝Ntotal/N\propto N_{\text{total}}/Nminimum nearly fixed3.7–4.4 epochs on CommonCrawl
Fixed D/N, f, RsubR_{\text{sub}} (§5.3)compute-matched∝Ntotal/N\propto N_{\text{total}}/N22–39% less at 1.2B33% and 21% less
Fixed U, T5 (Xue et al., 2023)compute-matched∝Ntotal\propto N_{\text{total}}larger models degrade morelarger models degrade more
R = 4, D ≈ 20N (Muennighoff et al., 2023)compute-matched≈ 0.6ΔL≈0.01\Delta L \approx 0.01 bpbnearly as good as fresh data
Small subset, many repeats (Hernandez et al., 2022)compute-matchedlargelarge excess lossmarked degradation

6. Beyond Token Counts

Partial repetition: ΔL∝f1.63zsub0.53\Delta L \propto f^{1.63} z_{\text{sub}}^{0.53} (R2=0.89R^2 = 0.89)—cheaper than full-corpus repetition suggests

Order effects: Consecutive replay raises loss by up to 0.46 bpb; shuffling changes it by <0.01 bpb

Source entropy: Lower-entropy sources degrade faster—GitHub ends 0.16 bpb above minimum at 16 epochs; Books doesn't rise (correlation with unigram entropy: r=−0.89r = -0.89 to −0.91-0.91)

Re-tokenization (BPE-dropout): Helps only at R≥8R \geq 8; at R=16R = 16 reduces loss rise from 0.218 to 0.054 bpb, but hurts at R≤4R \leq 4 (raises loss by 0.021 bpb at R=1R = 1)

Theoretical and Practical Implications

Practical Guidance

  1. Epoch counts must be transferred with unique data: Scaling unique data with model size keeps optimal epochs stable (~4); fixing U drops the minimum from 15 to 4 epochs between 127M and 2B

  2. Budget determines repetition tolerance: Larger TPP tolerates more epochs; RcR_c grows as (D/N)0.24(D/N)^{0.24}

  3. Allocation at fixed data: Grow model size and epochs together (roughly C\sqrt{C} each) until loss stops improving near RcR_c; additional compute beyond that point is wasted

  4. Layout matters: Space repeats across epochs, spread over samples, repeat lower-entropy sources less, re-tokenize only under heavy reuse

Methodological Implications

  • Report baselines and fixed quantities: Comparing repetition studies requires matching baseline (data-matched vs. compute-matched) and the quantity held fixed (U, U/N, or D/N)
  • Measured curves over fitted laws: Effective-data formulations cannot represent the largest repetition costs and amplify residuals when inverted
  • Separate annealing: Compare choices at separately annealed horizons since intermediate checkpoints can misrank

Conclusion

The paper establishes that the cost of data repetition in pretraining follows a simple geometric structure governed by the variable z=(R−1)Ntotal/Uz = (R-1)N_{\text{total}}/U:

  • Cost: ΔL=0.0203⋅z1.07\Delta L = 0.0203 \cdot z^{1.07} bpb
  • Critical value threshold: Rc≈2.7(D/N)0.24R_c \approx 2.7(D/N)^{0.24}
  • Size trends: Determined entirely by how z scales with size under each experimental design

Beyond counts, the value of repetition depends on ordering (spaced > consecutive), allocation (spread > concentrated), source entropy (higher entropy tolerates more), and tokenization (re-tokenization helps only under heavy reuse).

Limitations: Models far smaller than frontier systems; no weight decay or dropout in recipes (which could shift constants); held-out loss rather than downstream capabilities measured.

Future directions: Larger models with iso-FLOP runs locating loss minima directly, complex mixtures with per-source repetition rates, and downstream evaluation.

Related papers