# What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

> Excess loss from data repetition follows a single power law in (R-1)N/U, unifying conflicting size trends and enabling compute-optimal epoch allocation.

- **Source:** [arXiv](https://arxiv.org/abs/2610.05591)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/N4R6Hz
- **Whiteboard:** https://picx.dev/p/N4R6Hz/image

## Summary

# What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

## Summary (Overview)

- **Unified scaling law for repetition cost**: The paper demonstrates that excess loss from data repetition follows a single power law in the variable $z = (R-1)N_{\text{total}}/U$, where $R$ is epochs, $N_{\text{total}}$ is total parameters (including embeddings), and $U$ is unique tokens—no additional size-dependent terms are needed.

- **Critical epoch count**: Repeated tokens fall to half the value of fresh ones at $R_c \approx 2.7(D/N)^{0.24}$, which grows with training budget per parameter but is nearly independent of model size.

- **Resolution of conflicting size trends**: The direction of size-dependent repetition effects depends entirely on which quantity is held fixed (U, U/N, or D/N), explaining seemingly contradictory findings in prior literature.

- **Compute-optimal allocation at fixed data**: With unique data fixed, optimal runs grow model size and epochs together (each roughly as $\sqrt{C}$) until loss stops improving near the critical epoch count $R_c$.

- **Beyond token counts**: At identical counts, repetition order, allocation across samples, source entropy, and tokenization all significantly affect loss—consecutive replay can raise loss by up to 0.46 bits per byte.

## Introduction and Theoretical Foundation

### Background
Compute-optimal scaling laws (Kaplan et al., 2020) prescribe growing training sets with model size, but the finite stock of public human-generated text (Villalobos et al., 2024) forces pretraining to increasingly repeat data. With $U$ unique tokens, a model of $N$ parameters trained for $R$ epochs sees $D = RU$ training tokens.

### Key Questions
1. How many epochs should a run take?
2. How should this number change with model size?
3. Does anything besides the epoch count matter?

### Prior Conflicting Evidence
- **Muennighoff et al. (2023)**: Up to ~4 epochs nearly as valuable as fresh data
- **Hernandez et al. (2022)**: Repeating subsets degrades models
- **Xue et al. (2023)**: Degradation worsens with model size
- **Lovelace et al. (2026)**: Penalty term in $(R-1)N/U$, prescribing fewer epochs at large compute
- **Qin et al. (2026)**: Token effectiveness falls with tokens per parameter

### Theoretical Framework
The paper defines two baselines for evaluating repeated tokens:

**Data-matched baseline**: Same $N$ and $U$, extra compute from additional epochs. The value of a repeated token (fresh-token equivalence):

$$\eta = \frac{U_{\text{eq}} - U}{D - U}, \tag{1}$$

where $U_{\text{eq}}$ is the fresh budget whose one-epoch loss equals that of the repeated run. $\eta = 1$ means repeats are as good as fresh data; $\eta = 0$ means worthless.

**Compute-matched baseline**: Same $N$ and $D$, repeated data replace fresh data. The excess loss:

$$\Delta L = L(N, U, R) - L(N, D, 1) \tag{2}$$

## Methodology

### Experimental Setup
- **Models**: Decoder-only Transformers (23M to 2B parameters), trained with Muon optimizer for hidden matrices and Adam elsewhere
- **Data**: SlimPajama (cleaned/deduplicated RedPajama), GPT-2 byte-level BPE tokenization
- **Evaluation**: Bits per byte (bpb) on held-out test sets
- **Sources**: Seven subsets (CommonCrawl, GitHub, Books, etc.), trained separately and in natural mixtures

### Experimental Designs

**CommonCrawl Grid**: Model sizes 23M–361M, training budgets 5–300 TPP, epoch counts 1–16. Each size has a measured one-epoch curve providing both baselines.

**Fixed-U sweeps**: Eight sizes (14M–2B) on same 1B unique tokens

**Fixed-U/N sweeps**: Seven sizes with $U = 2.5N$

**Fixed-D/N sweeps**: Five sizes at $D = 20N$ with partial repetition

### Key Measurement Approach
Unlike prior work that inverts fitted scaling laws, the authors read $U_{\text{eq}}$ directly from **measured** one-epoch curves via interpolation. This avoids two problems:
1. Effective-data laws cannot represent runs that end above one-epoch loss (about 25% of repeated runs do)
2. Inverting fitted laws amplifies small loss residuals when curves are flat

## Empirical Validation / Results

### 1. Excess Loss Scaling Law

The central empirical finding across the entire grid:

$$\Delta L = A z^{p}, \qquad z = (R - 1) \frac{N_{\text{total}}}{U} = \frac{(R-1)R}{D/N} \cdot \frac{N_{\text{total}}}{N}, \tag{3}$$

with $A = 0.0203$ bpb and $p = 1.07$ ($R^2 = 0.96$). The exponent on an additional size term $N_{\text{total}}^{-g}$ is negligible ($g = -0.02$).

### 2. Value of Repeated Tokens
- **Second epoch**: Worth 0.87–1.02 of fresh tokens at every budget
- **Value decay**: Falls fastest at small budgets per parameter (from ~0.9 to 0 at $D/N = 5$ over 2–8 epochs)
- **At large budgets** ($D/N = 300$): Falls only from ~0.9 to 0.65

### 3. Critical Epoch Count

$$R_c \approx 2.7(D/N)^{0.24} \tag{4}$$

- About 4–5 epochs up to TPP of 20
- Rises to ~10 at TPP of 300
- Nearly independent of model size

### 4. Compute-Optimal Allocation at Fixed U

With $D/N$ fixed at $\tau$, $C = 6\tau N^2$ gives:

$$N_{\text{opt}} = \sqrt{C/(6\tau)}, \quad R_{\text{opt}} = \tau N_{\text{opt}}/U = \sqrt{\tau C/6}\big/U. \tag{5}$$

Fitted exponents: $N_{\text{opt}} \propto C^{0.52-0.53}$ and $R_{\text{opt}} \propto C^{0.47-0.48}$.

**Key results**:
- TPP held at 16–24 (below fresh-data optimum of ~25)
- Loss minimum at 164M params/5.3 epochs ($U = 0.5$B), 297M/5.6 epochs ($U = 1$B)
- Beyond the minimum, more compute raises loss

### 5. Size Trends by Design

**Table 1: Direction of size trends matches how z scales with model size**

| Design | Baseline | z with size | Predicted | Observed |
|--------|----------|------------|-----------|----------|
| Fixed U (§5.1) | data-matched | $\propto N_{\text{total}}$ | minimum at fewer epochs | 15.3 to 3.6 epochs, 127M to 2B |
| Fixed U/N (§5.2) | data-matched | $\propto N_{\text{total}}/N$ | minimum nearly fixed | 3.7–4.4 epochs on CommonCrawl |
| Fixed D/N, f, $R_{\text{sub}}$ (§5.3) | compute-matched | $\propto N_{\text{total}}/N$ | 22–39% less at 1.2B | 33% and 21% less |
| Fixed U, T5 (Xue et al., 2023) | compute-matched | $\propto N_{\text{total}}$ | larger models degrade more | larger models degrade more |
| R = 4, D ≈ 20N (Muennighoff et al., 2023) | compute-matched | ≈ 0.6 | $\Delta L \approx 0.01$ bpb | nearly as good as fresh data |
| Small subset, many repeats (Hernandez et al., 2022) | compute-matched | large | large excess loss | marked degradation |

### 6. Beyond Token Counts

**Partial repetition**: $\Delta L \propto f^{1.63} z_{\text{sub}}^{0.53}$ ($R^2 = 0.89$)—cheaper than full-corpus repetition suggests

**Order effects**: Consecutive replay raises loss by up to 0.46 bpb; shuffling changes it by <0.01 bpb

**Source entropy**: Lower-entropy sources degrade faster—GitHub ends 0.16 bpb above minimum at 16 epochs; Books doesn't rise (correlation with unigram entropy: $r = -0.89$ to $-0.91$)

**Re-tokenization (BPE-dropout)**: Helps only at $R \geq 8$; at $R = 16$ reduces loss rise from 0.218 to 0.054 bpb, but hurts at $R \leq 4$ (raises loss by 0.021 bpb at $R = 1$)

## Theoretical and Practical Implications

### Practical Guidance
1. **Epoch counts must be transferred with unique data**: Scaling unique data with model size keeps optimal epochs stable (~4); fixing U drops the minimum from 15 to 4 epochs between 127M and 2B

2. **Budget determines repetition tolerance**: Larger TPP tolerates more epochs; $R_c$ grows as $(D/N)^{0.24}$

3. **Allocation at fixed data**: Grow model size and epochs together (roughly $\sqrt{C}$ each) until loss stops improving near $R_c$; additional compute beyond that point is wasted

4. **Layout matters**: Space repeats across epochs, spread over samples, repeat lower-entropy sources less, re-tokenize only under heavy reuse

### Methodological Implications
- **Report baselines and fixed quantities**: Comparing repetition studies requires matching baseline (data-matched vs. compute-matched) and the quantity held fixed (U, U/N, or D/N)
- **Measured curves over fitted laws**: Effective-data formulations cannot represent the largest repetition costs and amplify residuals when inverted
- **Separate annealing**: Compare choices at separately annealed horizons since intermediate checkpoints can misrank

## Conclusion

The paper establishes that the cost of data repetition in pretraining follows a simple geometric structure governed by the variable $z = (R-1)N_{\text{total}}/U$:

- **Cost**: $\Delta L = 0.0203 \cdot z^{1.07}$ bpb
- **Critical value threshold**: $R_c \approx 2.7(D/N)^{0.24}$
- **Size trends**: Determined entirely by how z scales with size under each experimental design

Beyond counts, the value of repetition depends on ordering (spaced > consecutive), allocation (spread > concentrated), source entropy (higher entropy tolerates more), and tokenization (re-tokenization helps only under heavy reuse).

**Limitations**: Models far smaller than frontier systems; no weight decay or dropout in recipes (which could shift constants); held-out loss rather than downstream capabilities measured.

**Future directions**: Larger models with iso-FLOP runs locating loss minima directly, complex mixtures with per-source repetition rates, and downstream evaluation.

---

_Markdown view of https://picx.dev/p/N4R6Hz, served by PicX — AI-generated visual whiteboard summaries of research papers._
