# When do data mixtures improve scaling laws? Insights from high-dimensional regression

> Data mixtures provably accelerate scaling laws only when auxiliary data has heavier-tailed spectra and intermediate relative sample growth, with ridge regression achieving the optimal rate.

- **Source:** [arXiv](https://arxiv.org/abs/2609.38011)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/ogvNal
- **Whiteboard:** https://picx.dev/p/ogvNal/image

## Summary

# Summary

## Summary (Overview)

- **Core contribution**: This paper provides the first rigorous theoretical characterization of **when data mixtures improve scaling laws** in high-dimensional regression, identifying precise conditions under which mixing datasets yields faster scaling rates than using either dataset alone.

- **Key theoretical results**: The authors establish (i) the minimax risk under ellipsoidal parameter constraints for general covariance structures (Theorem 1), and (ii) deterministic equivalents for the test error of mixed ridge regression under commutative covariances (Theorem 2).

- **Scaling law identification**: For aligned power-law covariance spectra with target and auxiliary domains, the paper identifies a precise regime—**heavy-tailed auxiliary spectra (δ > 1) with intermediate relative sample growth (γ_c < γ_2 < 1)**—where data mixtures provably improve the scaling exponent over both single-dataset rates.

- **Ridge regression optimality**: Whenever the mixed minimax rate strictly improves upon either dataset alone, ridge regression achieves this improved rate (Corollary 1), even in regimes where target-only ridge suffers from saturation.

- **Empirical validation**: Numerical experiments on linear models, real image features, and **decoder-only transformers trained on SlimPajama** qualitatively confirm the theoretical findings, showing that appropriate data mixtures yield faster test-loss decreases than training on either domain alone.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Modern machine learning systems are trained on mixtures of data from multiple domains, and choosing the right mixture is a central design choice [GD+24, ZY+23]. Despite extensive empirical work on data mixing—including mixture weight transfer [XPD+23, FPJ24, LZM+25], budget-dependent optimization [KSW+25], and adaptive reweighting [CRB+23, CHL+25, JZF+25]—a fundamental question remains open: **when does auxiliary data genuinely improve scaling laws rather than merely providing more samples?**

### Problem Setting

The paper studies **high-dimensional regression** with $K$ training domains sharing a common regression function $\theta^*$. For domain $i \in [K]$:

$$y_i = X_i \theta^* + \varepsilon_i, \quad \varepsilon_i \sim N(0, \sigma^2_{\varepsilon_i} I_{n_i})$$

where $X_i \in \mathbb{R}^{n_i \times d}$ has rows drawn from distribution $P_i$ with $\mathbb{E}[x] = 0$ and $\mathbb{E}[xx^\top] = C_i$. The test distribution is a mixture $P = \sum_k \pi^*_k P_k$ of the domains.

The **excess test risk** for an estimator $\hat{\theta}$ is:

$$\mathcal{R}_\varepsilon(\hat{\theta}) = \mathbb{E}_{x \sim P}\left[\left(x^\top \hat{\theta} - x^\top \theta^*\right)^2\right] = \sum_{k=1}^{K} \pi^*_k \left\|C_k^{1/2}(\hat{\theta} - \theta^*)\right\|_2^2$$

### Key Theoretical Framework

**Mixed ridge estimator**:
$$\hat{\theta} = \arg\min_{\theta} \left\{\sum_{i=1}^{K} \|X_i\theta - y_i\|_2^2 + \lambda\|\theta\|_2^2\right\} = G \sum_{i=1}^{K} X_i^\top y_i$$

with $G = (\sum_{i=1}^{K} X_i^\top X_i + \lambda I)^{-1}$.

The paper operates under **Assumption 1** (concentration of designs), which covers sub-Gaussian coordinates and convex Lipschitz concentration [CM24, MS24]. The setting connects to kernel methods: with $x = \psi(z)$ for a feature map $\psi$, the model becomes kernel ridge regression with data mixture.

---

## Methodology

### Minimax Risk Characterization (Theorem 1)

Under an **ellipsoidal parameter constraint** $\Theta = \{\theta : \|S^{1/2}\theta\|_2 \leq R\}$, the minimax risk satisfies:

$$\mathcal{R}^*(n_1, \ldots, n_K) \asymp \mathcal{L}(M)$$

where $M = \sum_{i=1}^{K} \frac{n_i}{\sigma^2_{\varepsilon_i}} H_i$ (with $H_i = S^{-1/2}C_i S^{-1/2}$) is the **information operator**, and $\mathcal{L}(M)$ is the variational functional:

$$\mathcal{L}(M) = \sup_{A \succeq 0, \text{Tr}(A) \leq R^2} \text{Tr}\left[Q A^{1/2}\left(I + A^{1/2} M A^{1/2}\right)^{-1} A^{1/2}\right]$$

with $Q = \sum_k \pi^*_k H_k$.

**Key insight**: A heterogeneous dataset mixture is equivalent (at the minimax level) to a **single Gaussian sequence model** with information operator $M$. Domain $i$ contributes information proportional to $n_i/\sigma^2_{\varepsilon_i}$, weighted by its covariance in each direction.

For commuting covariances, the variational problem reduces to:

$$\mathcal{R}^*(n_1, \ldots, n_K) \asymp \inf_{m \geq 0} \left\{R^2 \max_{j>m} [Q]_{jj} + \sum_{j \leq m} \frac{[Q]_{jj}}{[M]_{jj}}\right\}$$

### Deterministic Equivalent (Theorem 2)

For commutative covariances, the risk of mixed ridge regression is characterized by a **fixed-point system**:

$$\mu_i = \frac{n_i}{1 + \text{Tr}(C_i \bar{G})}, \quad \bar{G} = \left(\sum_{i=1}^{K} \mu_i C_i + \lambda\right)^{-1}$$

This gives a deterministic proxy where each dataset is replaced by its population covariance scaled by an **effective sample size** $\mu_i \leq n_i$.

### Power-Law Model

Specialized to $K=2$ domains with aligned spectra:

$$[C_1]_{kk} = k^{-\alpha_1}, \quad [C_2]_{kk} = k^{-\alpha_2}, \quad n_1 = n, \quad n_2 = \lfloor n^{\gamma_2}\rfloor$$

with $\delta = \alpha_1 - \alpha_2 > 0$ (auxiliary spectrum has heavier tails). The parameter class imposes a **source condition** with regularity parameter $s$.

---

## Empirical Validation / Results

### Minimax Scaling Laws (Theorem 3)

Define the individual and mixed exponents:
- $\Gamma_{\text{tar}} = \frac{2s}{1+2s}$
- $\Gamma_{\text{aux}} = \begin{cases} \frac{2s\gamma_2}{1+2s-\delta}, & \delta < 1 \\ \gamma_2, & \delta \geq 1 \end{cases}$

**The mixed minimax rate improves on both single-dataset rates if and only if $\delta > 1$ and $\gamma_c < \gamma_2 < 1$**, where $\gamma_c = \frac{1-\delta}{1+2s}$.

Three regimes emerge:
1. **Too few auxiliary samples** ($\gamma_2 \leq \gamma_c$): mixing doesn't help—auxiliary data only aids coordinates already well-estimated
2. **Light auxiliary tails** ($\delta \leq 1$): no improvement—variance dominated by high-frequency coordinates
3. **Heavy auxiliary tails** ($\delta > 1$, $\gamma_c < \gamma_2 < 1$): **mixing strictly improves the rate**

### Ridge Regression Scaling Laws (Theorem 4)

With truncation parameters $a = \min\{s, \alpha_1\}$ and $b = \min\{s, \alpha_2\}$:

**Corollary 1**: Ridge regression achieves the mixed minimax rate **whenever the mixed minimax rate strictly improves** (δ > 1, γ_c < γ_2 < 1), even for $s > \alpha_1$ where target-only ridge suffers from saturation.

| Regime | Mixed Risk Rate | Improvement |
|--------|-----------------|-------------|
| $\gamma_2 \leq \gamma_c$ | $\tilde{\Theta}(n^{-\Gamma_{\text{tar}}})$ | Target-only optimal |
| $\gamma_2 \geq 1$, δ > 1 | $\tilde{\Theta}(n^{-\Gamma_{\text{aux}}})$ | Auxiliary-only optimal |
| $\gamma_c < \gamma_2 < 1$, δ > 1 | $\tilde{\Theta}(n^{-(\delta-1+\gamma_2)/\delta})$ | **Strict improvement** |

### Experimental Results

**Linear models** (Figure 1): Empirical risk of mixed ridge matches the deterministic equivalent and predicted scaling laws. For $\gamma_2 = 0.78$ with δ > 1, the mixed rate ($\propto n^{-0.890}$) improves on both target-only ($\propto n^{-0.773}$) and auxiliary-only ($\propto n^{-0.780}$) rates.

**Real image features** (Figure 2): Deterministic equivalent predictions hold for CIFAR-10 and ImageNet-100 features extracted via pretrained encoders, beyond the technical assumptions.

**Language models** (Figure 3): An 81.5M-parameter GPT-2-style model trained on SlimPajama shows:
- $\gamma_2 = 1.25$: data mixture improves the scaling law
- $\gamma_2 = 0.75$ or $1.75$: no noticeable improvement
- Token-frequency analysis reveals the auxiliary domain places more mass on tokens infrequent in the target domain—a **token-level analogue of the heavier-tailed auxiliary covariance**

---

## Theoretical and Practical Implications

### Theoretical Significance

1. **First characterization of when mixing helps**: The paper resolves the open question of whether data mixtures improve scaling rates or merely provide more samples, identifying the precise interplay between spectral decay, target regularity, and relative sample sizes.

2. **Positive distribution shift**: At equal sample sizes ($\gamma_2 = 1$), auxiliary data with heavier tails is *more informative* about the target than target data itself—a formal demonstration of positive transfer.

3. **Ridge regression is minimax optimal in the improvement regime**: Saturation effects (Tikhonov regularization) are compensated by the auxiliary data when its spectral tail is sufficiently heavy.

### Practical Implications

1. **Data mixture design guidelines**: Auxiliary data should have **heavier-tailed feature/token distributions** relative to the target, and its volume should grow at an **intermediate rate** relative to target data.

2. **Budget allocation**: The regime $\gamma_c < \gamma_2 < 1$ indicates the auxiliary budget should grow sub-linearly relative to target data—too little or too much auxiliary data both fail to improve scaling.

3. **Language model pretraining**: Token-level frequency distributions serve as a practical proxy for covariance spectra: mixing domains with complementary frequency coverage (heavier tails) at appropriate proportions yields measurable scaling improvements.

---

## Conclusion

### Main Takeaways

- Data mixtures **provably improve scaling laws** when: (i) the auxiliary spectrum is heavy-tailed relative to the target (δ > 1), meaning the two datasets cover complementary spectral directions; and (ii) the auxiliary sample size grows at an intermediate rate ($\gamma_c < \gamma_2 < 1$)

- In this regime, **ridge regression attains the minimax optimal rate**, achieving a scaling exponent strictly better than either dataset alone

- The qualitative predictions hold for **real language models**, where token-frequency distributions serve as a practical analogue of covariance spectra

### Future Directions

1. **Gradient descent dynamics**: Determining whether gradient descent methods attain the mixed minimax rates when ridge saturates (building on [WBK+26])

2. **Random-feature models**: Identifying how model size must scale with data budget to preserve mixing improvements [DLM24, BAP24]

3. **Feature learning**: Extending to features trained via gradient descent steps [BES+22, CPD+24]

4. **Joint optimization**: Optimizing domain proportions under combined data and compute constraints, and characterizing mixture transfer across model scales [GWB26]

The authors note open technical questions: extending the deterministic equivalent to non-commutative covariances (conjectured to hold) and achieving multiplicative (rather than additive) error guarantees in Theorem 2.

---

_Markdown view of https://picx.dev/p/ogvNal, served by PicX — AI-generated visual whiteboard summaries of research papers._
