Full text not available for this paper

Summary

Summary (Overview)

  • Core contribution: This paper provides the first rigorous theoretical characterization of when data mixtures improve scaling laws in high-dimensional regression, identifying precise conditions under which mixing datasets yields faster scaling rates than using either dataset alone.

  • Key theoretical results: The authors establish (i) the minimax risk under ellipsoidal parameter constraints for general covariance structures (Theorem 1), and (ii) deterministic equivalents for the test error of mixed ridge regression under commutative covariances (Theorem 2).

  • Scaling law identification: For aligned power-law covariance spectra with target and auxiliary domains, the paper identifies a precise regime—heavy-tailed auxiliary spectra (δ > 1) with intermediate relative sample growth (γ_c < γ_2 < 1)—where data mixtures provably improve the scaling exponent over both single-dataset rates.

  • Ridge regression optimality: Whenever the mixed minimax rate strictly improves upon either dataset alone, ridge regression achieves this improved rate (Corollary 1), even in regimes where target-only ridge suffers from saturation.

  • Empirical validation: Numerical experiments on linear models, real image features, and decoder-only transformers trained on SlimPajama qualitatively confirm the theoretical findings, showing that appropriate data mixtures yield faster test-loss decreases than training on either domain alone.


Introduction and Theoretical Foundation

Background and Motivation

Modern machine learning systems are trained on mixtures of data from multiple domains, and choosing the right mixture is a central design choice [GD+24, ZY+23]. Despite extensive empirical work on data mixing—including mixture weight transfer [XPD+23, FPJ24, LZM+25], budget-dependent optimization [KSW+25], and adaptive reweighting [CRB+23, CHL+25, JZF+25]—a fundamental question remains open: when does auxiliary data genuinely improve scaling laws rather than merely providing more samples?

Problem Setting

The paper studies high-dimensional regression with KK training domains sharing a common regression function θ∗\theta^*. For domain i∈[K]i \in [K]:

yi=Xiθ∗+εi,εi∼N(0,σεi2Ini)y_i = X_i \theta^* + \varepsilon_i, \quad \varepsilon_i \sim N(0, \sigma^2_{\varepsilon_i} I_{n_i})

where Xi∈Rni×dX_i \in \mathbb{R}^{n_i \times d} has rows drawn from distribution PiP_i with E[x]=0\mathbb{E}[x] = 0 and E[xx⊤]=Ci\mathbb{E}[xx^\top] = C_i. The test distribution is a mixture P=∑kπk∗PkP = \sum_k \pi^*_k P_k of the domains.

The excess test risk for an estimator θ^\hat{\theta} is:

Rε(θ^)=Ex∼P[(x⊤θ^−x⊤θ∗)2]=∑k=1Kπk∗∥Ck1/2(θ^−θ∗)∥22\mathcal{R}_\varepsilon(\hat{\theta}) = \mathbb{E}_{x \sim P}\left[\left(x^\top \hat{\theta} - x^\top \theta^*\right)^2\right] = \sum_{k=1}^{K} \pi^*_k \left\|C_k^{1/2}(\hat{\theta} - \theta^*)\right\|_2^2

Key Theoretical Framework

Mixed ridge estimator:

θ^=arg⁡min⁡θ{∑i=1K∥Xiθ−yi∥22+λ∥θ∥22}=G∑i=1KXi⊤yi\hat{\theta} = \arg\min_{\theta} \left\{\sum_{i=1}^{K} \|X_i\theta - y_i\|_2^2 + \lambda\|\theta\|_2^2\right\} = G \sum_{i=1}^{K} X_i^\top y_i

with G=(∑i=1KXi⊤Xi+λI)−1G = (\sum_{i=1}^{K} X_i^\top X_i + \lambda I)^{-1}.

The paper operates under Assumption 1 (concentration of designs), which covers sub-Gaussian coordinates and convex Lipschitz concentration [CM24, MS24]. The setting connects to kernel methods: with x=ψ(z)x = \psi(z) for a feature map ψ\psi, the model becomes kernel ridge regression with data mixture.


Methodology

Minimax Risk Characterization (Theorem 1)

Under an ellipsoidal parameter constraint Θ={θ:∥S1/2θ∥2≤R}\Theta = \{\theta : \|S^{1/2}\theta\|_2 \leq R\}, the minimax risk satisfies:

R∗(n1,…,nK)≍L(M)\mathcal{R}^*(n_1, \ldots, n_K) \asymp \mathcal{L}(M)

where M=∑i=1Kniσεi2HiM = \sum_{i=1}^{K} \frac{n_i}{\sigma^2_{\varepsilon_i}} H_i (with Hi=S−1/2CiS−1/2H_i = S^{-1/2}C_i S^{-1/2}) is the information operator, and L(M)\mathcal{L}(M) is the variational functional:

L(M)=sup⁡A⪰0,Tr(A)≤R2Tr[QA1/2(I+A1/2MA1/2)−1A1/2]\mathcal{L}(M) = \sup_{A \succeq 0, \text{Tr}(A) \leq R^2} \text{Tr}\left[Q A^{1/2}\left(I + A^{1/2} M A^{1/2}\right)^{-1} A^{1/2}\right]

with Q=∑kπk∗HkQ = \sum_k \pi^*_k H_k.

Key insight: A heterogeneous dataset mixture is equivalent (at the minimax level) to a single Gaussian sequence model with information operator MM. Domain ii contributes information proportional to ni/σεi2n_i/\sigma^2_{\varepsilon_i}, weighted by its covariance in each direction.

For commuting covariances, the variational problem reduces to:

R∗(n1,…,nK)≍inf⁡m≥0{R2max⁡j>m[Q]jj+∑j≤m[Q]jj[M]jj}\mathcal{R}^*(n_1, \ldots, n_K) \asymp \inf_{m \geq 0} \left\{R^2 \max_{j>m} [Q]_{jj} + \sum_{j \leq m} \frac{[Q]_{jj}}{[M]_{jj}}\right\}

Deterministic Equivalent (Theorem 2)

For commutative covariances, the risk of mixed ridge regression is characterized by a fixed-point system:

μi=ni1+Tr(CiGˉ),Gˉ=(∑i=1KμiCi+λ)−1\mu_i = \frac{n_i}{1 + \text{Tr}(C_i \bar{G})}, \quad \bar{G} = \left(\sum_{i=1}^{K} \mu_i C_i + \lambda\right)^{-1}

This gives a deterministic proxy where each dataset is replaced by its population covariance scaled by an effective sample size μi≤ni\mu_i \leq n_i.

Power-Law Model

Specialized to K=2K=2 domains with aligned spectra:

[C1]kk=k−α1,[C2]kk=k−α2,n1=n,n2=⌊nγ2⌋[C_1]_{kk} = k^{-\alpha_1}, \quad [C_2]_{kk} = k^{-\alpha_2}, \quad n_1 = n, \quad n_2 = \lfloor n^{\gamma_2}\rfloor

with δ=α1−α2>0\delta = \alpha_1 - \alpha_2 > 0 (auxiliary spectrum has heavier tails). The parameter class imposes a source condition with regularity parameter ss.


Empirical Validation / Results

Minimax Scaling Laws (Theorem 3)

Define the individual and mixed exponents:

  • Γtar=2s1+2s\Gamma_{\text{tar}} = \frac{2s}{1+2s}
  • Γaux={2sγ21+2s−δ,δ<1γ2,δ≥1\Gamma_{\text{aux}} = \begin{cases} \frac{2s\gamma_2}{1+2s-\delta}, & \delta < 1 \\ \gamma_2, & \delta \geq 1 \end{cases}

The mixed minimax rate improves on both single-dataset rates if and only if δ>1\delta > 1 and γc<γ2<1\gamma_c < \gamma_2 < 1, where γc=1−δ1+2s\gamma_c = \frac{1-\delta}{1+2s}.

Three regimes emerge:

  1. Too few auxiliary samples (γ2≤γc\gamma_2 \leq \gamma_c): mixing doesn't help—auxiliary data only aids coordinates already well-estimated
  2. Light auxiliary tails (δ≤1\delta \leq 1): no improvement—variance dominated by high-frequency coordinates
  3. Heavy auxiliary tails (δ>1\delta > 1, γc<γ2<1\gamma_c < \gamma_2 < 1): mixing strictly improves the rate

Ridge Regression Scaling Laws (Theorem 4)

With truncation parameters a=min⁡{s,α1}a = \min\{s, \alpha_1\} and b=min⁡{s,α2}b = \min\{s, \alpha_2\}:

Corollary 1: Ridge regression achieves the mixed minimax rate whenever the mixed minimax rate strictly improves (δ > 1, γ_c < γ_2 < 1), even for s>α1s > \alpha_1 where target-only ridge suffers from saturation.

RegimeMixed Risk RateImprovement
γ2≤γc\gamma_2 \leq \gamma_cΘ~(n−Γtar)\tilde{\Theta}(n^{-\Gamma_{\text{tar}}})Target-only optimal
γ2≥1\gamma_2 \geq 1, δ > 1Θ~(n−Γaux)\tilde{\Theta}(n^{-\Gamma_{\text{aux}}})Auxiliary-only optimal
γc<γ2<1\gamma_c < \gamma_2 < 1, δ > 1Θ~(n−(δ−1+γ2)/δ)\tilde{\Theta}(n^{-(\delta-1+\gamma_2)/\delta})Strict improvement

Experimental Results

Linear models (Figure 1): Empirical risk of mixed ridge matches the deterministic equivalent and predicted scaling laws. For γ2=0.78\gamma_2 = 0.78 with δ > 1, the mixed rate (∝n−0.890\propto n^{-0.890}) improves on both target-only (∝n−0.773\propto n^{-0.773}) and auxiliary-only (∝n−0.780\propto n^{-0.780}) rates.

Real image features (Figure 2): Deterministic equivalent predictions hold for CIFAR-10 and ImageNet-100 features extracted via pretrained encoders, beyond the technical assumptions.

Language models (Figure 3): An 81.5M-parameter GPT-2-style model trained on SlimPajama shows:

  • γ2=1.25\gamma_2 = 1.25: data mixture improves the scaling law
  • γ2=0.75\gamma_2 = 0.75 or 1.751.75: no noticeable improvement
  • Token-frequency analysis reveals the auxiliary domain places more mass on tokens infrequent in the target domain—a token-level analogue of the heavier-tailed auxiliary covariance

Theoretical and Practical Implications

Theoretical Significance

  1. First characterization of when mixing helps: The paper resolves the open question of whether data mixtures improve scaling rates or merely provide more samples, identifying the precise interplay between spectral decay, target regularity, and relative sample sizes.

  2. Positive distribution shift: At equal sample sizes (γ2=1\gamma_2 = 1), auxiliary data with heavier tails is more informative about the target than target data itself—a formal demonstration of positive transfer.

  3. Ridge regression is minimax optimal in the improvement regime: Saturation effects (Tikhonov regularization) are compensated by the auxiliary data when its spectral tail is sufficiently heavy.

Practical Implications

  1. Data mixture design guidelines: Auxiliary data should have heavier-tailed feature/token distributions relative to the target, and its volume should grow at an intermediate rate relative to target data.

  2. Budget allocation: The regime γc<γ2<1\gamma_c < \gamma_2 < 1 indicates the auxiliary budget should grow sub-linearly relative to target data—too little or too much auxiliary data both fail to improve scaling.

  3. Language model pretraining: Token-level frequency distributions serve as a practical proxy for covariance spectra: mixing domains with complementary frequency coverage (heavier tails) at appropriate proportions yields measurable scaling improvements.


Conclusion

Main Takeaways

  • Data mixtures provably improve scaling laws when: (i) the auxiliary spectrum is heavy-tailed relative to the target (δ > 1), meaning the two datasets cover complementary spectral directions; and (ii) the auxiliary sample size grows at an intermediate rate (γc<γ2<1\gamma_c < \gamma_2 < 1)

  • In this regime, ridge regression attains the minimax optimal rate, achieving a scaling exponent strictly better than either dataset alone

  • The qualitative predictions hold for real language models, where token-frequency distributions serve as a practical analogue of covariance spectra

Future Directions

  1. Gradient descent dynamics: Determining whether gradient descent methods attain the mixed minimax rates when ridge saturates (building on [WBK+26])

  2. Random-feature models: Identifying how model size must scale with data budget to preserve mixing improvements [DLM24, BAP24]

  3. Feature learning: Extending to features trained via gradient descent steps [BES+22, CPD+24]

  4. Joint optimization: Optimizing domain proportions under combined data and compute constraints, and characterizing mixture transfer across model scales [GWB26]

The authors note open technical questions: extending the deterministic equivalent to non-commutative covariances (conjectured to hold) and achieving multiplicative (rather than additive) error guarantees in Theorem 2.

Related papers