Full text not available for this paper
Summary
Summary (Overview)
-
Core contribution: This paper provides the first rigorous theoretical characterization of when data mixtures improve scaling laws in high-dimensional regression, identifying precise conditions under which mixing datasets yields faster scaling rates than using either dataset alone.
-
Key theoretical results: The authors establish (i) the minimax risk under ellipsoidal parameter constraints for general covariance structures (Theorem 1), and (ii) deterministic equivalents for the test error of mixed ridge regression under commutative covariances (Theorem 2).
-
Scaling law identification: For aligned power-law covariance spectra with target and auxiliary domains, the paper identifies a precise regime—heavy-tailed auxiliary spectra (δ > 1) with intermediate relative sample growth (γ_c < γ_2 < 1)—where data mixtures provably improve the scaling exponent over both single-dataset rates.
-
Ridge regression optimality: Whenever the mixed minimax rate strictly improves upon either dataset alone, ridge regression achieves this improved rate (Corollary 1), even in regimes where target-only ridge suffers from saturation.
-
Empirical validation: Numerical experiments on linear models, real image features, and decoder-only transformers trained on SlimPajama qualitatively confirm the theoretical findings, showing that appropriate data mixtures yield faster test-loss decreases than training on either domain alone.
Introduction and Theoretical Foundation
Background and Motivation
Modern machine learning systems are trained on mixtures of data from multiple domains, and choosing the right mixture is a central design choice [GD+24, ZY+23]. Despite extensive empirical work on data mixing—including mixture weight transfer [XPD+23, FPJ24, LZM+25], budget-dependent optimization [KSW+25], and adaptive reweighting [CRB+23, CHL+25, JZF+25]—a fundamental question remains open: when does auxiliary data genuinely improve scaling laws rather than merely providing more samples?
Problem Setting
The paper studies high-dimensional regression with training domains sharing a common regression function . For domain :
where has rows drawn from distribution with and . The test distribution is a mixture of the domains.
The excess test risk for an estimator is:
Key Theoretical Framework
Mixed ridge estimator:
with .
The paper operates under Assumption 1 (concentration of designs), which covers sub-Gaussian coordinates and convex Lipschitz concentration [CM24, MS24]. The setting connects to kernel methods: with for a feature map , the model becomes kernel ridge regression with data mixture.
Methodology
Minimax Risk Characterization (Theorem 1)
Under an ellipsoidal parameter constraint , the minimax risk satisfies:
where (with ) is the information operator, and is the variational functional:
with .
Key insight: A heterogeneous dataset mixture is equivalent (at the minimax level) to a single Gaussian sequence model with information operator . Domain contributes information proportional to , weighted by its covariance in each direction.
For commuting covariances, the variational problem reduces to:
Deterministic Equivalent (Theorem 2)
For commutative covariances, the risk of mixed ridge regression is characterized by a fixed-point system:
This gives a deterministic proxy where each dataset is replaced by its population covariance scaled by an effective sample size .
Power-Law Model
Specialized to domains with aligned spectra:
with (auxiliary spectrum has heavier tails). The parameter class imposes a source condition with regularity parameter .
Empirical Validation / Results
Minimax Scaling Laws (Theorem 3)
Define the individual and mixed exponents:
The mixed minimax rate improves on both single-dataset rates if and only if and , where .
Three regimes emerge:
- Too few auxiliary samples (): mixing doesn't help—auxiliary data only aids coordinates already well-estimated
- Light auxiliary tails (): no improvement—variance dominated by high-frequency coordinates
- Heavy auxiliary tails (, ): mixing strictly improves the rate
Ridge Regression Scaling Laws (Theorem 4)
With truncation parameters and :
Corollary 1: Ridge regression achieves the mixed minimax rate whenever the mixed minimax rate strictly improves (δ > 1, γ_c < γ_2 < 1), even for where target-only ridge suffers from saturation.
| Regime | Mixed Risk Rate | Improvement |
|---|---|---|
| Target-only optimal | ||
| , δ > 1 | Auxiliary-only optimal | |
| , δ > 1 | Strict improvement |
Experimental Results
Linear models (Figure 1): Empirical risk of mixed ridge matches the deterministic equivalent and predicted scaling laws. For with δ > 1, the mixed rate () improves on both target-only () and auxiliary-only () rates.
Real image features (Figure 2): Deterministic equivalent predictions hold for CIFAR-10 and ImageNet-100 features extracted via pretrained encoders, beyond the technical assumptions.
Language models (Figure 3): An 81.5M-parameter GPT-2-style model trained on SlimPajama shows:
- : data mixture improves the scaling law
- or : no noticeable improvement
- Token-frequency analysis reveals the auxiliary domain places more mass on tokens infrequent in the target domain—a token-level analogue of the heavier-tailed auxiliary covariance
Theoretical and Practical Implications
Theoretical Significance
-
First characterization of when mixing helps: The paper resolves the open question of whether data mixtures improve scaling rates or merely provide more samples, identifying the precise interplay between spectral decay, target regularity, and relative sample sizes.
-
Positive distribution shift: At equal sample sizes (), auxiliary data with heavier tails is more informative about the target than target data itself—a formal demonstration of positive transfer.
-
Ridge regression is minimax optimal in the improvement regime: Saturation effects (Tikhonov regularization) are compensated by the auxiliary data when its spectral tail is sufficiently heavy.
Practical Implications
-
Data mixture design guidelines: Auxiliary data should have heavier-tailed feature/token distributions relative to the target, and its volume should grow at an intermediate rate relative to target data.
-
Budget allocation: The regime indicates the auxiliary budget should grow sub-linearly relative to target data—too little or too much auxiliary data both fail to improve scaling.
-
Language model pretraining: Token-level frequency distributions serve as a practical proxy for covariance spectra: mixing domains with complementary frequency coverage (heavier tails) at appropriate proportions yields measurable scaling improvements.
Conclusion
Main Takeaways
-
Data mixtures provably improve scaling laws when: (i) the auxiliary spectrum is heavy-tailed relative to the target (δ > 1), meaning the two datasets cover complementary spectral directions; and (ii) the auxiliary sample size grows at an intermediate rate ()
-
In this regime, ridge regression attains the minimax optimal rate, achieving a scaling exponent strictly better than either dataset alone
-
The qualitative predictions hold for real language models, where token-frequency distributions serve as a practical analogue of covariance spectra
Future Directions
-
Gradient descent dynamics: Determining whether gradient descent methods attain the mixed minimax rates when ridge saturates (building on [WBK+26])
-
Random-feature models: Identifying how model size must scale with data budget to preserve mixing improvements [DLM24, BAP24]
-
Feature learning: Extending to features trained via gradient descent steps [BES+22, CPD+24]
-
Joint optimization: Optimizing domain proportions under combined data and compute constraints, and characterizing mixture transfer across model scales [GWB26]
The authors note open technical questions: extending the deterministic equivalent to non-commutative covariances (conjectured to hold) and achieving multiplicative (rather than additive) error guarantees in Theorem 2.
Related papers
- Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.
- Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Frozen weights do not guarantee safe saturation in agentic coding; auditors must monitor scaffold expansion, not just checkpoint freezing.