Full text not available for this paper
ScAn-Bench: Evaluating Scaling Analysis Methodology
Summary (Overview)
-
Introduction of novel surrogate benchmarks: The paper introduces ScAn-Bench-VLM and ScAn-Bench-LLM, the first surrogate benchmarks specifically designed for evaluating scaling analysis (ScAn) methodology across vision-language models (VLMs) and large language models (LLMs), based on 8,024 and 4,524 checkpoints respectively.
-
Systematic evaluation framework: The authors develop a unified framework decomposing scaling analysis into data acquisition (sampling configurations under a constrained search budget) and extrapolation (predicting optimal behavior at unseen compute scales), with three levels of extrapolation granularity: loss extrapolation, coarse parameter extrapolation, and full parameter extrapolation.
-
Key empirical findings: No single data acquisition and extrapolation combination is optimal across all benchmarks. CARBS excels at early exploitation but suffers from localized sampling, while deterministic grid-based approaches (Chinchilla A1/A2) provide broader Pareto frontier coverage. The parametric Chinchilla approach fails for VLM model size prediction due to invalid compute estimation heuristics (), while the empirical Power-Law fit remains robust across domains.
-
Computational efficiency: Surrogate querying takes only 29 CPU-seconds (VLM) and 17 CPU-seconds (LLM) for 100 evaluations, compared to 419.58 and 3,519.42 GPU-hours respectively for actual training, making ScAn research accessible to labs without large-scale GPU clusters.
-
Open-source availability: The benchmarks and evaluation code are publicly released at https://github.com/automl/scan_bench, enabling reproducible research on scaling analysis methodology.
Introduction and Theoretical Foundation
Background and Motivation
The rapid advancement of foundation models has made training at target scales prohibitively expensive. Researchers increasingly rely on scaling analysis (ScAn) to model empirical relationships between architecture, data, hyperparameters, compute, and loss, enabling:
- Extrapolation of performance to larger scales
- Prediction of optimal training configurations
- Derivation of scaling prescriptions
However, the authors identify a critical blind spot: no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. Existing evaluations are:
- Overwhelmingly focused on LLMs, neglecting other foundation model families
- Suffering from empirical biases due to inconsistent meta-choices during model training
- Largely studying static extrapolation while ignoring interactive data acquisition strategies
Theoretical Foundation: The Scaling Problem
The paper formalizes the scaling problem as follows. Given a pretraining task loss to be minimized at a target compute budget (typically FLOPs), with a joint hyperparameter search space where:
- : space of compute-scaling hyperparameters
- : space of non-scaling optimization hyperparameters
The theoretical goal is to find the optimal configuration that minimizes the objective subject to the compute constraint:
The compute budget is strictly a function of the compute-scaling hyperparameters, .
Three Levels of Extrapolation
The framework decomposes extrapolation into three granularity levels:
-
Loss extrapolation: Predicts minimal achievable loss for a given by extrapolating along the empirically observed Pareto optimal frontier.
-
Coarse parameter extrapolation: Treats model size as a scalar proxy for scale, predicting optimal model size and consequently optimal token count .
-
Full extrapolation: Directly models the optimal joint set of compute-scaling and non-scaling hyperparameters as a function of , eliminating reliance on heuristic priors.
Benchmarking Inspiration
The authors draw inspiration from neural architecture search (NAS) and hyperparameter optimization (HPO) fields, where tabular and surrogate benchmarks (e.g., NAS-Bench-101) have become key enabling research artifacts. Surrogate benchmarks fit a fast predictor over collected training data to approximate expensive training runs in seconds rather than GPU-hours.
Methodology
ScAn Spaces and Data Collection
VLM Pipeline
- Trains CLIP models using an open-source implementation on a 60M-sample subset of LAION-400M
- Scaling variables: vision encoder width, text encoder width, number of training samples
- Optimization hyperparameters: learning rate, weight decay, warmup fraction, optimizer coefficients (, , )
LLM Pipeline
- Trains decoder-only transformers on the SlimPajama dataset
- Scaling variables: number of layers, number of attention heads, embedding dimension, training tokens
- Optimization hyperparameters: learning rate, weight decay, cooldown steps, optimizer coefficients (, )
Table 2: Benchmark Summary
| Benchmark | #HPs | #Scale | FLOP Range | Cont. | Disc. | #Down. |
|---|---|---|---|---|---|---|
| ScAn-Bench-VLM | 6 | 3 | – | 7 | 2 | 40 |
| ScAn-Bench-LLM | 5 | 4 | – | 6 | 3 | - |
A controlled checkpointing scheme records training dynamics, with 8,024 checkpoints for VLMs and 4,524 for LLMs, evaluating both upstream and downstream performance at each checkpoint.
Surrogate Modeling
The authors evaluate multiple surrogate candidates:
- AutoGluon
- XGBoost
- LightGBM
- TabPFN (selected as best)
- Mixed ensemble (Linear Regression, Ridge Regression, Random Forests, XGBoost, LightGBM)
Table 3: Surrogate Performance Comparison
| Surrogate | VLM Spearman ↑ | VLM RMSE ↓ | LLM Spearman ↑ | LLM RMSE ↓ |
|---|---|---|---|---|
| TabPFN | 0.98 | 0.21 | 0.95 | 0.27 |
| AutoGluon | 0.98 | 0.26 | 0.91 | 0.30 |
| Ensemble (XGB) | 0.96 | 0.35 | 0.85 | 0.35 |
| Ensemble (Mix) | 0.95 | 0.37 | 0.86 | 0.33 |
| Ensemble (LGB) | 0.96 | 0.42 | 0.87 | 0.33 |
TabPFN consistently yields the strongest performance and is selected as the default surrogate predictor. A separate binary surrogate predicts whether a VLM configuration will diverge.
Table 1: Runtime Comparison (100 evaluations)
| Model family | Surrogate query (CPU-s) | Training (GPU-h) |
|---|---|---|
| VLM | 29 | 419.58 |
| LLM | 17 | 3,519.42 |
Evaluated ScAn Methodology
Data Acquisition Strategies
Four generalized heuristic-free approaches are evaluated:
- Iso-Parameter Profiling (Chinchilla Approach 1): Deterministic grid-based sampling at fixed model sizes
- Iso-FLOP Profiling (Chinchilla Approach 2): Deterministic grid-based sampling at fixed compute budgets
- CARBS (Cost-Aware Robust Bayesian Search): Dynamic sequential sampling using adapted Expected Improvement acquisition function
- Random Search: Baseline for isolating structured allocation strategy gains
Extrapolation Techniques
- Loss extrapolation: Kaplan-style power law fit vs. Chinchilla-style parametric loss function
- Coarse parameter extrapolation: Chinchilla parametric loss estimate under cost constraints vs. direct power law fitting between macro-parameters and total compute
- Full extrapolation: Independent linear regression per hyperparameter dimension (as in CARBS)
Experimental Setup
- Compute constraints: FLOPs for OpenCLIP, FLOPs for LLM
- Target budget: (anchored within benchmark compute scale for ground-truth verification)
- Loss function: Pretraining validation loss (monotonically decreasing, well-suited for power-law fitting)
- Seed aggregation: 10 independent seeds, reporting mean performance with standard error
Empirical Validation / Results
Research Questions Addressed
- RQ1: Trade-offs during data acquisition among maximizing immediate performance, exploring Pareto frontier, and ensuring accurate scaling law fit
- RQ2: Consistent loss extrapolation + data acquisition combinations across benchmarks
- RQ3: Consistent coarse parameter extrapolation + data acquisition combinations
- RQ4: Reliability of full extrapolation approaches
- RQ5: Systematic preferences of acquisition approaches for extrapolation techniques
Data Acquisition Results
A clear exploration-exploitation trade-off emerges (Figure 1):
- CARBS efficiently exploits high-performing regions, achieving the lowest Pareto Estimation Regret (PER) and superior early-stage Incumbent Loss
- Iso-Parameter (Chinchilla A1) ensures broad exploration, yielding the highest Hypervolume
- Random Search surpasses CARBS in best loss found as compute budget expands, suggesting CARBS's exploitation bias restricts its search space during later acquisition phases
Figure 2 reveals that CARBS-acquired data points exhibit a highly localized and narrow distribution, yet the resulting scaling trajectory closely mirrors the optimal empirical trend.
Loss Extrapolation Results
Under Iso-Parameter Profiling (Figure 3):
- The parametric Chinchilla Approach 3 achieves closer predictions to the empirical target loss
- Kaplan exhibits more stable, monotonic improvement with increasing compute
Under CARBS acquisition (Figure 4):
- Chinchilla Approach 3 significantly outperforms Kaplan in anytime prediction on OpenCLIP
- Kaplan yields more reliable and accurate predictions on the LLM benchmark
Key finding: No single data acquisition + extrapolation combination is optimal across all benchmarks.
Coarse Parameter Extrapolation Results
The parametric Chinchilla approach demonstrates fundamental brittleness when compute assumptions are violated:
- The strict reliance on compute estimation causes failure in predicting optimal OpenCLIP model size (Figure 5a)
- Both Chinchilla and empirical Power-Law projection methods reliably predict for the standard LLM benchmark (Figure 5b)
- The purely empirical Power-Law Fit is a more robust and generalized estimator across diverse architectural domains
Figure 6 analysis: Except for CARBS (which yields poor predictions due to localized exploitation), most profiling strategies facilitate accurate predictions at the given search budget.
Full Parameter Extrapolation Results
Table 4: Actual Loss at Predicted Configurations
| Strategy | LLM 0.1x | LLM 0.5x | LLM Best | OpenCLIP 0.1x | OpenCLIP 0.5x | OpenCLIP Best |
|---|---|---|---|---|---|---|
| CARBS | 1.31 | 0.41 | ||||
| RS | 1.31 | 0.41 |
Key observations:
- Random Search outperforms CARBS across all evaluated budgets
- Random Search predictions monotonically improve as more data is acquired
- The approach is highly sensitive to the exact composition of the empirical Pareto frontier
- The naive method assumes dimensional independence, ignoring coupled scaling correlations between hyperparameters (e.g., OpenCLIP text and image tower interactions)
Interaction Between Data Acquisition and Prediction
The parametric model (Chinchilla 3) consistently outperforms empirical Power-Law fits for loss extrapolation due to its highly flexible five-parameter formulation. However, this flexibility introduces severe identifiability issues:
"Because the parametric derivation of strictly depends on the precise ratio of specific fitted coefficients, it is highly brittle under poor identifiability."
The efficacy of prediction is dictated by whether the data acquisition phase explicitly traces a usable compute-optimal frontier:
- Structured profiling methods (Chinchilla A1/A2) map the envelope → Power-Law approach is highly accurate
- Unstructured methods (Random Search) scatter allocations without targeting the envelope → Power-Law ineffective
- Exploitative methods (CARBS) sample too narrowly → degenerate traces fail to map the Pareto frontier
Theoretical and Practical Implications
Theoretical Contributions
-
Unified framework for ScAn: The paper provides the first systematic decomposition of scaling analysis into distinct methodological components (data acquisition, loss extrapolation, coarse parameter extrapolation, full extrapolation), enabling principled comparison.
-
Identification of methodological brittleness: Demonstrates that the standard compute estimation fails for multi-modal architectures, highlighting the need for architecture-aware compute models.
-
Exploration-exploitation trade-off formalization: Quantifies the trade-off between exploitation-focused strategies (CARBS) and exploration-focused strategies (deterministic profiling) in the context of scaling law derivation.
-
Identifiability vs. flexibility tension: Reveals that flexible parametric models (Chinchilla 3) may fit loss surfaces well while failing to identify true underlying parameters needed for prediction.
Practical Implications
-
Democratization of ScAn research: The surrogate benchmarks reduce evaluation cost from GPU-hours to CPU-seconds (orders of magnitude faster), enabling researchers without large-scale GPU clusters to advance ScAn methodology.
-
Method selection guidance: Provides evidence that:
- Empirical Power-Law fits are more robust than parametric approaches for cross-architecture model size prediction
- Random Search can outperform sophisticated Bayesian methods for full parameter extrapolation
- No universal "best" acquisition strategy exists; the choice depends on the extrapolation goal
-
Reproducibility: The controlled evaluation framework decouples strategies from domain-specific heuristics, enabling apples-to-apples comparisons across model families.
-
Template for future benchmarks: The benchmark construction methodology serves as a template for creating additional benchmarks across other model families, architectures, and training recipes.
Conclusion
Main Takeaways
The paper introduces ScAn-Bench-VLM and ScAn-Bench-LLM, the first surrogate benchmarks for systematic evaluation of scaling analysis methodology. Through controlled evaluation, the authors demonstrate:
- Scaling law projections are highly sensitive to underlying design choices in both data acquisition and extrapolation phases
- No universally optimal combination exists across benchmarks and model families
- Parametric approaches are brittle when their underlying compute assumptions are violated
- Empirical approaches are more robust but depend critically on data acquisition strategies properly tracing the Pareto frontier
Future Directions
The authors identify several limitations and opportunities:
- Scale limitation: Current compute scales remain substantially smaller than state-of-the-art models; observed trends may not fully transfer to significantly larger regimes
- Hyperparameter sensitivity: The sensitivity of ScAn methods to their own hyperparameters is not evaluated
- Cross-architecture generalization: Results do not always align across different architectures and training recipes, highlighting the risk of drawing overly general conclusions from single pipelines
- Downstream task extension: The evaluation is restricted to upstream scaling behavior, though downstream task evaluation is supported by the surrogate benchmarks
The authors hope ScAn-Bench enables the community to develop and evaluate new ScAn methodology in a more systematic and reproducible manner, and that the benchmarks serve as a template encouraging the creation of additional benchmarks across diverse model families.
Related papers
- When do data mixtures improve scaling laws? Insights from high-dimensional regression
Data mixtures provably accelerate scaling laws only when auxiliary data has heavier-tailed spectra and intermediate relative sample growth, with ridge regression achieving the optimal rate.
- LLM sequential decision making under uncertainty in biochemical domains
LLMs overreact to new data and fail to explore in Bayesian optimization due to context-stickiness, a competence gap, not an intention gap, fixable by removing in-context history.
- How Linear Attention Remembers
Linear attention memory stores facts via concentrated content-specific writes and retrieves them via focused query-time reads, with hybrid models shifting recall to full-attention KV caches.