# ScAn-Bench: Evaluating Scaling Analysis Methodology

> ScAn-Bench introduces the first surrogate benchmarks for scaling analysis, revealing no universally optimal data acquisition and extrapolation strategy exists across LLMs and VLMs.

- **Source:** [arXiv](https://arxiv.org/abs/2609.35707)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/dzUnhy
- **Whiteboard:** https://picx.dev/p/dzUnhy/image

## Summary

# ScAn-Bench: Evaluating Scaling Analysis Methodology

## Summary (Overview)

- **Introduction of novel surrogate benchmarks**: The paper introduces ScAn-Bench-VLM and ScAn-Bench-LLM, the first surrogate benchmarks specifically designed for evaluating scaling analysis (ScAn) methodology across vision-language models (VLMs) and large language models (LLMs), based on 8,024 and 4,524 checkpoints respectively.

- **Systematic evaluation framework**: The authors develop a unified framework decomposing scaling analysis into data acquisition (sampling configurations under a constrained search budget) and extrapolation (predicting optimal behavior at unseen compute scales), with three levels of extrapolation granularity: loss extrapolation, coarse parameter extrapolation, and full parameter extrapolation.

- **Key empirical findings**: No single data acquisition and extrapolation combination is optimal across all benchmarks. CARBS excels at early exploitation but suffers from localized sampling, while deterministic grid-based approaches (Chinchilla A1/A2) provide broader Pareto frontier coverage. The parametric Chinchilla approach fails for VLM model size prediction due to invalid compute estimation heuristics ($C \approx 6ND$), while the empirical Power-Law fit remains robust across domains.

- **Computational efficiency**: Surrogate querying takes only 29 CPU-seconds (VLM) and 17 CPU-seconds (LLM) for 100 evaluations, compared to 419.58 and 3,519.42 GPU-hours respectively for actual training, making ScAn research accessible to labs without large-scale GPU clusters.

- **Open-source availability**: The benchmarks and evaluation code are publicly released at https://github.com/automl/scan_bench, enabling reproducible research on scaling analysis methodology.

---

## Introduction and Theoretical Foundation

### Background and Motivation

The rapid advancement of foundation models has made training at target scales prohibitively expensive. Researchers increasingly rely on **scaling analysis (ScAn)** to model empirical relationships between architecture, data, hyperparameters, compute, and loss, enabling:

- Extrapolation of performance to larger scales
- Prediction of optimal training configurations
- Derivation of scaling prescriptions

However, the authors identify a critical blind spot: **no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types**. Existing evaluations are:

1. **Overwhelmingly focused on LLMs**, neglecting other foundation model families
2. **Suffering from empirical biases** due to inconsistent meta-choices during model training
3. **Largely studying static extrapolation** while ignoring interactive data acquisition strategies

### Theoretical Foundation: The Scaling Problem

The paper formalizes the scaling problem as follows. Given a pretraining task loss $f$ to be minimized at a target compute budget $C_{\text{target}}$ (typically FLOPs), with a joint hyperparameter search space $\Lambda = \Lambda_c \times \Lambda_o$ where:

- $\Lambda_c$: space of compute-scaling hyperparameters
- $\Lambda_o$: space of non-scaling optimization hyperparameters

The theoretical goal is to find the optimal configuration $(\lambda^*_c, \lambda^*_o)$ that minimizes the objective subject to the compute constraint:

$$(\lambda^*_c, \lambda^*_o) = \arg\min_{(\lambda_c, \lambda_o) \in \Lambda} f(\lambda_c, \lambda_o) \quad \text{subject to} \quad C(\lambda_c) \leq C_{\text{target}}$$

The compute budget is strictly a function of the compute-scaling hyperparameters, $C(\lambda_c)$.

### Three Levels of Extrapolation

The framework decomposes extrapolation into three granularity levels:

1. **Loss extrapolation**: Predicts minimal achievable loss for a given $C_{\text{target}}$ by extrapolating $f$ along the empirically observed Pareto optimal frontier.

2. **Coarse parameter extrapolation**: Treats model size $N$ as a scalar proxy for scale, predicting optimal model size $N^*$ and consequently optimal token count $D^*$.

3. **Full extrapolation**: Directly models the optimal joint set of compute-scaling and non-scaling hyperparameters $(\hat{\lambda}^*_c, \hat{\lambda}^*_o)$ as a function of $C_{\text{target}}$, eliminating reliance on heuristic priors.

### Benchmarking Inspiration

The authors draw inspiration from **neural architecture search (NAS)** and **hyperparameter optimization (HPO)** fields, where tabular and surrogate benchmarks (e.g., NAS-Bench-101) have become key enabling research artifacts. Surrogate benchmarks fit a fast predictor over collected training data to approximate expensive training runs in seconds rather than GPU-hours.

---

## Methodology

### ScAn Spaces and Data Collection

#### VLM Pipeline
- Trains **CLIP models** using an open-source implementation on a 60M-sample subset of LAION-400M
- Scaling variables: vision encoder width, text encoder width, number of training samples
- Optimization hyperparameters: learning rate, weight decay, warmup fraction, optimizer coefficients ($\beta_1$, $\beta_2$, $\epsilon$)

#### LLM Pipeline
- Trains **decoder-only transformers** on the SlimPajama dataset
- Scaling variables: number of layers, number of attention heads, embedding dimension, training tokens
- Optimization hyperparameters: learning rate, weight decay, cooldown steps, optimizer coefficients ($\beta_1$, $\beta_2$)

**Table 2: Benchmark Summary**

| Benchmark | #HPs | #Scale | FLOP Range | Cont. | Disc. | #Down. |
|-----------|------|--------|------------|-------|-------|--------|
| ScAn-Bench-VLM | 6 | 3 | $3.6 \times 10^{13}$ – $7.8 \times 10^{18}$ | 7 | 2 | 40 |
| ScAn-Bench-LLM | 5 | 4 | $3.2 \times 10^{16}$ – $2.9 \times 10^{20}$ | 6 | 3 | - |

A controlled checkpointing scheme records training dynamics, with 8,024 checkpoints for VLMs and 4,524 for LLMs, evaluating both upstream and downstream performance at each checkpoint.

### Surrogate Modeling

The authors evaluate multiple surrogate candidates:
- **AutoGluon**
- **XGBoost**
- **LightGBM**
- **TabPFN** (selected as best)
- Mixed ensemble (Linear Regression, Ridge Regression, Random Forests, XGBoost, LightGBM)

**Table 3: Surrogate Performance Comparison**

| Surrogate | VLM Spearman ↑ | VLM RMSE ↓ | LLM Spearman ↑ | LLM RMSE ↓ |
|-----------|---------------|------------|----------------|------------|
| **TabPFN** | **0.98** | **0.21** | **0.95** | **0.27** |
| AutoGluon | 0.98 | 0.26 | 0.91 | 0.30 |
| Ensemble (XGB) | 0.96 | 0.35 | 0.85 | 0.35 |
| Ensemble (Mix) | 0.95 | 0.37 | 0.86 | 0.33 |
| Ensemble (LGB) | 0.96 | 0.42 | 0.87 | 0.33 |

**TabPFN** consistently yields the strongest performance and is selected as the default surrogate predictor. A separate binary surrogate predicts whether a VLM configuration will diverge.

**Table 1: Runtime Comparison (100 evaluations)**

| Model family | Surrogate query (CPU-s) | Training (GPU-h) |
|--------------|------------------------|------------------|
| VLM | 29 | 419.58 |
| LLM | 17 | 3,519.42 |

### Evaluated ScAn Methodology

#### Data Acquisition Strategies
Four generalized heuristic-free approaches are evaluated:

1. **Iso-Parameter Profiling (Chinchilla Approach 1)**: Deterministic grid-based sampling at fixed model sizes
2. **Iso-FLOP Profiling (Chinchilla Approach 2)**: Deterministic grid-based sampling at fixed compute budgets
3. **CARBS (Cost-Aware Robust Bayesian Search)**: Dynamic sequential sampling using adapted Expected Improvement acquisition function
4. **Random Search**: Baseline for isolating structured allocation strategy gains

#### Extrapolation Techniques
- **Loss extrapolation**: Kaplan-style power law fit vs. Chinchilla-style parametric loss function
- **Coarse parameter extrapolation**: Chinchilla parametric loss estimate under cost constraints vs. direct power law fitting between macro-parameters and total compute
- **Full extrapolation**: Independent linear regression per hyperparameter dimension (as in CARBS)

### Experimental Setup

- **Compute constraints**: $C_{\text{max}} = 2 \times 10^{16}$ FLOPs for OpenCLIP, $C_{\text{max}} = 1.0 \times 10^{19}$ FLOPs for LLM
- **Target budget**: $C_{\text{target}} = 20 \times C_{\text{max}}$ (anchored within benchmark compute scale for ground-truth verification)
- **Loss function**: Pretraining validation loss (monotonically decreasing, well-suited for power-law fitting)
- **Seed aggregation**: 10 independent seeds, reporting mean performance with standard error

---

## Empirical Validation / Results

### Research Questions Addressed

- **RQ1**: Trade-offs during data acquisition among maximizing immediate performance, exploring Pareto frontier, and ensuring accurate scaling law fit
- **RQ2**: Consistent loss extrapolation + data acquisition combinations across benchmarks
- **RQ3**: Consistent coarse parameter extrapolation + data acquisition combinations
- **RQ4**: Reliability of full extrapolation approaches
- **RQ5**: Systematic preferences of acquisition approaches for extrapolation techniques

### Data Acquisition Results

A clear **exploration-exploitation trade-off** emerges (Figure 1):

- **CARBS** efficiently exploits high-performing regions, achieving the lowest **Pareto Estimation Regret (PER)** and superior early-stage **Incumbent Loss**
- **Iso-Parameter (Chinchilla A1)** ensures broad exploration, yielding the highest **Hypervolume**
- **Random Search** surpasses CARBS in best loss found as compute budget expands, suggesting CARBS's exploitation bias restricts its search space during later acquisition phases

Figure 2 reveals that CARBS-acquired data points exhibit a **highly localized and narrow distribution**, yet the resulting scaling trajectory closely mirrors the optimal empirical trend.

### Loss Extrapolation Results

**Under Iso-Parameter Profiling (Figure 3)**:
- The parametric **Chinchilla Approach 3** achieves closer predictions to the empirical target loss
- **Kaplan** exhibits more stable, monotonic improvement with increasing compute

**Under CARBS acquisition (Figure 4)**:
- **Chinchilla Approach 3** significantly outperforms Kaplan in anytime prediction on OpenCLIP
- **Kaplan** yields more reliable and accurate predictions on the LLM benchmark

**Key finding**: No single data acquisition + extrapolation combination is optimal across all benchmarks.

### Coarse Parameter Extrapolation Results

The parametric Chinchilla approach demonstrates **fundamental brittleness** when compute assumptions are violated:

- The strict reliance on $C \approx 6ND$ compute estimation causes **failure in predicting optimal OpenCLIP model size** (Figure 5a)
- Both Chinchilla and empirical Power-Law projection methods **reliably predict $N^*$ for the standard LLM benchmark** (Figure 5b)
- The **purely empirical Power-Law Fit** is a more robust and generalized estimator across diverse architectural domains

**Figure 6 analysis**: Except for CARBS (which yields poor predictions due to localized exploitation), most profiling strategies facilitate accurate $N^*$ predictions at the given search budget.

### Full Parameter Extrapolation Results

**Table 4: Actual Loss at Predicted Configurations**

| Strategy | LLM 0.1x | LLM 0.5x | LLM Best | OpenCLIP 0.1x | OpenCLIP 0.5x | OpenCLIP Best |
|----------|----------|----------|----------|---------------|---------------|---------------|
| CARBS | $2.02 \pm 0.23$ | $1.76 \pm 0.17$ | 1.31 | $2.92 \pm 0.41$ | $2.49 \pm 0.14$ | 0.41 |
| RS | $1.71 \pm 0.18$ | $1.89 \pm 0.37$ | 1.31 | $2.28 \pm 0.29$ | $1.81 \pm 0.30$ | 0.41 |

Key observations:
- **Random Search outperforms CARBS** across all evaluated budgets
- Random Search predictions **monotonically improve** as more data is acquired
- The approach is **highly sensitive to the exact composition of the empirical Pareto frontier**
- The naive method assumes **dimensional independence**, ignoring coupled scaling correlations between hyperparameters (e.g., OpenCLIP text and image tower interactions)

### Interaction Between Data Acquisition and Prediction

The parametric model (Chinchilla 3) consistently outperforms empirical Power-Law fits for **loss extrapolation** due to its highly flexible five-parameter formulation. However, this flexibility introduces **severe identifiability issues**:

> "Because the parametric derivation of $N^*$ strictly depends on the precise ratio of specific fitted coefficients, it is highly brittle under poor identifiability."

The efficacy of $N^*$ prediction is dictated by whether the data acquisition phase explicitly traces a usable compute-optimal frontier:

- **Structured profiling methods** (Chinchilla A1/A2) map the envelope → Power-Law approach is highly accurate
- **Unstructured methods** (Random Search) scatter $(N, D)$ allocations without targeting the envelope → Power-Law ineffective
- **Exploitative methods** (CARBS) sample too narrowly → degenerate traces fail to map the Pareto frontier

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Unified framework for ScAn**: The paper provides the first systematic decomposition of scaling analysis into distinct methodological components (data acquisition, loss extrapolation, coarse parameter extrapolation, full extrapolation), enabling principled comparison.

2. **Identification of methodological brittleness**: Demonstrates that the standard $C \approx 6ND$ compute estimation fails for multi-modal architectures, highlighting the need for architecture-aware compute models.

3. **Exploration-exploitation trade-off formalization**: Quantifies the trade-off between exploitation-focused strategies (CARBS) and exploration-focused strategies (deterministic profiling) in the context of scaling law derivation.

4. **Identifiability vs. flexibility tension**: Reveals that flexible parametric models (Chinchilla 3) may fit loss surfaces well while failing to identify true underlying parameters needed for $N^*$ prediction.

### Practical Implications

1. **Democratization of ScAn research**: The surrogate benchmarks reduce evaluation cost from GPU-hours to CPU-seconds (orders of magnitude faster), enabling researchers without large-scale GPU clusters to advance ScAn methodology.

2. **Method selection guidance**: Provides evidence that:
   - Empirical Power-Law fits are more robust than parametric approaches for cross-architecture model size prediction
   - Random Search can outperform sophisticated Bayesian methods for full parameter extrapolation
   - No universal "best" acquisition strategy exists; the choice depends on the extrapolation goal

3. **Reproducibility**: The controlled evaluation framework decouples strategies from domain-specific heuristics, enabling apples-to-apples comparisons across model families.

4. **Template for future benchmarks**: The benchmark construction methodology serves as a template for creating additional benchmarks across other model families, architectures, and training recipes.

---

## Conclusion

### Main Takeaways

The paper introduces **ScAn-Bench-VLM** and **ScAn-Bench-LLM**, the first surrogate benchmarks for systematic evaluation of scaling analysis methodology. Through controlled evaluation, the authors demonstrate:

1. **Scaling law projections are highly sensitive to underlying design choices** in both data acquisition and extrapolation phases
2. **No universally optimal combination exists** across benchmarks and model families
3. **Parametric approaches are brittle** when their underlying compute assumptions are violated
4. **Empirical approaches are more robust** but depend critically on data acquisition strategies properly tracing the Pareto frontier

### Future Directions

The authors identify several limitations and opportunities:

- **Scale limitation**: Current compute scales remain substantially smaller than state-of-the-art models; observed trends may not fully transfer to significantly larger regimes
- **Hyperparameter sensitivity**: The sensitivity of ScAn methods to their own hyperparameters is not evaluated
- **Cross-architecture generalization**: Results do not always align across different architectures and training recipes, highlighting the risk of drawing overly general conclusions from single pipelines
- **Downstream task extension**: The evaluation is restricted to upstream scaling behavior, though downstream task evaluation is supported by the surrogate benchmarks

The authors hope ScAn-Bench enables the community to develop and evaluate new ScAn methodology in a more systematic and reproducible manner, and that the benchmarks serve as a template encouraging the creation of additional benchmarks across diverse model families.

---

_Markdown view of https://picx.dev/p/dzUnhy, served by PicX — AI-generated visual whiteboard summaries of research papers._
