Summary (Overview)
- This paper presents a systematic procedure for deriving robust compute-optimal scaling laws for high-energy physics (HEP) machine learning models, specifically for jet flavor tagging using transformer architectures on the ~11 billion-jet ATLAS JetSet2 dataset.
- The key methodological contribution is a "hyperparameter recipe" that makes the optimal learning rate and batch size predictable as closed-form functions of model size and token budget, using μP/Complete(d)P parameterization for width/depth transferability.
- The authors recover a near-equal dependence of model size and dataset size on compute (, ), matching the Chinchilla scaling result.
- Auxiliary multi-task objectives consistently lower the primary jet-classification loss at equal compute budget, while richer input representations (adding tracks, electrons, muons, particle-flow objects) systematically lower loss without changing the scaling exponent.
- The onset of the power-law scaling regime requires a minimum dataset size (a few hundred million jets); below this threshold, scaling fits carry little information about high-compute behavior.
Introduction and Theoretical Foundation
Jet flavor tagging—identifying jets originating from b-quarks, c-quarks, τ-leptons, and light quarks/gluons—is a central component of the LHC physics program. ATLAS has progressively moved toward more powerful neural-network-based taggers, from recurrent networks to graph and transformer models. The latest GN family uses transformers trained directly on low-level tracks and particle-flow constituents.
Neural scaling laws [24, 25] quantify the dependence of achievable loss on model size , dataset size , and compute through power-law relations. These laws can predict performance at larger training budgets and determine how a budget should be divided between and . The compute budget for transformers is:
where is the token budget (per-jet constituents), and is the number of jets.
The central problem: At fixed model size and shape, the optimal learning rate and batch size change systematically with the training budget. Fixing hyperparameters at a reference budget mixes genuine scaling behavior with the movement of the training optimum. A robust scaling analysis must account for this movement explicitly.
The scaling law framework: The loss decreases as a power law in , , and . The compute-optimal frontier is a two-level optimum:
The canonical "Chinchilla" parameterization of the full loss surface is:
with the compute-optimal allocation given by .
Methodology
Architecture: A transformer flavor tagger based on the ATLAS GN-family, ingesting three complementary input modalities: charged-particle tracks (with impact-parameter and quality features), particle-flow objects, and soft electron/muon information. The model uses attention pooling and multiple task-specific heads.
Training objectives: Five loss functions summed directly:
including jet classification (cross-entropy with class weights), vertexing (binary cross-entropy on track pairs), track-origin classification, track-type classification, and jet regression (L1 loss).
Hyperparameter recipe (Section 5):
- Remove model-size dependence: Use the Complete(d)P parameterization with per-parameter-type learning-rate, AdamW-ε, and weight-decay multipliers (Table 1), making width-invariant. A residual depth dependence is factored out:
with fitted .
- Joint fit: The optimal learning rate follows a separable power law:
-
Optimal batch size: For the multi-task objective, , while for classification-only the loss is flat in except in the few-step corner.
-
Optimal aspect ratio: —width should grow faster than depth.
Three fitting approaches for scaling laws:
- Training-curve envelope: Read the lower envelope of loss-vs-compute curves directly.
- IsoFLOP profiles: Sweep at fixed and locate the minimum.
- Full-surface fit: Fit of Eq. (5) jointly to all cells.
Data repetition model (Section 7): Early-stopped runs targeting passes over unique jets are modeled with a fresh-equivalent pass count:
with saturating at .
Empirical Validation / Results
Warm-up studies: On a teacher-student problem with known floor (), all three fitting approaches agree on the allocation. The compute-optimal envelope passes through three regimes: initial data-driven descent, a clean power-law scaling regime, and a tail approaching the floor. The full-surface fit (Approach 3) mis-fits across regimes and flattens onto a biased floor.
Compute-optimal allocation (Section 6): All three approaches agree on near-equal dependence:
| Method | () | () |
|---|---|---|
| This work (Approach 1) | 0.52 ± 0.01 | 0.48 ± 0.01 |
| This work (Approach 2) | 0.55 ± 0.02 | 0.45 ± 0.02 |
| This work (Approach 3) | 0.49 ± 0.03 | 0.51 ± 0.03 |
| Carpe Datum [42] | 0.10 | 0.90 |
| Kaplan et al. [24] | 0.73 | 0.27 |
| Hoffmann et al. [25] | 0.49 | 0.51 |
Onset of scaling: The scaling regime only sets in past ~a few hundred million jets. Below this threshold, the envelope is dominated by a single small model and fits return biased exponents.
Auxiliary tasks: At fixed compute, the multi-task variant achieves lower jet-classification loss than classification-only (Fig. 24), and shifts the compute-optimal allocation toward larger models and smaller datasets.
Input variants: Each added modality (track features, soft electrons, muons, particle-flow objects) lowers the compute-optimal loss while leaving the scaling exponent nearly unchanged.
Data repetition: The best training horizon grows with unique dataset size (5 passes on jets to 17 on for a 31.5M model). Model-wise double descent survives early stopping. At large compute, the allocation reverts to larger models targeting fewer passes.
Physics performance: Projections using translate loss into background rejection. At 70% b-jet efficiency, compute-optimal models projected to higher compute budgets show substantial gains over existing GN3 and TN25-86M models.
Theoretical and Practical Implications
- Methodological: Reliable scaling laws require training dynamics (learning rate, batch size) to be scaled together with model and dataset. Predictive hyperparameter transfer removes bias that would otherwise be absorbed into fitted scaling exponents and floors.
- Architecture comparison: The compute-optimal frontier provides a fair basis for comparing architectures, objectives, and input representations—each must be evaluated at its own optimal allocation rather than at fixed or .
- Practical guidance: The multi-task objective wins at scale despite higher per-jet cost; richer inputs are always beneficial; the scaling exponent is approximately invariant to input changes, so gains from better inputs and from scaling are multiplicative.
- Data requirements: The onset of scaling requires large, high-quality full-simulation datasets. Fits based on smaller datasets can return biased exponents even when curves appear well-described by power laws.
- Resource scale: The full study required ~97k A100 GPU-hours—substantial by HEP standards but modest compared to industrial-scale models. The observed scaling suggests larger budgets could be productively invested, particularly toward foundation models.
Conclusion
The paper presents a complete recipe for deriving robust compute-optimal scaling laws in HEP, demonstrated on multi-task jet-tagging transformers. Key findings:
- Near-equal scaling of model and dataset size (Chinchilla-like behavior)
- Auxiliary objectives improve primary task at scale and shift optimal allocation toward larger models
- Richer inputs lower loss without changing scaling exponents
- The power-law regime requires a minimum dataset scale to emerge
The framework serves as a practical blueprint for future scaling studies in HEP, suggesting a natural progression from small-scale studies to large-scale, full-simulation benchmarks, and ultimately toward generic foundation models pretrained at scale and fine-tuned for specific experiments, detector configurations, and physics tasks.
Related papers
- Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.
- iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
iS-KV compresses KV caches via block-incremental SVD, jointly updating basis and coordinates to retain all tokens, achieving near-original accuracy at 4-7x compression, outperforming eviction methods.