Summary of "Fault-tolerant foundation models"

Summary (Overview)

  • Core finding: Large language models (LLMs) trained with the same hardware faults they will encounter at inference become increasingly error-resilient as they scale, rather than degrading—a behavior opposite to conventionally trained models.
  • Key contribution: The authors derive modified neural scaling laws (from ~40,000 GPU-hours of training on simulated faulty hardware) that include a new curvature coefficient α2\alpha_2, which drives capacity recovery beyond a critical model size N∗N^*.
  • Theoretical conjecture: The scaling behavior suggests fault-hardened models learn to compute within "good" error-correcting codes (analogous to the brain's grid cell code), implying they may be formally fault-tolerant—retaining finite capacity as N→∞N \to \infty for fault rates below a threshold pthp_{\text{th}}.
  • Practical implication: If the conjecture holds, running AI inference on low-energy, faulty hardware (e.g., Low Energy Neural Network Accelerators, or LENNAs) could yield orders-of-magnitude energy savings over current perfectly-reliable hardware.
  • Contrast with fault-blind models: Models trained without faults are catastrophically fragile under inference-time errors, exhibiting divergent scaling laws, logit "heating" (temperature rise), and irreversible logit scrambling.

Introduction and Theoretical Foundation

Background and Motivation

Modern AI systems demand perfectly reliable hardware, which carries an energy penalty: chips suppress variability between logic gates by operating at high voltage, and energy scales with voltage squared. This energy–reliability tradeoff is fundamental to the thermodynamics of computation.

Key motivation: Biology demonstrates this reliance on pristine components is not fundamental. The brain—built from heterogeneous, noisy neurons—leverages fault-tolerant circuit architectures to perform complex reasoning with orders of magnitude less energy than artificial neural networks. Grid cells, for instance, compute within an error-correcting code: position is represented redundantly across modules of different spatial periods, so noise in any one module can be corrected.

The Problem with Current Models

Quantization (rounding weights to fewer bits) already degrades commercial LLMs significantly. Real low-energy hardware presents a far harder challenge: random errors of arbitrary magnitude. This fragility stems partly from "super weights"—outsized parameters critical to model function that act as single points of failure.

Theoretical Framework

The authors draw on two key theoretical ideas:

  1. Coding theory: Nature protects information in two ways:

    • Repetition codes: Keep many copies (crude, overhead grows with scale).
    • "Good" codes: Distributed redundancy where the overhead stays a fixed fraction regardless of scale (e.g., grid cells).
  2. Error-correction-enhanced evolvability [39]: Adaptive processes (learning or evolution) are drawn toward fault-tolerant solutions because fault tolerance allows more configurations to be explored without catastrophic damage.

Key Definitions

The paper measures computational capacity via the equivalent fault-free model size η\eta:

L(0,η,D)=L(p,N,D)(1)L(0, \eta, D) = L(p, N, D) \tag{1}

where L(p,N,D)L(p, N, D) is the loss of a model with NN parameters, fault rate pp, trained on DD tokens. The relative capacity η/N\eta/N is the neural network analog of a code's rate: the fraction of parameters doing useful work rather than absorbing errors.

The central conjecture is formal fault tolerance:

lim⁡N,D→∞ηN≥c,∀p<pth(3)\lim_{N, D \to \infty} \frac{\eta}{N} \geq c, \quad \forall p < p_{\text{th}} \tag{3}

for pth>0p_{\text{th}} > 0 and 0<c≤10 < c \leq 1, where pthp_{\text{th}} is the fault-tolerance threshold (analogous to Eigen's error threshold for replicating genomes).


Methodology

Dataset

  • Training data: 350-billion-token slice of FineWeb (web text corpus), tokenized with an 8192-token byte-pair-encoding vocabulary, packed into 1024-token sequences.
  • Validation: Held-out split of 2 billion tokens from the same web crawls, with near-duplicates removed.

Model Ladder

All models are Llama2-style decoder-only transformers (rotary position embeddings, RMSNorm, SwiGLU feed-forward blocks). Key architectural choices:

  • Attention head dimension: 32 at every rung
  • Width-to-depth ratio: held near 32
  • Context length: 1024 tokens
  • Sixteen logarithmically spaced (width, depth) pairs ranging from 7.9×1067.9 \times 10^6 to 9.3×1089.3 \times 10^8 total parameters

One architecture (N=5.0×107N = 5.0 \times 10^7) was additionally trained at eight token budgets, eight random seeds per condition, and six training fault rates.

Fault Injection Model

The error model emulates a Low Energy Neural Network Accelerator (LENNA)—a chip that runs arithmetic at voltages too low for perfect reliability but detects its own errors:

  • Hardware mechanism: Operands carry a redundant residue (remainder modulo small bases) in a Redundant Residue Number System (RRNS). After every k=4k = 4 fused multiply–add (FMA) cells, a cheap modulus check tests the running sum for consistency.
  • Error conversion: A failed check discards those kk products and carries the previous partial sum forward—equivalent to zeroing a block of kk elements in one operand.
  • Fault application: Random blocks of k=4k = 4 matrix elements within attention and feed-forward operations are zeroed with probability pp (17 values between 0 and 0.2). Fresh masks are drawn for every sequence on every forward pass.
  • Scope: Faults target matrix multiplications (which dominate >99% of FLOPs for large models); embeddings, norms, nonlinearities, and residual additions remain fault-free.

Training

  • Token budgets ranged from 7.9×1077.9 \times 10^7 to 3.5×10103.5 \times 10^{10} tokens (fixed multiple of parameter count).
  • The two axes (size and duration) were not fully crossed: ten architectures at or below 1.1×1081.1 \times 10^8 parameters were trained at 10, 20, 80, and 320 tokens per parameter; six larger ones only at 20 and below.
  • Every combination of size, duration, and fault rate was trained at two random seeds.
  • "Conservative" hyperparameters from a standard reference implementation were used to maintain clean scaling-law behavior.

Scaling-Law Fitting

Equation 4 was fit to each fault-rate cohort separately on log-loss with a Huber loss, from multiple starting points. Confidence intervals came from 1,500 bootstrap resamples per cohort.


Empirical Validation / Results

Modified Scaling Laws for Fault-Hardened Models

The central empirical finding is that fault-hardened transformers obey a modified scaling law with a new curvature coefficient α2\alpha_2:

L(p,N,D)=E(p)+LN+LDLN(p,N)=eα0(p)−α1(p)n−α2(p)n2LD(p,D)=eβ0(p)−β1(p)dn=ln⁡NN0,d=ln⁡DD0(4)\begin{array}{c} L(p, N, D) = E(p) + L_N + L_D \\ L_N(p, N) = e^{\alpha_0(p) - \alpha_1(p) n - \alpha_2(p) n^2} \\ L_D(p, D) = e^{\beta_0(p) - \beta_1(p) d} \\ n = \ln \frac{N}{N_0}, \quad d = \ln \frac{D}{D_0} \end{array} \tag{4}

where EE is the irreducible loss floor, αk\alpha_k the model-size coefficients, βk\beta_k the data coefficients, and N0=D0=4×107N_0 = D_0 = 4 \times 10^7.

Key coefficient behaviors (Fig. 2):

  • α1\alpha_1 (typical model-size coefficient) shrinks as pp increases.
  • α2\alpha_2 (curvature term, pinned to zero for p=0p = 0) grows with pp—this drives capacity recovery beyond N∗N^*.
  • β1\beta_1 (data coefficient) decreases with pp, meaning models become more expensive to train as fault rates rise.

Capacity Recovery

The relative capacity η/N\eta/N follows a parabola in log model size:

ln⁡ηN=−c0(p)−c1(p)n+c2(p)n2−A(p,N)n=ln⁡NN0,N∗=N0exp⁡(c1(p)2c2(p))(2)\begin{array}{c} \ln \frac{\eta}{N} = -c_0(p) - c_1(p) n + c_2(p) n^2 - \mathcal{A}(p, N) \\ n = \ln \frac{N}{N_0}, \quad N^* = N_0 \exp\left(\frac{c_1(p)}{2 c_2(p)}\right) \end{array} \tag{2}

Capacity first falls with size, then recovers beyond a critical size N∗N^*—behavior impossible for repetition-based error correction, strongly suggesting models learn "good" codes.

The coefficients ckc_k relate to the scaling-law coefficients via:

c0=α0(p)−α0(0)α1(0),c1=1−α1(p)α1(0),c2=α2(p)α1(0)A(p,N)=1α1(0)ln⁡(Δ(p)LN(p,N)+1)Δ(p)=E(p)−E(0)(11)\begin{array}{c} c_0 = \frac{\alpha_0(p) - \alpha_0(0)}{\alpha_1(0)}, \quad c_1 = 1 - \frac{\alpha_1(p)}{\alpha_1(0)}, \quad c_2 = \frac{\alpha_2(p)}{\alpha_1(0)} \\ \mathcal{A}(p, N) = \frac{1}{\alpha_1(0)} \ln\left(\frac{\Delta(p)}{L_N(p, N)} + 1\right) \\ \Delta(p) = E(p) - E(0) \end{array} \tag{11}

Two Regimes of Fault-Hardened Scaling

The pp-dependence of αk\alpha_k hints at two regimes, reminiscent of crossing a fault-tolerance threshold:

  • Small pp: Decreases in α1\alpha_1 are mirrored by increases in α2\alpha_2 (models learn to correct increasingly severe errors).
  • Large pp: α2\alpha_2 plateaus while α1\alpha_1 continues to shrink (models become overwhelmed).

Fault-Blind Models Collapse Under Inference-Time Errors

In stark contrast, fault-blind models (ptrain=0p_{\text{train}} = 0, peval>0p_{\text{eval}} > 0) exhibit:

  1. Divergent scaling: Performance can actually decrease with training data DD past a threshold, lacking convergent scaling laws altogether (Fig. 3a).
  2. Logit heating: The effective temperature τ\tau of their output distribution rises rapidly with pevalp_{\text{eval}}:
P(xt=k∣xt−1=kt−1,…,x0=k0)=eak/τ∑jeaj/τ(7)\mathbb{P}(x_t = k | x_{t-1} = k_{t-1}, \dots, x_0 = k_0) = \frac{e^{a_k / \tau}}{\sum_j e^{a_j / \tau}} \tag{7}
  1. Irreversible logit scrambling: Even after removing temperature effects, fault-blind models lose capacity rapidly. The KL-divergence decomposes as:
D(Q∣P)=D(Q∣P∗)+D(P∗∣P)(8)D(Q|P) = D(Q|P^*) + D(P^*|P) \tag{8}

where P∗P^* is the nearest temperature-shifted distribution to QQ. Fault-hardened models lose almost no capacity to the residual (unrescalable) term until pevalp_{\text{eval}} exceeds ptrainp_{\text{train}}, while fault-blind models deteriorate immediately.

Training Cost of Robustness

Robustness is not free. The effective training budget δ\delta of a faulty model (tokens needed to match a fault-free model trained on δ\delta tokens) follows:

ln⁡Dδ=β0(p)−β0(0)β1(0)+(1−β1(p)β1(0))d(6)\ln \frac{D}{\delta} = \frac{\beta_0(p) - \beta_0(0)}{\beta_1(0)} + \left(1 - \frac{\beta_1(p)}{\beta_1(0)}\right) d \tag{6}

Since β1(p)\beta_1(p) decreases with pp, the second term's coefficient is positive, meaning the cost grows with model size.


Theoretical and Practical Implications

Theoretical Significance

  1. Evidence for "good" codes: The capacity recovery beyond N∗N^* strongly suggests models learn distributed, grid-cell-like codes rather than repetition schemes. This is the largest-scale evidence to date for error-correction-enhanced evolvability.

  2. Formal fault tolerance conjecture: The authors conjecture fault-hardened models may be formally fault-tolerant (Eq. 3). Two asymptotic behaviors are possible, with pp controlling a smooth transition:

    • Collapse (if floor gap Δ≫LN\Delta \gg L_N at scale): ln⁡(η/N)∝−n\ln(\eta/N) \propto -n (Eq. 9)
    • Saturation (if η/N→1\eta/N \to 1): A(p,N)≈Δ(p)α1(0)LN(p,N)\mathcal{A}(p, N) \approx \frac{\Delta(p)}{\alpha_1(0) L_N(p, N)} (Eq. 10)
  3. Hardware-software codesign: The LENNA architecture demonstrates that error detection (cheap in RRNS arithmetic) can be delegated to hardware while error correction is learned by the model—echoing the brain's division of labor between grid modules and the grid–hippocampal loop.

Practical Implications

  • Energy efficiency: If fault-tolerant models can run on low-energy faulty hardware, inference could be orders of magnitude more energy efficient than current systems.
  • Hardware roadmap: LENNAs add only a few percent overhead for error detection (a 3-bit residue on 12-bit FMAs), making them a practical near-term path to energy-efficient AI.
  • Training cost: Robustness requires more training tokens, but the cost is a fixed multiplier that may be acceptable given inference energy savings.

Conclusion

Main Takeaways

  1. Fault-hardened LLMs (trained with the same faults they encounter at inference) become more error-resilient as they scale—opposite to fault-blind models, which are catastrophically fragile.
  2. The modified scaling law with curvature term α2\alpha_2 explains this recovery and suggests models learn "good" error-correcting codes.
  3. The results support a formal conjecture of fault tolerance (Eq. 3), though experiments at N≈109N \approx 10^9 could not resolve the asymptotic regime.

Future Directions

  • Large-scale validation: Establishing Eq. 3 at commercial scales (~101110^{11} parameters) would cost 102610^{26} FLOPs (10810^8 GPU-hours, similar to Llama-3 405B's total budget)—intractable within academia but suggested to be worth doing.
  • De-risking experiments: Smaller experiments on natural or synthetic datasets with tunable irreducible floors EE could serve as "model organisms" to test whether fault tolerance persists as EE increases.
  • Mechanistic analysis: The trained models (multiple terabytes of weights, publicly available) could be studied to understand how fault-hardened representations differ from fault-blind ones—paralleling how networks trained to navigate spontaneously develop grid-cell-like codes.

Closing Perspective

The brain is the existence proof that fault tolerance, energy efficiency, and intelligence are compatible. The principle of error-correction-enhanced evolvability predicts fault tolerance is an expected outcome of adaptation in noisy environments. This work provides the largest-scale evidence to date that AI trained under the constraints that shaped nervous systems may come to share their robustness, potentially enabling orders-of-magnitude energy savings on future low-energy hardware.

Related papers