# Fault-tolerant foundation models

> Fault-hardened language models become more error-resilient as they scale, unlike fault-blind models, suggesting they learn good error-correcting codes that may enable energy-efficient inference on faulty hardware.

- **Source:** [arXiv](https://arxiv.org/abs/2610.10311)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/zpJb63
- **Whiteboard:** https://picx.dev/p/zpJb63/image

## Summary

# Summary of "Fault-tolerant foundation models"

## Summary (Overview)

- **Core finding**: Large language models (LLMs) trained with the same hardware faults they will encounter at inference become *increasingly error-resilient as they scale*, rather than degrading—a behavior opposite to conventionally trained models.
- **Key contribution**: The authors derive modified neural scaling laws (from ~40,000 GPU-hours of training on simulated faulty hardware) that include a new curvature coefficient $\alpha_2$, which drives capacity recovery beyond a critical model size $N^*$.
- **Theoretical conjecture**: The scaling behavior suggests fault-hardened models learn to compute within "good" error-correcting codes (analogous to the brain's grid cell code), implying they may be *formally fault-tolerant*—retaining finite capacity as $N \to \infty$ for fault rates below a threshold $p_{\text{th}}$.
- **Practical implication**: If the conjecture holds, running AI inference on low-energy, faulty hardware (e.g., Low Energy Neural Network Accelerators, or LENNAs) could yield orders-of-magnitude energy savings over current perfectly-reliable hardware.
- **Contrast with fault-blind models**: Models trained without faults are catastrophically fragile under inference-time errors, exhibiting divergent scaling laws, logit "heating" (temperature rise), and irreversible logit scrambling.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Modern AI systems demand perfectly reliable hardware, which carries an energy penalty: chips suppress variability between logic gates by operating at high voltage, and energy scales with voltage squared. This energy–reliability tradeoff is fundamental to the thermodynamics of computation.

**Key motivation**: Biology demonstrates this reliance on pristine components is not fundamental. The brain—built from heterogeneous, noisy neurons—leverages fault-tolerant circuit architectures to perform complex reasoning with orders of magnitude less energy than artificial neural networks. Grid cells, for instance, compute within an error-correcting code: position is represented redundantly across modules of different spatial periods, so noise in any one module can be corrected.

### The Problem with Current Models

Quantization (rounding weights to fewer bits) already degrades commercial LLMs significantly. Real low-energy hardware presents a far harder challenge: *random errors of arbitrary magnitude*. This fragility stems partly from "super weights"—outsized parameters critical to model function that act as single points of failure.

### Theoretical Framework

The authors draw on two key theoretical ideas:

1. **Coding theory**: Nature protects information in two ways:
   - **Repetition codes**: Keep many copies (crude, overhead grows with scale).
   - **"Good" codes**: Distributed redundancy where the overhead stays a fixed fraction regardless of scale (e.g., grid cells).

2. **Error-correction-enhanced evolvability** [39]: Adaptive processes (learning or evolution) are drawn toward fault-tolerant solutions because fault tolerance allows more configurations to be explored without catastrophic damage.

### Key Definitions

The paper measures computational capacity via the **equivalent fault-free model size** $\eta$:

$$
L(0, \eta, D) = L(p, N, D) \tag{1}
$$

where $L(p, N, D)$ is the loss of a model with $N$ parameters, fault rate $p$, trained on $D$ tokens. The **relative capacity** $\eta/N$ is the neural network analog of a code's rate: the fraction of parameters doing useful work rather than absorbing errors.

The central conjecture is formal fault tolerance:

$$
\lim_{N, D \to \infty} \frac{\eta}{N} \geq c, \quad \forall p < p_{\text{th}} \tag{3}
$$

for $p_{\text{th}} > 0$ and $0 < c \leq 1$, where $p_{\text{th}}$ is the fault-tolerance threshold (analogous to Eigen's error threshold for replicating genomes).

---

## Methodology

### Dataset

- **Training data**: 350-billion-token slice of FineWeb (web text corpus), tokenized with an 8192-token byte-pair-encoding vocabulary, packed into 1024-token sequences.
- **Validation**: Held-out split of 2 billion tokens from the same web crawls, with near-duplicates removed.

### Model Ladder

All models are Llama2-style decoder-only transformers (rotary position embeddings, RMSNorm, SwiGLU feed-forward blocks). Key architectural choices:

- Attention head dimension: 32 at every rung
- Width-to-depth ratio: held near 32
- Context length: 1024 tokens
- Sixteen logarithmically spaced (width, depth) pairs ranging from $7.9 \times 10^6$ to $9.3 \times 10^8$ total parameters

One architecture ($N = 5.0 \times 10^7$) was additionally trained at eight token budgets, eight random seeds per condition, and six training fault rates.

### Fault Injection Model

The error model emulates a **Low Energy Neural Network Accelerator (LENNA)**—a chip that runs arithmetic at voltages too low for perfect reliability but detects its own errors:

- **Hardware mechanism**: Operands carry a redundant residue (remainder modulo small bases) in a Redundant Residue Number System (RRNS). After every $k = 4$ fused multiply–add (FMA) cells, a cheap modulus check tests the running sum for consistency.
- **Error conversion**: A failed check discards those $k$ products and carries the previous partial sum forward—equivalent to zeroing a block of $k$ elements in one operand.
- **Fault application**: Random blocks of $k = 4$ matrix elements within attention and feed-forward operations are zeroed with probability $p$ (17 values between 0 and 0.2). Fresh masks are drawn for every sequence on every forward pass.
- **Scope**: Faults target matrix multiplications (which dominate >99% of FLOPs for large models); embeddings, norms, nonlinearities, and residual additions remain fault-free.

### Training

- Token budgets ranged from $7.9 \times 10^7$ to $3.5 \times 10^{10}$ tokens (fixed multiple of parameter count).
- The two axes (size and duration) were not fully crossed: ten architectures at or below $1.1 \times 10^8$ parameters were trained at 10, 20, 80, and 320 tokens per parameter; six larger ones only at 20 and below.
- Every combination of size, duration, and fault rate was trained at two random seeds.
- "Conservative" hyperparameters from a standard reference implementation were used to maintain clean scaling-law behavior.

### Scaling-Law Fitting

Equation 4 was fit to each fault-rate cohort separately on log-loss with a Huber loss, from multiple starting points. Confidence intervals came from 1,500 bootstrap resamples per cohort.

---

## Empirical Validation / Results

### Modified Scaling Laws for Fault-Hardened Models

The central empirical finding is that fault-hardened transformers obey a modified scaling law with a new curvature coefficient $\alpha_2$:

$$
\begin{array}{c} 
L(p, N, D) = E(p) + L_N + L_D \\
L_N(p, N) = e^{\alpha_0(p) - \alpha_1(p) n - \alpha_2(p) n^2} \\
L_D(p, D) = e^{\beta_0(p) - \beta_1(p) d} \\
n = \ln \frac{N}{N_0}, \quad d = \ln \frac{D}{D_0}
\end{array} \tag{4}
$$

where $E$ is the irreducible loss floor, $\alpha_k$ the model-size coefficients, $\beta_k$ the data coefficients, and $N_0 = D_0 = 4 \times 10^7$.

**Key coefficient behaviors** (Fig. 2):
- $\alpha_1$ (typical model-size coefficient) **shrinks** as $p$ increases.
- $\alpha_2$ (curvature term, pinned to zero for $p = 0$) **grows** with $p$—this drives capacity recovery beyond $N^*$.
- $\beta_1$ (data coefficient) **decreases** with $p$, meaning models become more expensive to train as fault rates rise.

### Capacity Recovery

The relative capacity $\eta/N$ follows a parabola in log model size:

$$
\begin{array}{c} 
\ln \frac{\eta}{N} = -c_0(p) - c_1(p) n + c_2(p) n^2 - \mathcal{A}(p, N) \\
n = \ln \frac{N}{N_0}, \quad N^* = N_0 \exp\left(\frac{c_1(p)}{2 c_2(p)}\right)
\end{array} \tag{2}
$$

Capacity first falls with size, then **recovers beyond a critical size $N^*$**—behavior impossible for repetition-based error correction, strongly suggesting models learn "good" codes.

The coefficients $c_k$ relate to the scaling-law coefficients via:

$$
\begin{array}{c} 
c_0 = \frac{\alpha_0(p) - \alpha_0(0)}{\alpha_1(0)}, \quad 
c_1 = 1 - \frac{\alpha_1(p)}{\alpha_1(0)}, \quad 
c_2 = \frac{\alpha_2(p)}{\alpha_1(0)} \\
\mathcal{A}(p, N) = \frac{1}{\alpha_1(0)} \ln\left(\frac{\Delta(p)}{L_N(p, N)} + 1\right) \\
\Delta(p) = E(p) - E(0)
\end{array} \tag{11}
$$

### Two Regimes of Fault-Hardened Scaling

The $p$-dependence of $\alpha_k$ hints at two regimes, reminiscent of crossing a fault-tolerance threshold:
- **Small $p$**: Decreases in $\alpha_1$ are mirrored by increases in $\alpha_2$ (models learn to correct increasingly severe errors).
- **Large $p$**: $\alpha_2$ plateaus while $\alpha_1$ continues to shrink (models become overwhelmed).

### Fault-Blind Models Collapse Under Inference-Time Errors

In stark contrast, fault-blind models ($p_{\text{train}} = 0$, $p_{\text{eval}} > 0$) exhibit:

1. **Divergent scaling**: Performance can actually *decrease* with training data $D$ past a threshold, lacking convergent scaling laws altogether (Fig. 3a).
2. **Logit heating**: The effective temperature $\tau$ of their output distribution rises rapidly with $p_{\text{eval}}$:

$$
\mathbb{P}(x_t = k | x_{t-1} = k_{t-1}, \dots, x_0 = k_0) = \frac{e^{a_k / \tau}}{\sum_j e^{a_j / \tau}} \tag{7}
$$

3. **Irreversible logit scrambling**: Even after removing temperature effects, fault-blind models lose capacity rapidly. The KL-divergence decomposes as:

$$
D(Q|P) = D(Q|P^*) + D(P^*|P) \tag{8}
$$

where $P^*$ is the nearest temperature-shifted distribution to $Q$. Fault-hardened models lose almost no capacity to the residual (unrescalable) term until $p_{\text{eval}}$ exceeds $p_{\text{train}}$, while fault-blind models deteriorate immediately.

### Training Cost of Robustness

Robustness is not free. The effective training budget $\delta$ of a faulty model (tokens needed to match a fault-free model trained on $\delta$ tokens) follows:

$$
\ln \frac{D}{\delta} = \frac{\beta_0(p) - \beta_0(0)}{\beta_1(0)} + \left(1 - \frac{\beta_1(p)}{\beta_1(0)}\right) d \tag{6}
$$

Since $\beta_1(p)$ decreases with $p$, the second term's coefficient is positive, meaning the cost grows with model size.

---

## Theoretical and Practical Implications

### Theoretical Significance

1. **Evidence for "good" codes**: The capacity recovery beyond $N^*$ strongly suggests models learn distributed, grid-cell-like codes rather than repetition schemes. This is the largest-scale evidence to date for error-correction-enhanced evolvability.

2. **Formal fault tolerance conjecture**: The authors conjecture fault-hardened models may be formally fault-tolerant (Eq. 3). Two asymptotic behaviors are possible, with $p$ controlling a smooth transition:
   - **Collapse** (if floor gap $\Delta \gg L_N$ at scale): $\ln(\eta/N) \propto -n$ (Eq. 9)
   - **Saturation** (if $\eta/N \to 1$): $\mathcal{A}(p, N) \approx \frac{\Delta(p)}{\alpha_1(0) L_N(p, N)}$ (Eq. 10)

3. **Hardware-software codesign**: The LENNA architecture demonstrates that error *detection* (cheap in RRNS arithmetic) can be delegated to hardware while error *correction* is learned by the model—echoing the brain's division of labor between grid modules and the grid–hippocampal loop.

### Practical Implications

- **Energy efficiency**: If fault-tolerant models can run on low-energy faulty hardware, inference could be orders of magnitude more energy efficient than current systems.
- **Hardware roadmap**: LENNAs add only a few percent overhead for error detection (a 3-bit residue on 12-bit FMAs), making them a practical near-term path to energy-efficient AI.
- **Training cost**: Robustness requires more training tokens, but the cost is a fixed multiplier that may be acceptable given inference energy savings.

---

## Conclusion

### Main Takeaways

1. Fault-hardened LLMs (trained with the same faults they encounter at inference) become *more* error-resilient as they scale—opposite to fault-blind models, which are catastrophically fragile.
2. The modified scaling law with curvature term $\alpha_2$ explains this recovery and suggests models learn "good" error-correcting codes.
3. The results support a formal conjecture of fault tolerance (Eq. 3), though experiments at $N \approx 10^9$ could not resolve the asymptotic regime.

### Future Directions

- **Large-scale validation**: Establishing Eq. 3 at commercial scales (~$10^{11}$ parameters) would cost ~$10^{26}$ FLOPs (~$10^8$ GPU-hours, similar to Llama-3 405B's total budget)—intractable within academia but suggested to be worth doing.
- **De-risking experiments**: Smaller experiments on natural or synthetic datasets with tunable irreducible floors $E$ could serve as "model organisms" to test whether fault tolerance persists as $E$ increases.
- **Mechanistic analysis**: The trained models (multiple terabytes of weights, publicly available) could be studied to understand how fault-hardened representations differ from fault-blind ones—paralleling how networks trained to navigate spontaneously develop grid-cell-like codes.

### Closing Perspective

The brain is the existence proof that fault tolerance, energy efficiency, and intelligence are compatible. The principle of error-correction-enhanced evolvability predicts fault tolerance is an expected outcome of adaptation in noisy environments. This work provides the largest-scale evidence to date that AI trained under the constraints that shaped nervous systems may come to share their robustness, potentially enabling orders-of-magnitude energy savings on future low-energy hardware.

---

_Markdown view of https://picx.dev/p/zpJb63, served by PicX — AI-generated visual whiteboard summaries of research papers._
