Back to News

Dust: Zeroth-Order Pretraining Method Rivals Backprop for Transformer LMs

87 points · 13 comments#zeroth-order#pretraining#backprop#evolutionary-strategies

A research correspondence introduces Dust, the first zeroth-order pretraining method for transformer language models that is competitive with backpropagation. Dust perturbs activations at every token, treating each token as a virtual population member and evaluating them in parallel in a single forward pass, and is orders of magnitude more efficient than weight-space evolutionary strategies. The method even exceeds backprop in some settings, and larger models (e.g., 243M parameters) appear more population-efficient, challenging the assumption that zeroth-order methods do not scale.

Coverage timeline

  1. Hacker NewsE-Reverance

    October 2026 Correspondence to s@qlabs.sh·Code· ## TL;DR * We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a _virtual_ population member and one forward pass evaluates them all in parallel. * Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even _exceeds_ it. This hints that in a compute-rich regime we might be able to surpass backprop. * Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $1 0^{3}$ to $1 0^{4}$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations. * Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model