ResearchOctober 05, 2026
32
๐ถ๏ธ๐ถ๏ธ
Dust Takes On Backprop
Dust pretrains transformers without backprop โ competitive at scale.
#Backprop#Zeroth-Order#Pretraining#Scaling#Optimization

๐ฅ What happened
Researchers unveiled Dust, a zeroth-order method that competes with backprop for pretraining transformer language models. Dust perturbs activations per token, treating each token as a virtual population member evaluated in parallel during one forward pass.
๐ก Why it matters
From 1M tokens up, Dust is extrapolated to be 10ยณโ10โดร more efficient than EGGROLL, a state-of-the-art ES method. Strikingly, larger models are more population-efficient: a 243M-parameter model beats a 120ร smaller one at most population sizes.
โก Our take
If compute keeps winning, backprop's differentiability crutch may become optional. Dust is still expensive, but the scaling trend is a warning shot to gradient orthodoxy.
The title, summary and analysis of this item were produced automatically by an AI system and have not been editorially reviewed. They may contain errors, bias or omissions โ when in doubt, read the linked original source.
