ResearchOctober 05, 2026
32
๐ŸŒถ๏ธ๐ŸŒถ๏ธ

Dust Takes On Backprop

Dust pretrains transformers without backprop โ€” competitive at scale.

#Backprop#Zeroth-Order#Pretraining#Scaling#Optimization
Dust fordert Backprop heraus
Share Article
๐Ÿ”ฅ What happened Researchers unveiled Dust, a zeroth-order method that competes with backprop for pretraining transformer language models. Dust perturbs activations per token, treating each token as a virtual population member evaluated in parallel during one forward pass. ๐Ÿ’ก Why it matters From 1M tokens up, Dust is extrapolated to be 10ยณโ€“10โดร— more efficient than EGGROLL, a state-of-the-art ES method. Strikingly, larger models are more population-efficient: a 243M-parameter model beats a 120ร— smaller one at most population sizes. โšก Our take If compute keeps winning, backprop's differentiability crutch may become optional. Dust is still expensive, but the scaling trend is a warning shot to gradient orthodoxy.
The title, summary and analysis of this item were produced automatically by an AI system and have not been editorially reviewed. They may contain errors, bias or omissions โ€” when in doubt, read the linked original source.