Dust: Pretraining Transformers Without Backpropagation
- ID
- 32245
- Status
- summarized
- Published
- 06 Oct 2026, 5:15 AM
- Fetched
- 06 Oct 2026, 7:28 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://qlabs.sh/research/dust
- Source URL
- https://hnrss.org/best
Summary
- Score
- 6.0
- Created
- 06 Oct 2026, 7:28 PM
- Tags
- Audience
- developersai_ml_learners
What happened
Q Labs Research (Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna, October 2026) published 'Dust', a zeroth-order pretraining method that perturbs activations independently at every token so each token acts as a virtual population member evaluated in one forward pass. The authors claim it is the first zeroth-order method competitive with backprop at pretraining transformer LMs, and report that from 1M tokens up it is roughly 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, a state-of-the-art evolution-strategies method, based on their extrapolations. They also report larger models are more population-efficient, with a 243M-parameter model beating a 120x smaller one at most population sizes, and alignment with backprop gradients holding up to 1B tokens tested.
Why it matters
Treat the headline numbers as claims to verify, not facts: the EGGROLL comparison is explicitly extrapolated and competitiveness with backprop requires a 'substantially more compute' large-population regime, so there is no cheaper training run to switch to today. The concrete detail worth tracking is the scaling direction - 243M parameters outperforming a 120x smaller model at most population sizes, and gradient alignment holding to 1B tokens - because it contradicts the standard assumption that zeroth-order methods collapse at scale. If you rent GPU time for pretraining, the memory-per-step profile of a method that needs no backward pass is the thing to watch; no Malaysia-specific angle is present in the text.
Discussion angle
Which parts of the Dust result are measured versus extrapolated - specifically the 10^3-10^4x efficiency claim against EGGROLL - and whether the 'larger models are more population-efficient' finding survives independent replication before anyone treats backprop-free pretraining as a real option.