AI Weekly Malaysia

Back to items Summaries

Dust: Pretraining Transformers Without Backpropagation

ID
32245
Status
summarized
Published
06 Oct 2026, 5:15 AM
Fetched
06 Oct 2026, 7:28 PM
Provider
Hacker News
Category
dev-community
Original URL
https://qlabs.sh/research/dust
Source URL
https://hnrss.org/best

Summary

Score
6.0
Created
06 Oct 2026, 7:28 PM
Tags
Audience
developersai_ml_learners

What happened

Q Labs Research (Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna, October 2026) published 'Dust', a zeroth-order pretraining method that perturbs activations independently at every token so each token acts as a virtual population member evaluated in one forward pass. The authors claim it is the first zeroth-order method competitive with backprop at pretraining transformer LMs, and report that from 1M tokens up it is roughly 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, a state-of-the-art evolution-strategies method, based on their extrapolations. They also report larger models are more population-efficient, with a 243M-parameter model beating a 120x smaller one at most population sizes, and alignment with backprop gradients holding up to 1B tokens tested.

Why it matters

Treat the headline numbers as claims to verify, not facts: the EGGROLL comparison is explicitly extrapolated and competitiveness with backprop requires a 'substantially more compute' large-population regime, so there is no cheaper training run to switch to today. The concrete detail worth tracking is the scaling direction - 243M parameters outperforming a 120x smaller model at most population sizes, and gradient alignment holding to 1B tokens - because it contradicts the standard assumption that zeroth-order methods collapse at scale. If you rent GPU time for pretraining, the memory-per-step profile of a method that needs no backward pass is the thing to watch; no Malaysia-specific angle is present in the text.

Discussion angle

Which parts of the Dust result are measured versus extrapolated - specifically the 10^3-10^4x efficiency claim against EGGROLL - and whether the 'larger models are more population-efficient' finding survives independent replication before anyone treats backprop-free pretraining as a real option.

Top