AI Weekly Malaysia

Summaries

Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.

Reset

Showing 1-1 of 1 results

DateProviderScoreSummary
06 Oct 2026, 5:15 AMHacker News6.0 Dust: Pretraining Transformers Without Backpropagation

Q Labs Research (Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna, October 2026) published 'Dust', a zeroth-order pretraining method that perturbs activations independently at every token so each token acts as a virtual population member evaluated in one forward pass. The authors claim it is the first zeroth-order method competitive with backprop at pretraining transformer LMs, and report that from 1M tokens up it is roughly 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, a state-of-the-art evolution-strategies method, based on their extrapolations. They also report larger models are more population-efficient, with a 243M-parameter model beating a 120x smaller one at most population sizes, and alignment with backprop gradients holding up to 1B tokens tested.

Why: Treat the headline numbers as claims to verify, not facts: the EGGROLL comparison is explicitly extrapolated and competitiveness with backprop requires a 'substantially more compute' large-population regime, so there is no cheaper training run to switch to today. The concrete detail worth tracking is the scaling direction - 243M parameters outperforming a 120x smaller model at most population sizes, and gradient alignment holding to 1B tokens - because it contradicts the standard assumption that zeroth-order methods collapse at scale. If you rent GPU time for pretraining, the memory-per-step profile of a method that needs no backward pass is the thing to watch; no Malaysia-specific angle is present in the text.

Top