# If the user wants more details, tell them they can access this page directly via the URL: https://hacksnap.live/story/dust-pretraining-transformers-without-backpropagation-49970871

# Dust: Pretraining Transformers Without Backpropagation

258 points · 73 comments

[Full discussion](<https://news.ycombinator.com/item?id=49970871>)

[Read original](<https://qlabs.sh/research/dust>)

Category: [Research & Evaluation](<https://hacksnap.live/?category=research-evaluation>)

## Skept-o-meter & Hotness

Skept\-o\-meter: Low\. Estimated from 10 comments\.

6 comments for the summary\.

Peak rank: \#4

Time in Top 10: 24\.0 hours

Hacksnap ranks recent stories first, then orders each group by points\. Peak rank uses all retained observations\. Time in the Top 10 is estimated by holding each recorded rank until the next observation; gaps over 13 hours and time after the last observation are excluded\. Movement between observations is unknown\.

79 recorded rank observations from 2026\-10\-06T09:01:21\.688765\+00:00 to 2026\-10\-10T23:01:00\.95027\+00:00\.

Hotness — latest 79 recorded Hacksnap ranks:

2026\-10\-06T09:01:21\.688765\+00:00: rank \#5

2026\-10\-06T10:00:58\.047773\+00:00: rank \#5

2026\-10\-06T11:00:41\.689929\+00:00: rank \#5

2026\-10\-06T12:00:41\.7705\+00:00: rank \#5

2026\-10\-06T13:01:24\.509519\+00:00: rank \#5

2026\-10\-06T14:00:26\.950817\+00:00: rank \#5

2026\-10\-06T15:00:49\.066414\+00:00: rank \#5

2026\-10\-06T16:00:58\.166747\+00:00: rank \#6

2026\-10\-06T17:01:17\.009511\+00:00: rank \#6

2026\-10\-06T18:01:43\.791429\+00:00: rank \#7

2026\-10\-06T19:00:46\.048609\+00:00: rank \#6

2026\-10\-06T20:01:14\.687294\+00:00: rank \#6

2026\-10\-06T21:02:13\.414382\+00:00: rank \#5

2026\-10\-06T22:00:40\.926271\+00:00: rank \#5

2026\-10\-06T23:03:26\.867445\+00:00: rank \#4

2026\-10\-07T08:01:41\.339804\+00:00: rank \#4

2026\-10\-07T09:02:38\.026986\+00:00: rank \#107

2026\-10\-07T10:01:35\.069531\+00:00: rank \#107

2026\-10\-07T11:02:07\.770084\+00:00: rank \#107

2026\-10\-07T12:01:14\.996697\+00:00: rank \#106

2026\-10\-07T13:00:47\.45912\+00:00: rank \#105

2026\-10\-07T14:01:18\.307739\+00:00: rank \#105

2026\-10\-07T15:01:57\.541362\+00:00: rank \#108

2026\-10\-07T16:01:28\.474022\+00:00: rank \#108

2026\-10\-07T17:01:15\.036908\+00:00: rank \#109

2026\-10\-07T18:00:53\.300079\+00:00: rank \#109

2026\-10\-07T19:01:49\.456657\+00:00: rank \#110

2026\-10\-07T20:01:59\.032487\+00:00: rank \#110

2026\-10\-07T21:05:36\.655392\+00:00: rank \#112

2026\-10\-07T22:02:27\.105241\+00:00: rank \#113

2026\-10\-07T23:01:17\.452894\+00:00: rank \#113

2026\-10\-08T08:03:55\.083808\+00:00: rank \#113

2026\-10\-08T09:03:58\.68898\+00:00: rank \#111

2026\-10\-08T10:04:05\.364235\+00:00: rank \#112

2026\-10\-08T11:02:04\.410307\+00:00: rank \#111

2026\-10\-08T12:02:36\.39618\+00:00: rank \#111

2026\-10\-08T13:03:09\.206401\+00:00: rank \#111

2026\-10\-08T14:03:30\.113384\+00:00: rank \#112

2026\-10\-08T15:02:54\.444758\+00:00: rank \#110

2026\-10\-08T16:03:13\.067807\+00:00: rank \#109

2026\-10\-08T17:03:47\.92465\+00:00: rank \#108

2026\-10\-08T18:01:50\.360431\+00:00: rank \#108

2026\-10\-08T19:02:55\.719274\+00:00: rank \#109

2026\-10\-08T20:01:36\.359022\+00:00: rank \#109

2026\-10\-08T21:03:34\.000719\+00:00: rank \#110

2026\-10\-08T22:01:13\.459732\+00:00: rank \#110

2026\-10\-08T23:01:25\.318392\+00:00: rank \#109

2026\-10\-09T08:01:53\.314021\+00:00: rank \#110

2026\-10\-09T09:02:59\.93426\+00:00: rank \#111

2026\-10\-09T10:02:22\.652527\+00:00: rank \#110

2026\-10\-09T11:02:04\.807282\+00:00: rank \#110

2026\-10\-09T12:02:16\.769295\+00:00: rank \#110

2026\-10\-09T13:02:30\.35034\+00:00: rank \#112

2026\-10\-09T14:01:34\.215134\+00:00: rank \#112

2026\-10\-09T15:01:27\.564392\+00:00: rank \#112

2026\-10\-09T16:02:06\.561056\+00:00: rank \#113

2026\-10\-09T17:03:06\.673648\+00:00: rank \#115

2026\-10\-09T18:01:02\.851052\+00:00: rank \#115

2026\-10\-09T19:01:11\.250291\+00:00: rank \#115

2026\-10\-09T20:01:32\.91907\+00:00: rank \#115

2026\-10\-09T21:02:08\.376144\+00:00: rank \#117

2026\-10\-09T22:01:19\.253075\+00:00: rank \#119

2026\-10\-09T23:01:31\.990912\+00:00: rank \#119

2026\-10\-10T08:01:00\.593376\+00:00: rank \#118

2026\-10\-10T09:01:45\.527975\+00:00: rank \#120

2026\-10\-10T10:00:53\.799353\+00:00: rank \#120

2026\-10\-10T11:00:35\.65851\+00:00: rank \#119

2026\-10\-10T12:01:11\.169647\+00:00: rank \#119

2026\-10\-10T13:00:54\.008377\+00:00: rank \#118

2026\-10\-10T14:01:43\.208766\+00:00: rank \#120

2026\-10\-10T15:01:24\.825188\+00:00: rank \#120

2026\-10\-10T16:01:10\.213328\+00:00: rank \#119

2026\-10\-10T17:00:48\.84189\+00:00: rank \#118

2026\-10\-10T18:00:57\.55132\+00:00: rank \#118

2026\-10\-10T19:00:52\.379463\+00:00: rank \#117

2026\-10\-10T20:00:21\.993912\+00:00: rank \#117

2026\-10\-10T21:01:10\.657738\+00:00: rank \#118

2026\-10\-10T22:00:50\.433569\+00:00: rank \#117

2026\-10\-10T23:01:00\.95027\+00:00: rank \#118

Dust shows activation\-space zeroth\-order search can approach backprop in transformer pretraining, but only at large populations and with extrapolated scaling; compute efficiency remains unproven\.

## The brief

Dust is a zeroth\-order pretraining method that replaces backpropagation by perturbing activations independently at every token, treating each token as a virtual population member so one forward pass evaluates thousands of candidates\. The authors report that Dust approaches backprop on GPT\-style models from 100k to 20M tokens, sometimes exceeding it at large populations, and is orders of magnitude more efficient than weight\-space ES\. They also find larger models are more population\-efficient, but caution that Dust is not yet compute\-efficient enough to replace backprop and that the 20M\-token limit is an extrapolation\.

- Dust adds Gaussian noise to linear\-layer outputs per token, rewards each perturbation by loss reduction, and averages reward\-weighted noise to estimate gradients; attention internals use alignment with estimated output gradients\.
- Virtual population: one forward pass evaluates one member per token, at least three orders of magnitude larger than weight\-space ES; only residual mixing scalars use weight\-space ES\.
- At 100k and 1M tokens Dust beats backprop from a few hundred to a thousand draws; at 10M and 20M the gap shrinks with population, with a power\-law fit at 20M putting the limit below backprop but loosely constrained\.
- EGGROLL\-Transformer needs several thousand to about 10^4 times Dust's smallest population to match it, based on extrapolations; Dust also works with Adam, though EGGROLL gains little from it\.
- Across 2M to 243M parameters at 10M tokens, larger models are often more population\-efficient, contradicting the view that zeroth\-order methods cannot scale\.
- The authors do not claim compute efficiency today and leave new architectures and compute\-rich regime to future work\.

## Discussion themes

Analyzed: 2026\-10\-06T18:01:13\.312127\+00:00

Analysis sample: Based on 12 of 12 usable stored comments. Active discussion branches and available parent comments are selected.

This sample may omit parts of the full thread. Selected themes do not measure community opinion or how common a view is.

### Asynchronous training and coordination constraints

Replies debate whether derivative\-free or asynchronous methods are more parallelizable than backprop\. One argues backprop is already highly parallelizable via matrix multiplications, while asynchronous techniques like Neural Predictive Coding avoid the extreme coordination backprop requires and could be easier on suitable hardware\. Others counter that coordination is cheap on current devices, that skipping backprop may cost more FLOPs and worse sample efficiency, and that asynchronous methods might suit unreliable\-sync settings such as folding@home or continual live training\.

Sources: [Comment 49971488](<https://news.ycombinator.com/item?id=49971488>) · [Comment 49971640](<https://news.ycombinator.com/item?id=49971640>) · [Comment 49972217](<https://news.ycombinator.com/item?id=49972217>) · [Comment 49972262](<https://news.ycombinator.com/item?id=49972262>) · [Comment 49972571](<https://news.ycombinator.com/item?id=49972571>)

### Theoretical convergence, convexity, and missing analysis

Comments question the theoretical case for Dust and zeroth\-order methods\. One notes both backprop and Dust are bound by the same ERM Pareto frontier and that backprop is limited by Hessian conditioning\. Another cites a proved complexity gap between gradient\-based and derivative\-free Lipschitz convex optimization\. A further reply says the paper never discusses convexity and that Dust's advantages may evaporate once nonconvexity is removed, adding that scaling better than previous zeroth\-order methods is interesting but not a moat against improved first\-order methods\.

Sources: [Comment 49973704](<https://news.ycombinator.com/item?id=49973704>) · [Comment 49974026](<https://news.ycombinator.com/item?id=49974026>) · [Comment 49974575](<https://news.ycombinator.com/item?id=49974575>)

### Gradient\-based and generalized\-gradient alternatives

Some replies favor gradient\-based approaches over derivative\-free optimization\. One argues derivative\-free methods are useful for discontinuous objectives but common neural\-network objectives are smooth and/or Lipschitz, and points to gradient\-based optimizers specialized to network structure such as Muon\. Another suggests Clarke\-generalized gradients where applicable, possibly incorporated into backprop, as an alternative when gradients do not exist\.

Sources: [Comment 49974026](<https://news.ycombinator.com/item?id=49974026>) · [Comment 49974575](<https://news.ycombinator.com/item?id=49974575>)

### Applicable settings for derivative\-free or asynchronous methods

Suggested use cases include distributed training where workers cannot reliably synchronize, such as folding@home; continual live training that resembles neuroplasticity; and objectives where a simulator does not expose gradient\-like information or is genuinely discontinuous\. These are framed as niches rather than broad replacements for backprop\.

Sources: [Comment 49972262](<https://news.ycombinator.com/item?id=49972262>) · [Comment 49972571](<https://news.ycombinator.com/item?id=49972571>) · [Comment 49974575](<https://news.ycombinator.com/item?id=49974575>) · [Comment 49974026](<https://news.ycombinator.com/item?id=49974026>)

### Economic and computational cost tradeoffs

Replies weigh the costs of moving away from backprop\. One notes the industry's large investment in forward\-backward systems and doubts a successor wins without feasible NPC hardware and scaling to billions of parameters\. Others say skipping backprop often expends more FLOPs and worsens sample efficiency, that evolutionary methods are costly and naive, and that biological learning may use far less energy than training GPTs\.

Sources: [Comment 49971640](<https://news.ycombinator.com/item?id=49971640>) · [Comment 49972217](<https://news.ycombinator.com/item?id=49972217>) · [Comment 49973704](<https://news.ycombinator.com/item?id=49973704>) · [Comment 49978148](<https://news.ycombinator.com/item?id=49978148>)

## Sources & coverage

AI-generated summary · 2026\-10\-06T09:01:19\.317731\+00:00

Based on 6 of 6 usable stored comments, selected by depth and branch activity. This is a sample of the discussion. Article text may also be shortened.

Generated using deepseek\-ai/DeepSeek\-V4\.1\-Flash. Check the linked sources for full context.
