Skip to content
Rota Nacional

Radar ·

Dust: training transformers without gradients through activation perturbations

Radar: the publication describes Dust, a zeroth-order optimization method that uses parallel perturbations of activations as a virtual population to pretrain transformers, with scale experiments and comparisons with backpropagation and evolution strategies.

The publication, dated October 5, 2026, describes Dust, a zeroth-order optimization method that uses parallel perturbations of activations as a "virtual population" to pretrain transformers. According to the text, the work reports scale experiments and comparisons with backpropagation and with an evolution strategies method.

The work presents Dust as a practical alternative to backpropagation and reports how zeroth-order training behaves as model size and compute increase. Rota Nacional does not offer gradient-free training; this item is informational only. To check the results, consult the original publication through the source recorded in Radar and verify the metrics, methodology and experimental conditions. If you use AI to study the material, do not paste excerpts containing personal data.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 10,00.

Try free