Dust: training transformers without gradients through activation perturbations
Radar: the publication describes Dust, a zeroth-order optimization method that uses parallel perturbations of activations as a virtual population to pretrain transformers, with scale experiments and comparisons with backpropagation and evolution strategies.
The publication, dated October 5, 2026, describes Dust, a zeroth-order optimization method that uses parallel perturbations of activations as a "virtual population" to pretrain transformers. According to the text, the work reports scale experiments and comparisons with backpropagation and with an evolution strategies method.
The work presents Dust as a practical alternative to backpropagation and reports how zeroth-order training behaves as model size and compute increase. Rota Nacional does not offer gradient-free training; this item is informational only. To check the results, consult the original publication through the source recorded in Radar and verify the metrics, methodology and experimental conditions. If you use AI to study the material, do not paste excerpts containing personal data.