Published on September 27, 2025, the report describes an approach that iteratively reduces the energy assigned to next-token candidates. In tests with 44-million-parameter models on RedPajama-Data-v2, it outperformed same-sized vanilla transformers on three of four benchmarks.
A report published on September 27, 2025 presents Energy-Based Transformers as an alternative to predicting the next token in a single step. The method assigns an energy score to candidates and iteratively reduces it through gradient steps.
In the reported tests, using 44-million-parameter models on RedPajama-Data-v2, the approach outperformed same-sized vanilla transformers on three of four benchmarks. To assess the finding’s scope, consult the original material and verify the test setup, benchmarks, and reported comparisons; the available summary does not detail these elements.