Skip to content
Rota Nacional

Radar ·

Mixture-of-Depths allocates compute per token

An article published September 22, 2026 presents a method for Transformers to allocate compute to sequence positions while respecting a per-layer budget.

Published on September 22, 2026, the article presents Mixture-of-Depths, a method for Transformer language models to allocate compute to specific sequence positions across layers. The proposal limits how many tokens participate in self-attention and MLP calculations in each layer, allowing compute to vary by token within a per-layer budget.

The material describes the approach; the available summary does not report evaluation results. To verify the method and its evidence, consult the original article and examine the methodology and results it presents. If you use AI to study or apply the idea, avoid sending personal data or internal information unnecessarily, and follow your organization’s data policy.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free