Skip to content
Rota Nacional

Radar ·

Study compares attention levels in 13 LLMs

Experiments reported with 13 LLMs found the best result at 17% attention, equivalent to two of 12 layers. The comparison may inform architecture choices, but does not establish a universal rule.

Published on October 12, 2025, the account describes experiments in which the author trained 13 LLMs with different proportions of attention and DeltaNet linear attention. Among the tested configurations, 17% attention — two of 12 layers — achieved the best result, according to the text. The study may help engineers weigh full attention against linear attention in model architectures.

The available summary does not provide metrics, datasets, or experimental conditions, so it does not establish how broadly the result applies. To verify the conclusion, consult the original study and check how the tests and comparison criteria were defined. If you use AI to study or apply the material, avoid submitting personal data or confidential documents without authorization, and follow your organization’s policies.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free