Skip to content
Rota Nacional

Radar ·

HOLA proposes an exact KV cache for linear attention

Published on July 3, 2026, the article summary says HOLA combines a recurrent state based on the delta rule with a bounded, exact key-value cache. It reports improved Wikitext perplexity and robust needle recall on RULER up to 32k tokens.

Published on July 3, 2026, the article describes HOLA, which combines a recurrent state based on the delta rule with a bounded, exact key-value cache for linear attention. According to the summary, the design aims to improve long-range recall while keeping the state compressed.

The article reports improved Wikitext perplexity and robust needle recall on RULER up to 32k tokens. To check the findings, consult the original article and verify its test setup, metrics, and comparisons; the available summary does not provide those details.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free