HOLA combines linear attention with an exact KV cache
HOLA combines a delta-rule recurrent state with a bounded key-value cache. The paper reports improved Wikitext perplexity and robust needle recall on RULER up to 32k tokens.
HOLA combines a delta-rule-based recurrent state with an exact, bounded key-value (KV) cache used in attention. The proposal aims to improve long-range retrieval while retaining the compressed state of linear attention.
The paper reports better perplexity on Wikitext and robust needle recall on the RULER benchmark up to 32k tokens. These results describe the reported tests; they do not guarantee the same performance on other tasks or systems.
To evaluate the idea, identify the baseline configuration, cache size, and evaluation criteria. Compare perplexity and recall under equivalent conditions, including different context lengths, and record memory and latency costs.
If you use AI to study or apply the material, do not submit personal data or internal information without authorization. Review content before sharing it, and follow your organization’s policy for sensitive data.