A publication reports that keeping four initial tokens and a 64-token window, without training, matched or outperformed distilled linear-attention models on almost all benchmarks.
The publication describes a token-retention strategy as an alternative to the KV cache: keep the first four tokens and a 64-token window, without additional training. According to the report, the method matched or outperformed distilled linear-attention models on almost all benchmarks evaluated.
The result suggests a way to reduce reliance on the KV cache without post-training, but it does not show that the strategy will work equally well for every task or deployment. If using AI to study or test the idea, avoid entering personal or confidential data: use synthetic or anonymized examples, and check the tool’s retention and access policies.