A post published on September 23, 2026 describes two levels of KV sharing intended to reduce prefill cost and cache size, while improving retrieval in long-context workloads.
On September 23, 2026, a post introduced HySparse2, combining KV sharing across decoders with reuse between sparse and full-attention layers. The design aims to reduce prefill FLOPs and KV cache size in long-context inference.
According to the post, in agentic workloads with growing context, HySparse2 achieved better retrieval scores than the MiMo-V2.6 Hybrid SWA architecture in an evaluation with 1 million tokens. Consult the original post for its method and evaluation details; treat these as source-reported results and check configurations and metrics before comparing them with other systems.