A post claims KV cache use fell from 389,000 bytes to 890 bytes per token in three years. It says a linked article explains the approach with diagrams.
Published on September 13, 2026, the post attributes a reduction in KV cache from 389,000 bytes to 890 bytes per token over three years to an AI system. These figures are claims made by the post; the supplied material offers no independent measurements or methodological details to confirm them.
The text says the linked article explains the approach with diagrams and notes that KV cache size can affect memory use and the cost of serving long-context agents. To check the claim, consult the article cited in the post and examine how the values were measured, under what conditions, and against what comparison; those details are not in the supplied material.