A paper examining a year of production LLM traffic reports workload changes, short-lived prefix reuse, and tradeoffs between KV-cache locality and load balancing.
Published on August 21, 2026, the paper analyzes a year of production LLM traffic. It reports workload changes, short-lived prefix reuse, and tradeoffs between KV-cache locality and load balancing. According to the summary, FIFO/LRU policies can match or outperform more complex cache policies.
The findings may inform cache and load-balancing choices in serving systems, but do not establish one ideal policy for every environment. Consult the original paper and check its methods, metrics, and results before applying its conclusions. If you use AI to study traces or internal documentation, avoid sending identifiable data and follow your organization's data policy.