Published on December 24, 2025, the report describes a training-free method that reuses key patterns across layers. It reports speed gains on H100 GPUs and little accuracy loss on long-context benchmarks.
On December 24, 2025, a publication introduced Kascade, a training-free sparse-attention method that reuses key patterns across layers. According to the report, on H100 GPUs the method achieved generation up to 4.1 times faster and prefill 2.2 times faster, with little accuracy loss on long-context benchmarks.
The reported results may interest engineers working on long-context inference, but do not by themselves establish performance on other hardware or in other settings. Consult the original publication and check its benchmarks, test conditions, and metric definitions before drawing conclusions. If you use AI to study or apply the method, avoid including personal data or internal information without authorization.