Published on September 28, 2026, the post reports that kernel design agents generated and optimized Kimi Delta Attention kernels, reaching up to 2.96× the performance of FlashKDA on B300 with one-tenth the state error.
On September 28, 2026, an NVIDIA blog described using kernel design agents to generate and optimize Kimi Delta Attention kernels. It reports up to 2.96× the performance of FlashKDA on B300, with one-tenth the state error.
The post also covers kernel optimization and diagnosis in CUDA and specialized kernel languages. To check the result, consult the original post and verify its comparison conditions and reported metrics. If you use AI to study or apply the material, avoid submitting confidential organizational data.