Published on March 11, 2026, the note describes new pipelines, reformulated computations, and memory optimizations for FlashAttention-4 on Blackwell GPUs.
The publication reports that FlashAttention-4 uses new pipelines, reformulated computations, and memory optimizations designed for NVIDIA Blackwell GPUs. On B200 GPUs, it reports speed gains over cuDNN 9.13 and Triton, but the available summary gives no figures.
The result may interest teams optimizing attention workloads on these GPUs; performance and compilation-time gains should be checked in each team’s environment. Consult the original publication and verify its configurations, metrics, and test conditions before comparing results or applying the optimizations.