Published on September 26, 2025, Modal’s post describes reverse engineering FlashAttention-4 and attributes a 20% performance gain to changes in computation and execution.
On September 26, 2025, Modal published a reverse-engineering analysis of FlashAttention-4 for engineers exploring kernel optimization for attention workloads. The post highlights asynchrony, fast approximate exponents, and a more efficient softmax.
According to the post, a cubic polynomial, improved numerical stability, and increased asynchrony contributed to a 20% performance gain. Consult Modal’s original post to verify its methods and context; these results alone do not show that the same optimization will work in other kernels or environments.