Skip to content
Rota Nacional

Radar ·

FlashAttention from scratch in CuPy and CUDA

Published on March 1, 2026, this project implements FlashAttention with tiling and online softmax without materializing the NxN score matrix. The author reports mathematical verification, but no speedup over cuBLAS.

Published on March 1, 2026, the project presents a practical implementation of FlashAttention from scratch, using CuPy and custom CUDA kernels. It employs tiling and online softmax, avoiding materialization of the NxN score matrix. The author reports verifying the implementation's mathematics, but did not achieve a speedup over cuBLAS.

The result highlights that a correct implementation is not necessarily faster than a reference library. To assess the work, consult the original publication and check its implementation details, test conditions, and measurements; these help interpret the comparison without assuming causes the report does not provide. If you use AI to study or adapt the code, avoid sending proprietary code, credentials, or personal data unnecessarily, and follow your organization's data policies.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free