Skip to content
Rota Nacional

Radar ·

CUDA/PTX kernel beats cuBLAS on B200

In a post dated April 27, 2026, Paul Chan says a custom B200 matrix multiplication kernel was 6% faster than cuBLAS in one specific configuration.

On April 27, 2026, Paul Chan described a matrix multiplication (matmul) kernel written in pure CUDA/PTX for the B200 GPU. The author reports that, with M=N=K=8192, the kernel was 6% faster than cuBLAS. This reported result applies to that configuration; the post does not establish that the advantage holds for other sizes or conditions.

The post points to a technical blog and a code repository for engineers to consult and analyze. To verify the result, review the original materials and check the code, test configuration, and comparison method before drawing conclusions or applying the optimization.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free