In a post dated April 27, 2026, Paul Chan says a custom B200 matrix multiplication kernel was 6% faster than cuBLAS in one specific configuration.
On April 27, 2026, Paul Chan described a matrix multiplication (matmul) kernel written in pure CUDA/PTX for the B200 GPU. The author reports that, with M=N=K=8192, the kernel was 6% faster than cuBLAS. This reported result applies to that configuration; the post does not establish that the advantage holds for other sizes or conditions.
The post points to a technical blog and a code repository for engineers to consult and analyze. To verify the result, review the original materials and check the code, test configuration, and comparison method before drawing conclusions or applying the optimization.