Skip to content
Rota Nacional

Radar ·

CUDA-L2 optimizes HGEMM kernels with reinforcement learning

Published on December 15, 2025, the paper presents an approach combining language models and reinforcement learning to optimize CUDA half-precision matrix multiplication kernels. It evaluates kernels across 1,000 configurations and reports results above matmul baselines.

Published on December 15, 2025, the paper describes CUDA-L2, an approach combining large language models and reinforcement learning to optimize CUDA half-precision matrix multiplication (HGEMM) kernels. Execution speed serves as the reward during optimization.

The work evaluates kernels across 1,000 configurations and reports results above matrix multiplication baselines. To assess the scope of the finding, consult the original paper and check how it defines the baselines, configurations, and speed measurements; the available summary does not detail these methods or quantify the performance difference.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free