Skip to content
Rota Nacional

Radar ·

CUDA matrix transposition with shared memory

Published on February 26, 2026, Mark Harris’s article discusses how shared memory can improve matrix transposition performance in CUDA C/C++.

On February 26, 2026, Mark Harris published an article on the NVIDIA blog about matrix transposition in CUDA C/C++. It discusses potential performance gains from using GPU shared memory for this common operation. The available source text does not give specific numerical results.

To check the scope and technical details, consult the original article on the NVIDIA blog and verify its examples, execution conditions, and any measurements in the text. If you use AI to study or apply the material, avoid submitting confidential code or documents until you have checked your organization’s policies.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free