An article on techniques for improving GPU matrix-transposition performance with shared memory in CUDA C/C++.
Mark Harris’s article, published on the NVIDIA blog, discusses potential performance gains from using shared memory to transpose matrices in CUDA C/C++.
Transposition rearranges a matrix’s rows and columns. It is a common operation in GPU computing, and the article explains how shared memory can improve its performance.
If you use an AI tool to study or adapt this material, avoid submitting code that contains credentials, personal data, or internal information. Use synthetic examples or remove such data before sharing code.
Evaluate any implementation in your own context: the article concerns GPU performance, and a technique it describes does not guarantee the same results in other programs or on other hardware.