A report published on September 12, 2026 describes running a model on an RTX 5090 with 125.7 GiB of RAM and part of a 502 GB GGUF file on NVMe.
A report published on September 12, 2026 describes running a language model on an RTX 5090 GPU with 125.7 GiB of RAM. Part of a 502 GB GGUF file remained on NVMe storage rather than fitting entirely in GPU memory or RAM.
The author reports 5.12 tokens per second for new content and up to 21.27 tokens per second when data was already resident. These are reported results for that configuration, not evidence of universal performance. To assess the method, consult the original post and check its hardware, test conditions, and definition of resident data before comparing results.