Skip to content
Rota Nacional

Radar ·

RTX 3090 tests compare LLM throughput and efficiency

A briefing published on May 8, 2026 compares eight local LLMs on one RTX 3090, at power limits from 100 W to 450 W. Average throughput rose at the higher tested power, while energy efficiency fell.

Published on May 8, 2026, the briefing reports tests of eight local LLMs running on a single RTX 3090, at power limits ranging from 100 W to 450 W. Its focus is comparing throughput and energy efficiency during local inference.

Reported average throughput was 90.4 tokens per second at 225 W and 107.1 tokens per second at 450 W. Efficiency fell from 0.4167 to 0.2731 tokens per second per watt. To check the context and methods, consult the original publication and verify how models, measurements, and test conditions were defined; those details are not included in the supplied briefing.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free