Skip to content
Rota Nacional

Radar ·

DFlash and DDTree run Qwen3.6-27B on an RTX 3090

A post published April 23, 2026 reports 73 tokens per second for Qwen3.6-27B on a single RTX 3090, using speculative decoding with DFlash and DDTree. The author says the stack loads the model because its architecture and layer and head dimensions match Qwen3.5, but throughput is lower.

Published April 23, 2026, the post reports 73 tokens per second for Qwen3.6-27B on a single RTX 3090, using speculative decoding with DFlash and DDTree. According to the author, the stack can load the model because its architecture and layer and head dimensions match Qwen3.5; even so, throughput is lower.

The case illustrates how architectural compatibility can enable speculative decoding on consumer GPUs before dedicated upstream support is available. To verify the figures and test conditions, consult the original post and check the hardware, configuration, and metric it reports; the information available here does not detail those parameters.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free