A post published April 23, 2026 reports 73 tokens per second for Qwen3.6-27B on a single RTX 3090, using speculative decoding with DFlash and DDTree. The author says the stack loads the model because its architecture and layer and head dimensions match Qwen3.5, but throughput is lower.
Published April 23, 2026, the post reports 73 tokens per second for Qwen3.6-27B on a single RTX 3090, using speculative decoding with DFlash and DDTree. According to the author, the stack can load the model because its architecture and layer and head dimensions match Qwen3.5; even so, throughput is lower.
The case illustrates how architectural compatibility can enable speculative decoding on consumer GPUs before dedicated upstream support is available. To verify the figures and test conditions, consult the original post and check the hardware, configuration, and metric it reports; the information available here does not detail those parameters.