Skip to content
Rota Nacional

Guides ·

DFlash and DDTree run Qwen3.6-27B on an RTX 3090

A report claims 73 tokens per second for Qwen3.6-27B on one RTX 3090 using speculative decoding with DFlash and DDTree. The result depends on architectural compatibility and is not the same as dedicated upstream support.

A report published on April 23, 2026 says DFlash and DDTree ran Qwen3.6-27B on a single RTX 3090, reaching 73 tokens per second with speculative decoding. This figure is attributed to the author’s case, not presented as an independent benchmark.

According to the report, the stack can load the model because its architecture and layer and head dimensions match Qwen3.5. Even so, throughput is lower than with that earlier model.

The example shows how architectural compatibility can enable speculative decoding on consumer GPUs before dedicated upstream support arrives. To assess whether the result is useful, check the full configuration, workload, and measurement conditions; do not treat one number as a performance guarantee.

If you use AI to study or apply this approach, avoid sending prompts, logs, or documents containing personal data or operational secrets. Prefer anonymized, authorized material, and check the chosen tool’s retention and access rules. Rota applies personal-data policies before inference; this does not replace controls for confidential information that is not personally identifying.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free