A report published on April 23, 2026 says DFlash and DDTree ran Qwen3.6-27B on a single RTX 3090, reaching 73 tokens per second with speculative decoding. This figure is attributed to the author’s case, not presented as an independent benchmark.
According to the report, the stack can load the model because its architecture and layer and head dimensions match Qwen3.5. Even so, throughput is lower than with that earlier model.
The example shows how architectural compatibility can enable speculative decoding on consumer GPUs before dedicated upstream support arrives. To assess whether the result is useful, check the full configuration, workload, and measurement conditions; do not treat one number as a performance guarantee.
If you use AI to study or apply this approach, avoid sending prompts, logs, or documents containing personal data or operational secrets. Prefer anonymized, authorized material, and check the chosen tool’s retention and access rules. Rota applies personal-data policies before inference; this does not replace controls for confidential information that is not personally identifying.