A llama.cpp pull request adds support for MTP heads. Its author reports that Qwen3.6-27B reached 65 tokens per second on an RTX 3090, versus 38 with MTP enabled.
Published on May 17, 2026, the report says a llama.cpp pull request adds support for Multi-Token Prediction (MTP) heads, which the post says can predict multiple tokens per execution. Engineers serving compatible models can evaluate the technique as a way to increase inference throughput.
The author reports 65 tokens per second for Qwen3.6-27B on an RTX 3090 with MTP enabled, versus 38 in the comparison presented. To verify the result, consult the original pull request and post and check the test conditions; the report does not detail other benchmark parameters.