Skip to content
Rota Nacional

Radar ·

llama.cpp adds Multi-Token Prediction support

A llama.cpp pull request adds support for MTP heads. Its author reports that Qwen3.6-27B reached 65 tokens per second on an RTX 3090, versus 38 with MTP enabled.

Published on May 17, 2026, the report says a llama.cpp pull request adds support for Multi-Token Prediction (MTP) heads, which the post says can predict multiple tokens per execution. Engineers serving compatible models can evaluate the technique as a way to increase inference throughput.

The author reports 65 tokens per second for Qwen3.6-27B on an RTX 3090 with MTP enabled, versus 38 in the comparison presented. To verify the result, consult the original pull request and post and check the test conditions; the report does not detail other benchmark parameters.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free