Skip to content
Rota Nacional

Radar ·

Luce PFlash adds prefix caching and cold-start tuning

An update to a speculative inference server reports prefix caching, cold-start tuning, and a CUDA VMM fix. The author reports performance gains for Qwen3.6 27B.

On May 3, 2026, the Luce PFlash entry announced prefix caching, cold-start tuning, and a CUDA VMM fix. The project is described as a speculative inference server for language models on heterogeneous hardware and consumer GPUs.

The author reports performance about 10 times faster in warm runs and 2.5 times in cold runs, with block sparse attention autotuning, for Qwen3.6 27B. Consult the original repository to examine the implementation, and check its test conditions and methods before comparing the figures; the entry does not provide those details. If you use AI to study or apply the material, avoid submitting internal code or data without authorization and check your organization’s data policy.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free