Luce PFlash adds prefix caching and cold-start tuning
An update to a speculative inference server reports prefix caching, cold-start tuning, and a CUDA VMM fix. The author reports performance gains for Qwen3.6 27B.
On May 3, 2026, the Luce PFlash entry announced prefix caching, cold-start tuning, and a CUDA VMM fix. The project is described as a speculative inference server for language models on heterogeneous hardware and consumer GPUs.
The author reports performance about 10 times faster in warm runs and 2.5 times in cold runs, with block sparse attention autotuning, for Qwen3.6 27B. Consult the original repository to examine the implementation, and check its test conditions and methods before comparing the figures; the entry does not provide those details. If you use AI to study or apply the material, avoid submitting internal code or data without authorization and check your organization’s data policy.