Published on July 11, 2026, Colibri is described as a dependency-free C inference engine that loads MoE model experts from disk as needed. The post says it runs a 744-billion-parameter model on a laptop with 25 GB of RAM.
The post presents Colibri as a dependency-free inference engine written in C, loading MoE model experts from disk on demand. According to the July 11, 2026 publication, it runs GLM-5.2, described as a 744-billion-parameter model, on a laptop with 25 GB of RAM.
The approach aims to reduce the memory needed to run large MoE models by bringing experts into memory as they are requested. To assess the result and its limits, consult the original publication and verify the requirements, test conditions, and measurement method; the available summary does not detail them.