BITCOS reduces ternary weights to 1.485 bits per weight
A September 20, 2026 report on Intel’s BITCOS method says it stores exact ternary weights at 1.485 bits per weight and reports decode-throughput gains of up to 18% on CPUs and 27% on GPUs.
The report on BITCOS, Intel’s method for ternary LLMs, says zero weights make it possible to store exact weights at 1.485 bits per weight. According to the text, tests recorded up to 18% higher decode throughput on CPUs and 27% on GPUs. The reported gains may matter when serving ternary models under memory and compute constraints.
The publication is dated September 20, 2026. To assess the result, consult the original post and check how it defines the models, hardware, metrics, and comparison conditions; those details are not in the available summary. If you use AI to study or apply the technique, avoid submitting personal or internal data unless necessary, and verify responses against the original technical source.