BITCOS is a compression method for ternary LLMs. According to the report, it uses zero weights to store exact weights at 1.485 bits per weight.
The text also reports up to 18% higher decode throughput on CPU and up to 27% on GPU. These are reported results, not a guarantee for every model or device.
To assess the method, check the test conditions, including model, hardware, output quality, and memory use. Compare it with an equivalent baseline before deciding whether it suits your workload.
If you use AI to study or apply the material, do not submit personal data, secrets, or internal information without authorization. Review outputs and validate results in your own environment: the described technique is neither a privacy solution nor an implementation provided by Rota Nacional.