A paper compares distilled byte and token models trained at scales up to 1 trillion bytes and reports a higher performance ceiling for byte models as compute increases.
Published on September 14, 2026, the paper studies distilled byte and token models at scales of up to 1 trillion training bytes. It presents approximate and exact methods for converting token logits into byte logits, and reports that byte models reach a higher ceiling as compute increases.
The findings may inform engineers assessing tokenization and distillation when scaling language models, but they do not establish a universal choice. Consult the original paper and check its methods, experimental conditions, and results before applying the conclusion. If you use AI to study it, avoid entering confidential organizational data and verify responses against the original text.