Published on December 23, 2025, the account credits the Squinch algorithm with using one quarter of IB bandwidth while maintaining SOTA MFU in pretraining on GCP H100-TCPX.
A Character.AI post says the company used Squinch, a gradient-compression algorithm, during pretraining on GCP H100-TCPX. According to the account, the method maintained SOTA MFU while using one quarter of IB bandwidth.
The case illustrates how gradient compression can contribute to training efficiency when network bandwidth is constrained. To check the method, conditions, and scope of the claims, consult Character.AI's original post and assess the results in the stated context; the available summary provides no further evaluation details.