Skip to content
Rota Nacional

Radar ·

Squinch compresses gradients in pretraining

Published on December 23, 2025, the account credits the Squinch algorithm with using one quarter of IB bandwidth while maintaining SOTA MFU in pretraining on GCP H100-TCPX.

A Character.AI post says the company used Squinch, a gradient-compression algorithm, during pretraining on GCP H100-TCPX. According to the account, the method maintained SOTA MFU while using one quarter of IB bandwidth.

The case illustrates how gradient compression can contribute to training efficiency when network bandwidth is constrained. To check the method, conditions, and scope of the claims, consult Character.AI's original post and assess the results in the stated context; the available summary provides no further evaluation details.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free