A post reports reducing the Hy4 preview from 1.5 TB to 200 GB with per-layer precision settings, and compares results on two benchmarks.
According to the post, Tencent reduced the Hy4 preview from 1.5 TB to 200 GB, using calibration data to set each layer’s precision; some layers were set to 1.31 bits. The post reports scores of 83.2 on MCP Atlas and 81.3 on SWE-Bench Multi, compared with 82.9 as a reference.
The result illustrates an engineering trade-off: per-layer quantization can shrink a model, but should be evaluated alongside benchmark performance and application needs. These figures are the post’s claims and do not, by themselves, show that the technique will have the same effect on other models or tasks. When using AI to study or apply the material, do not submit internal or personal data without authorization, and follow your organization’s data-protection policy.