Published on October 22, 2025, the summary presents Ring-1T, a reasoning MoE model with one trillion parameters and about 50 billion active per token, and describes its three-stage training.
A post published on October 22, 2025 describes Ring-1T, a reasoning model based on a mixture of experts (MoE), with one trillion parameters and about 50 billion active per token. The summary reports three training stages: supervised fine-tuning with long chains of thought (long-CoT), reasoning reinforcement learning with verifiable rewards, and general reinforcement learning from human feedback (RLHF).
The text also names IcePop, C3PO++ and ASystem as components related to system approaches and scaling reinforcement learning for large reasoning models. To consult and verify the claims and technical details, read the original publication and check how it defines each stage and component; the available summary does not provide quantitative evaluation results.