A publication dated October 27, 2025 describes two models combining linear and softmax attention and claims efficiency gains from a proprietary FP8 operator library.
On October 27, 2025, the publication presented Ring-mini-linear-2.0, with 16 billion parameters, and Ring-flash-linear-2.0, with 104 billion. The models combine linear and softmax attention; according to the text, they reduce inference costs and improve training efficiency by 50% using a proprietary library of FP8 operators.
The attention design and FP8 training approach may interest teams optimizing inference and long-context training. To assess the claims, consult the original publication and check its conditions, metrics, and comparisons; the supplied information does not detail the methodology. If you use AI to study or apply the material, avoid submitting personal data or internal documents without authorization.