Published on July 11, 2026, the report presents a 9-billion-parameter hybrid architecture for long contexts and reports good benchmark results, with scalable inference up to 1 million tokens.
The report describes MiniCPM-SALA, a 9-billion-parameter hybrid architecture combining sparse attention and linear attention for long contexts. According to the text, the approach achieved good results on long-context benchmarks and supports scalable inference up to 1 million tokens.
The paper explores how combining attention mechanisms can balance context quality, compute, and memory use. To check the findings, consult the original paper and verify its methods, benchmarks, and evaluation conditions; the report does not specify those details.