Published on June 25, 2026, the report gives GLM 5.2 (Max reasoning) a score of 34.29%, versus 34.08% for Opus 4.8 Max, and reports zero failures across 84 GLM 5.2 attempts.
The report published on June 25, 2026 says GLM 5.2 (Max reasoning) scored 34.29% on PostTrainBench, slightly above Opus 4.8 Max at 34.08%. It also reports no failures across 84 GLM 5.2 attempts.
According to the publication, the ranking compares model scores and execution reliability. To check the result, consult the original publication and verify its figures and ranking context; the report does not detail other results or criteria.