The fact
On FrontierSWE benchmark for extended coding tasks, GLM-5.2 trails Anthropic's Claude Opus 4.8 by just 1 percentage point
The model still lags significantly behind closed-source rivals on reasoning capabilities
Click the link to read an article on the topic: