Hacker News new | past | comments | ask | show | jobs | submit
Check out the "Completed steps on..." graph in this [1] evaluation.

That graph gives a good perspective of what models they've tested, and roughly what "subject" each step covers. It is on that task that they note this:

> Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2.

[1] - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...