Story Detail of id 47679149 | Liveview Hacker News

jaggs20 hours ago | on: GLM-5.1: Towards Long-Horizon Tasks

How does it compare to Kimi 2.5 or Qwen 3.6 Plus?

General intelligence (not coding) comparison: https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-...

BoorishBears10 hours ago | root | parent

Is there really no rule that discourages 99% of your interactions with HN from being peddling some useless slop benchmark?

XCSme6 hours ago | root | parent

If it's relevant to the discussion, I hope not.

I've spent probably over100 hours working on this benchmarking/site platform, and all tests are manually written. For me (and many others that reached out to me) are not useless either. I use this myself regularly when choosing and comparing new models. I honestly beleive it is providing value to the conversation.

Let me know if you know of a better platform you can use to compare models, I built this one because I didn't find any with good enough UX.

loading story #47688005

eis20 hours ago | parent | next

The blog post has a benchmark comparison table with these two in it

jaggs19 hours ago | root | parent

Thanks, I missed that. It's very interesting. They're quite close, but I found Qwen 3.6 plus was just marginally better than Kimi 2.5. But looking at the stats I'll definitely give GLM 5.1 a try now. [edit: even though looking at it, it's not cheap and has a much smaller context size.And I can't tell about tool use.]

DeathArrow20 hours ago | parent

Compared to Kimi 2.5 or Qwen 3.6 Plus I don't know, but I ran GLM 5 (not 5.1) side by side with Qwen 3.5 Plus and it was visibly better.

#visit	13,257,283
#session	74,665
#live-session	0