Hacker News new | past | comments | ask | show | jobs | submit
I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

Terminal Bench 4.0

  GPT 6 Astra             59.6
  Claude Fable 5.1        55.1
  Claude Opus 5           49.0
  MiMo-V2.6-Pro           34.9
  MiMo-V2.6-Flash         28.8
  DeepSeek V4.1 Flash     26.8
  MiMo-V2.5-Pro            1.5
ExploitGym

  GPT 6 Astra             42.4
  Claude Fable 5.1        30.4
  Claude Opus 5           22.1
  MiMo-V2.6-Pro           17.8
  MiMo-V2.6-Flash          6.0
  MiMo-V2.5-Pro            0.1
DeepSWE v1.1

  DeepSeek V4.1 Flash     74.2
  Claude Opus 5           74.0
  GPT 6 Astra             74.0
  MiMo-V2.6-Pro           71.9
  Claude Fable 5          70.0
  MiMo-V2.6-Flash         67.9
  MiMo-V2.5-Pro           19.0
loading story #49793324
loading story #49794219
loading story #49793651