Xiaomi MiMo v2.6
https://mimo.xiaomi.com/mimo-v2-6The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
Pelicans for Pro: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
One simple task: I needed an LLM to go through and clean up a few thousand page descriptions and titles in my personal search engine index, where the human web page authors had put in no effort sigh. I did a shoot out between Claude, Luna, GLM 5.3 Flash and Deepseek. Despite the high cost, Claude's descriptions were terrible, and even Opus warned me that the descriptions coming back from Haiku were "generalized, not accurate". I expected I would choose Luna because of price, and occasionally it did have wonderful descriptions (one captured emotion in a way no other model did). But in the end, the GLM 5.3 Flash descriptions were the easiest to read, they flow well while also being accurate & including necessary keywords, and being highly affordable. So it won out. It's a task that is nowhere near frontier, but a task where somehow China is better than frontier.
Pro [2]:, 1.02T total / 42B activated parameters
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
GPT 6 Astra 59.6
Claude Fable 5.1 55.1
Claude Opus 5 49.0
MiMo-V2.6-Pro 34.9
MiMo-V2.6-Flash 28.8
DeepSeek V4.1 Flash 26.8
MiMo-V2.5-Pro 1.5
ExploitGym GPT 6 Astra 42.4
Claude Fable 5.1 30.4
Claude Opus 5 22.1
MiMo-V2.6-Pro 17.8
MiMo-V2.6-Flash 6.0
MiMo-V2.5-Pro 0.1
DeepSWE v1.1 DeepSeek V4.1 Flash 74.2
Claude Opus 5 74.0
GPT 6 Astra 74.0
MiMo-V2.6-Pro 71.9
Claude Fable 5 70.0
MiMo-V2.6-Flash 67.9
MiMo-V2.5-Pro 19.0Some features of the release I like:
- Demonstration of diverse tasks, such as using a DAW
- Graphs from various benchmarks and price ranges
- Real world use of the model in scientific environments
Averages ~25-35tok/s which isn't bad for a first attempt.
It's because offpeak electricity is cheaper?
Funnily it's perfect if you are in the Pacific Time Zone because you can use it daytime 9am to 5pm
Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.
And as these models get better the pace of training is quickly speeding up too.
This doesn't bode particularly well for anthropic/OAI after they go public.
so weird to acknowledge someone being on the front edge, but not name it
MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks
Perhaps with IQ2 flash will run on 128G M5?
but now I got my "proof".
It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.
I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.