Hacker News new | past | comments | ask | show | jobs | submit
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.

Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..

You have to test each model on your actual use case to see how well it really performs.

Yeah, totally. But... the problem is that I don't have time to test every single model that comes out. So I rely on reports like this to decide, should I even bother testing out Muse?
> Opus 5 is incredible at making games.

This is a bit vague. What sort of games with what technology?

My son was gifted an old Mac from his grandparents. It only supports OSX 10.13. I’ve been able to make several games that he genuinely likes (7 yo) Opus built them on my workstation and then pushed them to his computer and tested them over SSH. It handled all the asset creation or collection from CC0 licensed sources. I believe everything is built on the Godot engine. It’s really amazing to me. I don’t know what it would cost me to get someone to build custom games on a long deprecated computer architecture, but I paid Anthropic $20.
I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic as breakout or flappy bird can have wildly different quality between models. Yeah, you could call this animal on a bike benchmarking, but I don't think it is. IMO the problem space occupies an interesting area where you can ignore the pass/fail and focus on the actual level of the model to do something beyond that.

I doubt the OP meant something like creating the whole tech stack for WOW.

You seem to think I was disagreeing somehow.

I was just asking what kinds of games and with which technology.

Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.

Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
Or they did try to game the benchmarks and just didn’t do it well enough.

Benchmarks are one data point, not the only one, but the easiest one to compare.

Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.
If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.

Cherry picking the benchmarks you present is where the falsehoods lie.

Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
>We can throw benchmarks in the bin by now.

No, you really can't. This rhetoric on here is so profoundly boring and tired by now. People have been saying this noise about benchmarks for time eternal, usually because their pet didn't win.

Meta knows their models aren't as good -- demonstrated by their benchmark performance -- and their value proposition is a much lower price.