Hacker News new | past | comments | ask | show | jobs | submit
Only useful benchmarks are those you (in particular) don't have access to.
The only useful benchmarks are those you've created for your specific workflow. Only then can you assess whether a given model is better or worse for what you are using it for.

There are tools like promptfoo designed for this.