Only useful benchmarks are those you (in particular) don't have access to.
The only useful benchmarks are those you've created for your specific workflow. Only then can you assess whether a given model is better or worse for what you are using it for.
There are tools like promptfoo designed for this.