Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
That's not normally how people act when they're confident in their product
The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?
The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.
Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.
OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?