Ten advances in mathematics and theoretical computer science
https://openai.com/index/ten-advances-in-mathematics/The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
https://x.com/polynoamial/status/2083478171975082334
As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget
--
The linked tweet from Noam Brown at OpenAI reads:
> And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
> But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
Training the model is going to be amortized over other uses.
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
[edit: deleted a distracting comparison to Chess]
Which is less interesting work. And you probably need to do the hard grunt work by hand first to develop the skills and intuition to be able to verify an AI-generated result. So you can’t outsource everything to AI without loss of skill.
If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).
P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.
Was this different before chess computers were invented?
The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.
This was not a problem that was for sale
Yes. And there are many of them. I wonder what would help them come to terms with it. Seriously, people are going to be grieving over this. Loss of identity, loss of social standing, ideas of entire future lives that will now never happen. The greatest crime people may hold AI guilty of is taking away their dreams.
It's more like they've already automated the parts of the jobs that the humans most closely thought of as the "their job"
My guess from following this stuff quite closely is that these companies are still a couple years away from fully autonomous research staff.
[1]. https://www.anthropic.com/institute/recursive-self-improveme...
I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient
No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
I think you're over-estimating what a smarter highschooler could write.
A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:
Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor: repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)
Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.
which is infact a very important part of the prompt.
When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.
I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".
What is your mechanistic model of self awareness that yields this conclusion?
> It's a tool
Does your model suggest that tools can't have self awareness?