Hacker News new | past | comments | ask | show | jobs | submit

Ten advances in mathematics and theoretical computer science

https://openai.com/index/ten-advances-in-mathematics/
loading story #49160757
loading story #49158671
loading story #49133172
Pretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely.

The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.

loading story #49160352
loading story #49160288
loading story #49159750
loading story #49159420
loading story #49160443
Which straw man are you arguing against?
loading story #49135838
loading story #49135292
loading story #49158822
loading story #49160006
loading story #49134299
loading story #49158499
loading story #49159958
loading story #49159469
loading story #49160775
loading story #49134323
loading story #49138240
loading story #49133534
loading story #49144681
loading story #49160237
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I want to know:

1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?

It seems like they threw it a decently large battery of open math problems and probably limited it to something like $200-500 per problem:

https://x.com/polynoamial/status/2083478171975082334

As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget

--

The linked tweet from Noam Brown at OpenAI reads:

> And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).

> But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.

I believe we're seeing a new kind of mathematics that will require completely new formats for publication, a bit similar to those used in experimental sciences. AI-powered mathematics should be fully reproducible, so it's the authors' responsibility to disclose the exact model type, inference settings/seeds and the full prompt history leading to the result. Of course that would ideally require open weights models.

It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.

loading story #49136501
loading story #49136405
loading story #49136711
I guess people will always find something to gripe about.
Yeah I remember reading about something along the lines of Mathematics is now about the scaffolding around you find the problems/solutions not just the problems and solutions. For teaching purposes. This was before this ai craze
I don't think you want to bring cost into this argument.

Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.

Do you really think that if you paid that to humans, they will deliver the same results?

loading story #49133539
loading story #49135823
loading story #49134144
loading story #49135377
loading story #49134030
loading story #49134482
loading story #49133415
[flagged]
loading story #49132393
loading story #49132714
> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.

Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?

Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.

Many less important Erdos problems have been solved by amateurs prompting ChatGPT 5.{3,4,5,6} Pro using their $200 subscription.
loading story #49132342
loading story #49134850
loading story #49134684
loading story #49160717
Great stuff. Wonder how many of the ten problems where solved by independent mathematicians not linked to OpenAI
Zero. These were open problems.
loading story #49160709
loading story #49133823
loading story #49159789
loading story #49136109
loading story #49133465
loading story #49141668
I wonder if Erdos would be saying " It's fine that y'all are answering my questions, but who is asking better questions??"
I'm enjoying learning about these hard problems, but this line about credit made me chuckle:

> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness

Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?

well, a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle, see https://infosec.exchange/@0xabad1dea/117002106099986943 and https://lipn.info/@mevenlennonbertrand/116997917683191056
loading story #49132370
loading story #49132386
No, the correctness isn't for the "inside the Lean proofs", but for the translation of "human language math" and its formal Lean variant.
loading story #49132279
I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)
I wonder what the total cost of this research was, including the salary for their mathematicians and engineers.
Why would you factor in salary unless they had to baby it through. You would only count the hours for setting up the harness and prompt and checking the result.

Training the model is going to be amortized over other uses.

loading story #49133842
> The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices.

https://x.com/polynoamial/status/2083470822258467194

loading story #49132435
Given that OpenAI pays their employees with stock surely a breathtaking number, but not a very meaningful number now that the infrastructure is in place and the models are trained. AI could never get better and it would still be incredibly disruptive.
I don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians.

Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.

[edit: deleted a distracting comparison to Chess]

The old way of establishing career credibility is being destroyed, for better or worse. Accomplishments that used to be career-defining are hard to distinguish from AI, and correlate more with access to compute. Think about Bill Gates's math paper he wrote in college. That kind of thing is gone now as a path to credibility. There's still competitions and grades, but the diversity of paths is going away. Maybe new ones will open up. This is a competitive advantage for old people who have credible pre-2025 accomplishments they can point to.
If accomplishments can't be distinguished between talented people and untalented people with compute, is there really a point in trying? I suppose one can hope that talented people given compute will be more effective than untalented people with compute, but I despair that that may not be true for much longer.
{"deleted":true,"id":49132664,"parent":49132593,"time":1785576019,"type":"comment"}
Knowledgable people can confirm what the AI produces is correct. I could make ChatGPT produce a result on an open question and I would have zero way to verify its actual correctness.

Which is less interesting work. And you probably need to do the hard grunt work by hand first to develop the skills and intuition to be able to verify an AI-generated result. So you can’t outsource everything to AI without loss of skill.

loading story #49135798
Given that we were nowhere near this state even two years ago, I think it’s a question of velocity more so than just distance.
The chess analogy is awful. If you simply want to know the answer to a chess problem, give it to the engine. Chess only lives on because it's a competition between humans to test their skill (just like bicycles, cars, trains didn't eliminate foot races) ... the computer is largely factored out, but not entirely -- people train with the computer, use it to check whether they played correctly, ... and they cheat. A lot. Thus there are more and more sophisticated mechanisms to detect and prevent cheating.

If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).

P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.

I’ve deleted it but no it’s not awful anymore than saying “we survived WWII, we can survive this.” The point was that change happens but humans find a way forward.
Every time someone makes a comparison to chess I die inside. Chess is a spectator sport primarily funded by a few eccentric billionaires. Players artificially constrain themselves in timed environments knowing that they will never be able to produce better moves than a smartphone because a select few people find it interesting. Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs. I shudder to imagine what will happen to the tens of thousands of non-Fields medalist caliber mathematicians if math goes the way of chess. Perhaps Terence Tao and a few other famous mathematicians will be funded by Peter Thiel to report on how well humanity can keep up with the machines? How do you expect any mathematician to be optimistic about this comparison.
>Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs.

Was this different before chess computers were invented?

Fully agreed. As someone who both loves chess and works on chess engines... these comparisons to chess needs to stop.
The distinction is mathematician vs mathematics. Mathematics is going to reach new heights beyond the wildest dreams of contemporary mathematicians. But perhaps without the participation of many paid mathematicians.
Just sounds dystopian,
Another noteworthy difference is that Stockfish is also gpl.
If there was any real money in it Stockfish would not be the best chess engine.
Thats not the point, if there were a better proprietary engine stockfish would still be there as a baseline. Anyone can access an engine as good as stockfish to practice against. Are any open models touting mathematical breakthroughs?
There is money in this, so of course the closed models are far ahead. The open models will likely catch up a bit at some point, just as Stockfish caught up to AlphaZero. That being said, there are already a couple. It seems Deepseek has a claimed proof to the "Ziegler's Cross-Polytope Conjecture" [0], but I can't speak to the significance of the result.

[0] https://arxiv.org/abs/2606.31640

loading story #49134009
As in chess and go and also coding for the past ~year there are two groups of people: the disappointed and the enthusiastic. The disappointed are sad that they lost their advantage and that the craft they honed for years or decades has rapidly lost its value; the enthusiastic are excited about the future and what computers can bring to their domain and how it will evolve. I’m a bit of both if it comes to programming, more enthusiastic than disappointed, but also more than a bit terrified about the pace of it all. I imagine that’s how Kasparov felt back then, that’s how Lee Sedol felt and now that’s how Terry Tao feels.

The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.

A fundamental difference being that no one was actually paid to find good moves in chess and go like they are to solve math problems and write code. You're comparing the digital camera and the automobile.
loading story #49134034
I am duly impressed by the powerl of the nameless internal AI, but not a single human contributor's name listed anywhere? Did someone at least make this model a coffee?
> The results were achieved by an internal version of Astra, our next major model.
surely some human regularly typed "think deeper, make no mistakes".
What happens when OpenAI et al stop being open about these things, and just pack it into the training?
Not much point to pure math being kept secret, in all honesty. There isn't really industrial value, its only purpose (to them) is showing off their model's capabilities. More realistically they'll just stop paying for it.

Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.

loading story #49132969
What does this even mean lol. These are not solved questions. The solution never existed.
{"deleted":true,"id":49132909,"parent":49132392,"time":1785578640,"type":"comment"}
loading story #49139716
On the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill.

Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.

For OpenAI, research is marketing. I’m sure they’ve got plenty of budget for that.
No such thing as free, even internally at a company. All such use of resources is accounted for, assigned a dollar value and billed to some department. Someone ran the numbers and figured that whatever they spent on these GPU cycles was worth it.
Presumably it's a rounding error compared to their full output, and they're making sure they have enough compute set aside for research by limiting public models. The more datacenters they build the less they have to limit them.
RL training can use all of them - idk what needed means.
I love how people come up with creative ideas to prove the bubble. This one is even more ridiculous - that OpenAI had spare compute to advance mathematics proves that data centres will not be needed. WHAT.

If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.

loading story #49136448
loading story #49137983
loading story #49136543
loading story #49136460
How much do you all think it would cost to "buy" these advances from PhDs, practicing scientists?
This isn't really a productive way to think about these things, IMO. It's quite possible it would take hundreds of years for any specific group of PhDs to solve them. Or one individual PhD could have the correct flash of insight and solve it in a month. There's absolutely no way to predict this, besides trying to gauge the apparent simplicity of the proof or counterexample (which is likely to be misleading). Until someone actually runs an experiment like this it's not a viable metric.
loading story #49132327
loading story #49132344
A friend’s PhD advisor has been chasing non-sofic groups for 25 years (and was shown a preprint of the results by openai to verify them). He believed a solution would be Fields-worthy

This was not a problem that was for sale

You couldn't. PhDs have been working on these problems for decades. It wasn't for lack of trying that none of them could figure these solutions out!
loading story #49133403
loading story #49133754
loading story #49133251
loading story #49158814
loading story #49141147
loading story #49136495
loading story #49133695
loading story #49133485
{"deleted":true,"id":49132325,"parent":49132058,"time":1785572706,"type":"comment"}
I don't know why, but when I saw the source of this particular headline it reminded me of the album title 26 Mixes for Cash
Ambient 0: Math for Airports
loading story #49133108
loading story #49136189
loading story #49159543
loading story #49139434
In a way the most remarkable thing about this is that it isn't even at the top of the HN homepage. Even if this is a step up from what we've seen before, we're no longer astonished by the idea that AI can make significant advances in mathematics and computer science.
This is not at the top as it is actively flagged by people that can't psychologically cope with the advances of AI. Hacker News is no longer a web site of an elite.
> people that can't psychologically cope with the advances of AI

Yes. And there are many of them. I wonder what would help them come to terms with it. Seriously, people are going to be grieving over this. Loss of identity, loss of social standing, ideas of entire future lives that will now never happen. The greatest crime people may hold AI guilty of is taking away their dreams.

loading story #49137391
loading story #49134221
loading story #49133383
loading story #49137606
loading story #49137896
loading story #49139510
loading story #49136774
loading story #49139375
loading story #49133349
loading story #49137087
loading story #49141171
loading story #49136427
loading story #49133566
loading story #49138278
loading story #49135021
loading story #49136471
I don't think that's it. Multiple or my friends from the target audience (academic mathematicians) admitted to scrolling past because the title made it sound like a review of last month's contributions, instead of 10 new ones.
loading story #49135658
What about AI research itself? Is OpenAI close to automating its human staff out of a job?
> What about AI research itself? Is OpenAI close to automating its human staff out of a job?

It's more like they've already automated the parts of the jobs that the humans most closely thought of as the "their job"

They and Anthropic have indicated that the models are substantially augmenting the research and doing large amounts of work autonomously at this point. Here is one of the many blog posts on it [1]. Many people would dismiss this as "marketing" so take it for what you will.

My guess from following this stuff quite closely is that these companies are still a couple years away from fully autonomous research staff.

[1]. https://www.anthropic.com/institute/recursive-self-improveme...

Yes, but they wouldn't publish that bit lest other companies steal the ideas.
It’s also very very divided (x companies, oss vs not and other interests)
This is one of the most impactful mathematical publications in history by all accounts

I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient

No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.

Honestly just a bit burnt out on posts like this
loading story #49160586
loading story #49158667
loading story #49158561
loading story #49158616
the real milestone isn't that AI solved ten math problems, it's that we now need a press release to tell us which ten problems count as important
I asked ChatGPT and it told me these aren’t not important /s
[flagged]
To clarify a bit due to the downvotes : this is not a research paper from a startup or a public frontier lab, it is just PR from a corporation, thus yes an advertisement. Downvote all you like it's still of no value.
loading story #49134790
loading story #49133318
> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work.

AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.

Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.

Sorry, OpenAI's take is correct here. If you're not convinced, here is how they prompted LLM: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... [0]

A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

[0]: Not one of the proofs in the linked article, but from OpenAI too.

> A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

I think you're over-estimating what a smarter highschooler could write.

A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:

  Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor:

  repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.

* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)

But why can’t we prompt the LLM “just do math research”? This is what I don’t understand.
loading story #49137034
loading story #49132871
The human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them.

Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.

loading story #49132491
loading story #49137854
I don't disagree with you but there's no need for exaggeration; ain't no high school student writing this:

> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.

which is infact a very important part of the prompt.

loading story #49133097
When you use a crane to do the "heavy lifting" for construction work, do you give full credit to the cranes?
loading story #49132588
Your brain also is physical. Electrochemical gradients flow between physical molecular constructs. Isn't it just chemistry? Do you attribute it to physics or some whole-is-greater-than-the-parts idea?
> Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.

When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.

I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".

> AI has no self-awareness

What is your mechanistic model of self awareness that yields this conclusion?

> It's a tool

Does your model suggest that tools can't have self awareness?

loading story #49132430
loading story #49132747
loading story #49132636
A better analogy would be a manufactured object, say 3d printed for simplicity. The 3d printer is given an input, and an object manifests itself after some time. We say that the creator of the object is the person turning on the machine, sending the data, and collecting the object. Not the machine itself.
loading story #49132337
loading story #49158456
loading story #49159127
loading story #49160428