Hacker News new | past | comments | ask | show | jobs | submit
Sorry, OpenAI's take is correct here. If you're not convinced, here is how they prompted LLM: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... [0]

A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

[0]: Not one of the proofs in the linked article, but from OpenAI too.

> A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

I think you're over-estimating what a smarter highschooler could write.

A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:

  Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor:

  repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.

* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)

The human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them.

Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.

What gives the intention and ability to the human?
Are you playing dumb? Using power tools to build furniture is very different than using an ai robot to carve a statue or whatever.
But why can’t we prompt the LLM “just do math research”? This is what I don’t understand.
100% agree. If the models are so capable that they're advancing math, it doesn't seem like a stretch to expect they should be able to determine with "doing math research" entails and the best way to use their capabilities towards that end. Why do we need to hand hold the models by telling them to do parallel research, keep threads independent, etc.
If there aren't thousands of TPUs doing that [0] right now I'd be quite surprised.

[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".

I don't disagree with you but there's no need for exaggeration; ain't no high school student writing this:

> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.

which is infact a very important part of the prompt.

the fact that such things have to be explicitly in the prompt points to the fact that the underlying system is still far from where it needs to be (basically, lacks basic understanding what a proof is)
No, that's not what this is. This is a warning to the LLM that coming back with partial results is not good enough.

Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusually cruel and unreasonable advisor who tells them, do not dare to talk to me until you've fully solved the problem. This paragraph is exactly that. It's there exactly because the underlying system is smart enough to know that real mathematicians do not work like that.

If that was the case, that elaborate listing of all things that might look like a proof to a naive student, but are actually not proofs (and not even just partial results, but fundamental misunderstandings of what constitutes a proof) could have been easily and equivalently replaced by 'I am interested only in a full proof, don't bother me with partial results'. Yet, they were not.

To a real mathematician you would not have to list those explicitly, he/she would have understood that implicitly from 'give me a full proof'. That listing makes sense to say only to somebody who pretends to be a mathematician, but has not true understanding of how the math works. The models are getting better and better in this pretension, but prompts like that reveal that it is still just a pretension, not a true understanding.

They don't have to be. At this point, we have multiple results from 3rd parties where the prompts are very basic.

To name a few:

- https://xcancel.com/DmitryRybin1/status/2079904005652893709

- https://archive.ph/2w4fi (https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...)

you cannot be serious
When you use a crane to do the "heavy lifting" for construction work, do you give full credit to the cranes?
Read the prompts in the PDF I link and see if your analogy makes sense in this context :)
Prompt quality should not matter. If a high-schooler operates the crane to lift a ton of weight 10 floors high, should the credit entirely go to the crane?
When I type 56789*23456 into my calculator and get the result I don't claim to have solved the problem, the calculator did it.
loading story #49140758