But why can’t we prompt the LLM “just do math research”? This is what I don’t understand.
100% agree. If the models are so capable that they're advancing math, it doesn't seem like a stretch to expect they should be able to determine with "doing math research" entails and the best way to use their capabilities towards that end. Why do we need to hand hold the models by telling them to do parallel research, keep threads independent, etc.
If there aren't thousands of TPUs doing that [0] right now I'd be quite surprised.
[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".