Hacker News new | past | comments | ask | show | jobs | submit
Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?
This isn't latency bound, it is trivially parallelize. So you want to run it on the most efficient compute you have, not the fastest.
loading story #49299471
That’s the whole point, just cost and compute limitations in your way (mostly).