Hacker News new | past | comments | ask | show | jobs | submit
Thank you for your response. Part c was especially insightful. Quite a smart way to do it and makes the possibilities of post training seem almost endless. Makes sense that you just need more time and compute.

A positive feedback loop then. RL->better model->better RL pipeline -> better model…

And we’ve only recently started getting into the much better RL pipelines