This seems really interesting - I was curious about this line from the website.
“The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”
How does soup auto tune the hyper parameters and make some of these more complex training decisions?
[flagged]
loading story #49172065