Hacker News new | past | comments | ask | show | jobs | submit
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
I dare not ask about the cost, having burned $60 on a task running for 1h 16min once.
I’m on the $100/month subscription; this session took about $500 in token-equivalent costs.

(Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)

loading story #49948243
loading story #49948210
loading story #49948385