9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
https://tfmbot.com is the link (discord and source links on the splash screen).
The results are fucking incredible to the point where people in discord are stating "I'm surprised this is working so well". I am too.
I feel like there's a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I've been in the industry for over 25years, highly respected and can't fathom the "AI dumb lololol" type of comments i see on HN. AI is superseeding all other ways to develop.
I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.
Pros know these are lower cost models.
In the meantime, I have cancelled my Anthropic subscription...
I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.
I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.
Then I ask each model to review and critique the other proposals.
By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.
Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...