Hacker News new | past | comments | ask | show | jobs | submit
You are making the mistake of classifying models on a single linear axis, or even a multi axis basis set of all benchmarks. That just isn’t true. Each model is unique in its skills and capabilities and the way it approaches problems, in a way that is not represented in benchmarks. Fable is better at reviewing things. I don’t know how to explain it well but it is true. I trust Fable to do thorough reviews (sometimes too thorough) and to present its information in a dense but ordered way. Its output is equivalent to what you used to get from security firms doing code reviews. Having Opus do the work, and have fable do reviews (of the plan and implementation) is a good combo.