I think the test above is about tool calling... That's how I read it.
The issue here is known as "out of distribution detection" in the old-timey classification world.
I am not sure how a micro model will fundamentally solve it. Would love to understand what dannyw and team did there?