If you use structured outputs they’ll usually stick to the program. Not to completely constrain the categories like TFA was saying, but something like
{ rationale, categories }
Where you don’t really care about the rationale but you’re using it as a pseudo thinking for models that don’t support it.Luna is surprising capable and cheap, and I haven’t done this type of thing since before GPT 5 so might not be such a useful trick now