Hacker News new | past | comments | ask | show | jobs | submit
Proof?

In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.

For customer support I don't think models have gotten better since gpt-4.1. The class of small models, with limited to no reasoning, that need to handle a complex issue with a touch of empathy, has not improved much.

I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following maximizing models seem to make worse free-form agents, but they're really all that some domains need.

I understand the point (I don't agree with it; tool calling has gotten much better/reliable and that is very important for customer support) but consider: If you can get same for a lot less, that's an improvement. If we found a way to supply fresh water and electricity for -90% cost after 2 years, that would be fantastic.

You can do many more things, when stuff is cheaper, even if the stuff were otherwise unchanged.

you are working on coding. they are working on things like "creative writing" remember that gpt 4o was popular among those who had ai as a romantic partnet?
> remember that gpt 4o was popular among those who had ai as a romantic partner

I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.

Dude it's not a system prompt, it's the training.
gpt4o & associated parasociality is considered an alignment failure and is actively trained out of the model, so that is a terrible example of regression
sycophancy

It wasn't "better" it was better at kissing your ass which matches what a lot of people want in a partner.

Well that's on purpose lol. OpenAI does not want you falling in love with their chatbot and have been deliberately training it to be less romantic.
There have been several cases of suicide and self-harm related to 4o, AI psychosis is a real risk and will probably be in the DSM
Do any of the big AI companies have a model that are good at tasks that require learning?

For example, every day people teach teenagers how to drive and with only dozens of hours of practice, they are on the road.

is this not essentially what ARC-AGI-3 is? i agree that in-context/continual learning is somewhere the models are still mostly weak at