It doesn’t believe it’s running on that chip, it’s arguing with me
It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
What would tokens/sec performance look like for a reasoning model? An order of magnitude slower?
Reasoning models are the same speed. They’re just post trained with RL to do CoT inside tags like <thinking></thinking> before a tag like <response></response>
There’s no difference in the inference implementation, parameter count, or speed.
There's a difference in the latency distribution between when you submit a query and you see the response, which is what the comment is (clumsily) asking about.
But yeah, there are a lot of factors, so it's hard to answer, and tokens/s isn't the right question.
Which model? Or how many active parameters?
{"deleted":true,"id":49202817,"parent":49202775,"time":1786051660,"type":"comment"}
loading story #49203580
loading story #49203916