Hacker News new | past | comments | ask | show | jobs | submit
For anyone wondering “how slow is this?”

IIUC, Kimi K3 on RTX 6000 Ada (48GB) takes 292 s/token

https://github.com/lyogavin/airllm/releases/tag/v3.1.0

loading story #49160834
I do think these “run a bigger model than will fit in VRAM” projects are necessary steps, but are they functionally useful or helpful to anyone currently? For example, is anyone out there running a big Qwen for coding on a 16-32GB machine with these techniques?
loading story #49160910
loading story #49155867
loading story #49157966
loading story #49160712
loading story #49156383
loading story #49159708
loading story #49158978
Ahaha thank you, I naively assumed the unlabeled graph in the readme was tps, not spt!
loading story #49156196
I wonder what this measures in J/token.
loading story #49156171
How many is that in tokens per Scaramucci?
loading story #49160002
that's 0.003 tokens/second. To get an hour's work done that's normally 30 tokens/second (108k output tokens in an hour) will take 416 days at this rate. And if you're using 100 watts, during that time you will spend $124.61 in electricity, as well as not being able to use your device for something else, plus the noise and heat from your device.

For $124, on Moonshot's official Kimi K3 API rates ($0.30 per 1M cached input, $3 per 1M fresh input, $15 per 1M fresh output), you can purchase 42 million fresh-input tokens, or 8.3 million generated output tokens, in whatever mix you want.

So what you get is 80x more expensive and you wait 416 days to get it.

{"deleted":true,"id":49155713,"parent":49155253,"time":1785764554,"type":"comment"}
loading story #49158749
Hah I was looking for it and couldn't work out how many years/token. 292s is pretty good.