Hacker News new | past | comments | ask | show | jobs | submit
even 30B model is too large to large on local device (low end). meta should provide free hosted model api to use it.
Meanwhile those of us with 128GB RAM plus some VRAM don't have any good modern (last 8 months) open weights models to make use of all that. I don't care if it would run 5 tok/s, I want a smarter model than Qwen3.6 which avoids loops and can handle more context than 80k before crashing.
why don't you use the quantized version of kimi-k3