Hacker News new | past | comments | ask | show | jobs | submit
The post suggests that you need an rtx 5090 use it, which is currently selling for around $5,000 USD. I wouldn't exactly call that "my device", since my device costs about 25% of that for the entire computer.

For the same cost, you could run on a frontier model on a pro plan for two years. The economics dont make a lot of sense for this to me, so I would love some input on why people want to do this instead (privacy, for fun, etc).

If you’re doing breakeven math on subscriptions, consider that your own rig can run 24/7 whereas you will get a fraction of that with sub rate limits. Even if you factor in PG&E residential rates, the breakeven is a lot closer to months for overnight long-running agentic coding a couple times a week.

And in terms of interesting use cases: recently pointed an agent at Blender and gave it vision. That setup can essentially iterate on a scene forever.

It seems exceedingly unlikely that the current Pro plan costs will hold for the next two years. The subsidization train is going to end eventually.
Is the assumption here that inference costs will stay roughly static, or that frontier models will keep getting more expensive quickly enough to offset efficiency gains?

Because I don’t think “the subsidization train is going to end” necessarily means current pricing becomes impossible.

If capital keeps pouring into frontier AI, companies still have an incentive to subsidize access while competing for users and market share. And if that subsidization starts drying up, there’s even more incentive to bring inference costs down by making smaller and cheaper models catch up to today’s frontier capabilities.

So either way, I’m not sure you can extrapolate from the cost of serving current frontier models to what equivalent capability will cost two years from now.

I think like you mentioned, the practical reasons are disproportionately oriented around either privacy (, a clear constrained workload (need to OCR files/transcribe audio, and there isn't really a clear or meaningful reason to switch out the model to chase new incremental gains), or regulatory compliance (e.g. source code, patient data, can't leave the country).
The post does not imply the 5090 is needed, that is just a common reference point.

A single six year old RTX 3090 works great: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...

I fully expect Meta will release other, smaller Muse models in the near future too.

The 5090 is also supposed to be a $2000 GPU, not a $5000 one. The entire market is utterly distorted right now, which will impact cloud inference more and more over time too. They are not immune to the absurdly high RAM prices, so their prices will have to go up over time too until the RAM supply chain goes back to normal.

I'm currently running it on an RTX 3090 (street price ~$1000 USD) with a long context and getting pretty good performance.

Prefill: ~1000 tok/s

Decode: 75-100 tok/s

It'll be far faster on a 5090, but I find the above performance to be acceptable. I've seen some claims that it even works OK on an AMD RX 7900XT (~$500USD)

The performance will be so-so, but you can buy an Intel Arc B70 for $1000. There are definitely ways to get going for less.