Hacker News new | past | comments | ask | show | jobs | submit
If you're dropping thousands on API tokens, you're going to be slowed down at least 10x trying to do everything on a single MBP.
But you could grab a 5090, and paired with some DRAM for MoE offloading of bigger models, and be a happy camper with 1.8TB/s of memory bandwidth.

Or just use Luna honestly. Worth considering if you’re ok with hosted APIs.