Hacker News new | past | comments | ask | show | jobs | submit
> Small llms are still way more efficiently server on big GPUs.

Yes, but the privacy aspect means that for many, many applications slower local will still be preferable to faster remote so long as the actual model performance is the same.