Hacker News new | past | comments | ask | show | jobs | submit
I am running it on a single RTX 3090 (24GB VRAM).

Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...

It uses an order of magnitude less VRAM at longer contexts which is a huge advantage over Qwen 3.6 27B

Seems like that's the tradeoff with this model. Close to 27b intelligence while using less vram.