https://xcancel.com/finkd/status/2086755195535413696
"... Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model..."
This is bigger news - good for self hosting enthusiasts and a strategically sound move for Meta. Any push towards 'anti Chinese' models will directly benefit Meta as the competition on the frontier open-weights American models is almost non-existent. Meta will have no problem being #1.
It wouldn’t surprise me if Meta does become the #1 American open weights provider, but I doubt it’ll be easy. Thinking Machines has a good amount of talent behind them as I understand it and their Inkling model was decent (admittedly not great though). I think Meta’s biggest problem is going to be internal as there’s be a bunch of headlines posted here on their talent retention issues.
Poolside Laguna was quite good too (if you look beyond some of the teething issues).
Had Deepseek V4 Flash 0731 not launched, their latest Laguna release was really intelligent at non-coding tasks and it would have been my go-to model for my local workloads.
For me Laguna frequently slightly corrupted text then it would be unable to notice the difference and get stuck making the dumbest conclusions. Thinks like typoed directory or function names. It was a great model other than that, but I ended up just going back to Qwen3.6
Yeah, but if it's huge, how many can run it? Many folks struggle to run 200B+ models
Two DGX sparks will run the full Deepseek V4 Flash full. (It is definitely expensive, but relatively easy and compact; and extremely power efficient)
Many people don't have $5000 for DGX Sparks. With that said, it doesn't take much to run Deepseek. I run it on a sub $1000 system 128gb 2 3060 at 6-7tk/sec and then on a $1000 system with 10 MI50 GPUs built when the price was cheap.
What about Inkling? It's a quite large model that for some reason isn't discussed much.