Muse Code and Muse Spark 1.2
https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2Does your screen have the message that you’re getting free tokens?
https://platform.openai.com/settings/organization/data-contr...
Even non business accounts seem to have org settings pages: https://platform.openai.com/settings/organization/data-contr...
But even with all data sharing enabled, I’m not seeing the free tokens message there.
Do others see a free tokens message at https://platform.openai.com/settings/organization/data-contr... after enabling all sharing there?
[1] https://platform.openai.com/settings/organization/data-contr...
you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...
But I cannot find this "variant" in OpenRouter.
"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?
Claude: I'm not going to help build this one."
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
This is a bit vague. What sort of games with what technology?
Benchmarks are one data point, not the only one, but the easiest one to compare.
Cherry picking the benchmarks you present is where the falsehoods lie.
No, you really can't. This rhetoric on here is so profoundly boring and tired by now. People have been saying this noise about benchmarks for time eternal, usually because their pet didn't win.
Meta knows their models aren't as good -- demonstrated by their benchmark performance -- and their value proposition is a much lower price.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
Hah, someone has https://www.felonybench.com up and running now.
Such mediocrity makes them the best of the mostly-terrible bunch.
https://futureoflife.org/ai-safety-index-summer-2026/#scorec...
There’s nothing wrong with that, given the landscape we’re all living in.
One of my professors told us about the time he did a request to Facebook to send him all his data. By law they had to send it on paper. They brought it in a big truck.
All the stuff he'd deleted was still there, just with "(deleted)" next to it.
They have a lot on people without Facebook accounts though, because their tracking stuff is all over the web.
I always found it weird that Instagram gives me much better ads than Google does... Google should know much better!
Any insiders know how Muse Code is doing internally?
> If there were, do you believe it would be in their interest to answer this publicly?
If it were being adopted like gangbusters in their organization, sure!
So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...
By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.
Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.
I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
Wasn't the previous one us only? This is probably the biggest part of the post
Anyone know if muse code is open source?
It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.
I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.
It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).
I mean, when Meta's engineer is creating some new DINOv4 or Segment Anything, with all the scaffolding around it, do they train on that?
After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.
Think twice before falling for this announcement and ask yourself what they are not telling you.
The use traces must be crucial to functionality which is why they’re keeping prices so low.
This is my new LLM benchmark when a provider cant use their own models to target a new platform, it cant be that good.
If allegedly a LLM could write a compiler or port a runtime, then this should be trivial to port. But they are not doing it.