Muse Spark benchmarks: Meta's numbers vs independent tests
Meta says Muse Spark 1.3 scores 75.4 on DeepSWE and uses 25% fewer tokens. Independent tests found a 12% cost increase, a #6 index rank and mixed results.
Meta reports that Muse Spark 1.3 leads DeepSWE v1.1 at 75.4, ties GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8, and uses about 25% fewer tokens than 1.2. Independent testing is less tidy: eesel puts it #6 on the Artificial Analysis Intelligence Index and behind Opus 5 on four of six agent evals, DataCamp measured a net 12% cost increase on real tasks, and MindStudio got weak results on an open-ended game build. On bug fixing and long context, though, some outside tests back Meta up.
Here's what each source says, and how to read the gap.
Meta's reported numbers for Muse Spark 1.3
These come from Meta's Spark 1.3 announcement, as summarized by MarkTechPost. They are Meta's claims.
| Benchmark | Muse Spark 1.3 | Comparison Meta gives |
|---|---|---|
| DeepSWE v1.1 | 75.4 | Opus 5 74.0, GPT-5.6 Sol 72.7 |
| Terminal-Bench 2.1 | 88.8 | Ties GPT-5.6 Sol |
| SWE-Atlas Codebase QnA | 59.4 | |
| MRCR v2 (256K–512K) | 98.5 | |
| MRCR v2 (512K–1M) | 98.1 | |
| OSWorld 2.0 (max) | 66.9 |
Meta also claims about 20% fewer tool calls and 25% fewer tokens than Spark 1.2. The model has a 1,048,576-token context and a max reasoning mode.
Independent results
Leaderboard rank. eesel reports Spark 1.3 at #6 on the Artificial Analysis Intelligence Index, leading long-context and coding rows but trailing Claude Opus 5 on four of six agent evals. MindStudio's write-up says it placed third on Artificial Analysis. The two may be reading different indexes or different dates. We haven't resolved which is current, so check Artificial Analysis directly.
Token efficiency. DataCamp's Josep Ferrer ran three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase. Uncached input fell 19–22%, not 25%.
Open-ended builds. MindStudio's game-clone test produced "a cube shooting at other cubes."
Bug fixing. Paweł Huryn planted 105 bugs across two real repos. Spark 1.3 at max fixed 33, tying Fable 5.1 (high) and ahead of Grok 4.6 and Opus 5 at 27 each. This one supports Meta's coding story.
Security. On a CVE rediscovery benchmark, Spark 1.3 found 19 of 32 on average at pass@1, versus 23.3 for Grok 4.6, and 24 of 32 when three runs were pooled.
Effort levels. Morgan Linton's VulcanBench v4 run across every effort level showed serious issues at lower effort. If you're comparing costs, you're comparing different quality levels too.
Writing. Louis-François Bouchard's team placed Spark 1.3 #24 of 87 models on their internal writing benchmark, up from #31 for Spark 1.1.
Vision. On a calorie-estimation test from food photos, Spark 1.3 led the field at 48% within 20% error. Earlier, Deedy found Spark's visual grounding and counting "far from perfect."
Where Meta's claims hold up and where they don't
| Claim | Independent evidence | Verdict |
|---|---|---|
| Top-tier agentic coding | 105-bug test ties the leader | Mostly holds |
| Long context | eesel says it leads long-context rows | Holds |
| 25% fewer tokens | DataCamp: net +12% cost on 3 tasks | Depends on task |
| Frontier overall | #6 (eesel) or #3 (MindStudio) on Artificial Analysis | Near, not at, the top |
| Open-ended creative builds | MindStudio's cube game | Weak in that test |
The pattern is familiar. Well-specified tasks with a test to pass (fix this bug, answer this question about a codebase) look like the benchmarks. Vague tasks with no spec ("make a fun game") expose the gap. Plenty of catalog builds show strong one-shot games anyway, such as a pirate platformer and AstraVanguard, which passed 25/25 checks at xhigh effort. The difference is often how specific the prompt was.
Why benchmarks and hands-on results diverge
- Effort settings. Some of Meta's figures, like OSWorld 2.0, are at max reasoning, and VulcanBench shows lower effort levels behave differently. Your default run may not match.
- Harness. Spark 1.2 was co-trained with the Muse Code harness. Results in opencode, Cursor or your own loop can differ. muse-spark-anywhere exists because Spark sends a non-standard SSE event some tools don't handle.
- Token accounting. "Fewer tokens" averaged across a suite can still mean more tokens on your refactor.
- Contributor vs standard. Muse Code defaults to the contributor model. Know which one you're measuring.
How Spark has moved since April
The version history matters when you read older tests. Spark launched on April 8 with API access limited to private-preview partners, so the first outside impressions were chat-based. Evergreen Capital spent a few hours with it and found it comparable to Opus 4.6 on web data search, PDF parsing and general tasks, though not quite frontier. Versions 1.1 (July 9), 1.2 (August 5) and 1.3 (September 2) followed. The writing benchmark above shows the direction: #31 for Spark 1.1, #24 for 1.3. A test result without a version number and effort level isn't much use, so note both when you share yours. The full sequence is in the Muse launch timeline.
What about Muse Glimmer's benchmarks
Glimmer's model card claims SWE-Bench Verified 76.0, MCP Atlas 75.5 and GPQA Diamond 83.5. Independent results are split; see Glimmer vs other open models.
How to test it yourself
Pick 10 to 30 tasks from your real backlog, run them at the effort level you'll actually pay for, and track pass rate and cost per pass. Evaluating agent builds has a checklist, and reducing agent costs covers the cost side.
Frequently asked questions
How good is Muse Spark 1.3 at coding?
Meta reports 75.4 on DeepSWE v1.1, above Opus 5's 74.0. An independent 105-bug test had Spark 1.3 at max tie for the lead with 33 fixes. Open-ended builds without a clear spec are weaker.
Is Muse Spark 1.3 really 25% more token-efficient?
Meta says so, averaged over its tests. DataCamp's three-task test found a net 12% cost increase because one refactor used 70% more tokens. Measure it on your own work.
Where does Muse Spark rank on Artificial Analysis?
Sources differ. eesel reports #6 on the Artificial Analysis Intelligence Index, and MindStudio says third. Check the live leaderboard.
Is Muse Spark better than Claude Opus 5?
On Meta's DeepSWE numbers and one planted-bug test, yes. On eesel's reading of agent evals, Opus 5 leads four of six. It depends on the task.
Numbers throughout are as reported by the build authors or by Meta, not verified by shipwithmuse. Official documentation lives at muse.ai/platform.
ChatForm
Tgmlabs