shipwithmuse

Catalog / Use case

Benchmarks & research

About 130 benchmark results and research notes on Muse Spark and Muse Glimmer: arena rankings, index scores, head-to-head tests and architecture teardowns.

139 builds · page 1 of 4

Alex Volkov

@altryne

Here's my full in depth review of @muse (muse.ai) Meta's new FREE AI agent with it's own browser, computer and connectors to anything from FB marketplace to Apple health! I go into setting it up, have it buy stuff securely with @link, and 3 great usecases that

X post · Benchmarks & research· ♥ 77

Hands-on Muse review with Link purchases

thehype.

@thehypedotnews

muse glimmer 30b vs nemotron 3.5 lightning – on agentic coding two 30b open-weight agent models, released three days apart, both sold as the model you leave running. same harness, same tools, same tasks, two of them: repair a broken landing page, then write breakout from a spec

X post · Benchmarks & research· ♥ 32

Muse Glimmer vs Nemotron on agentic coding

wh

@nrehiew_

Muse Glimmer was trained directly logit distilled from Muse Spark. This means that there isn't a traditional 'base model' in that it was trained from the start on agentic traces. Super cool and haven't seen this approach in a while Welcome back GPT OSS

X post · Benchmarks & research· ♥ 336

Glimmer was logit-distilled from Spark

Arena.ai

@arena

Muse Spark 1.1 has entered the Code Arena: Frontend at #9! Muse Spark 1.1 reshapes the cost-performance Pareto Frontier by scoring 1541 at a blended $3.5M ($1.25 per input MToken, $4.25 per output MToken). This is frontier performance at a fraction of the price. Congrats to

X post · Benchmarks & research· ♥ 302

Spark 1.1 in Code Arena: Frontend

Arena.ai

@arena

Muse Spark 1.3 Max by @AIatMeta has reshaped the Pareto frontier for Code Arena: WebDev! Meta's latest model at Max reasoning is doing something interesting on the Arena Pareto frontier: it's the only model holding down the wide price band between Qwen3.8-max ($5/MToken) and

X post · Benchmarks & research· ♥ 334

Spark 1.3 Max on the WebDev Pareto frontier

Artificial Analysis

@ArtificialAnlys

Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B (Reasoning), 1141 for Gemini 3.5 Flash-Lite, and 1004 for Kimi K2.5 (Reasoning), with Terminal-Bench v2.1 (52%) also behind Qwen3.6 27B (61%). The

X post · Benchmarks & research· ♥ 34

Where Glimmer trails on agentic evals

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@JacobStephens2

@JacobStephens2

A reproducible evaluation of a stochastic defect where Muse Spark 1.1, under some agent-like request envelopes, wrote to claude-smoke.txt when asked for muse-smoke.txt; a neutral control kept 80/80 filenames intact.

GitHub · Benchmarks & research

Muse Spark 1.1 filename-substitution repro

U

PathfinderTactician

u/PathfinderTactician

I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context, whilst Qwen3.6 27B can only run at 147,500 context - full GPU offload in both cases. Both models have been coding on an enterprise-grade web application. Detailed report of each model (warning - includes AI generated content): Diagnostic quality - comparable. Both have shown genuinely good root-cause work when they apply themselves. Qwen found coding issue and worked to fix things cleanly. Muse Glimmer correctly traced bugs and even caught something that a Frontier model missed after more than 10 rounds of review. Neither one is weak at diagnosis. Implementation reliability - Qwen ahead. Qwen did introduce real regressions into the coding along the way (eg. severe zone-scope refactor regression, and case-sensitivity regression) but each one eventually got fixed properly once caught, usually within one or two corrective rounds. Muse Glimmer did land fixes that were clean and verified true to spec. However, when working in a complex environment exceeding 200k context, Muse Glimmer fa

Reddit post · Benchmarks & research

BF16 Muse Glimmer vs Qwen3.6 27B on real code

AI at Meta

@AIatMeta

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth

X post · Benchmarks & research· ♥ 637

WildArtifactBench agent eval preview

Artificial Analysis

@ArtificialAnlys

Last week the Intelligence Index vs Cost Pareto frontier moved out substantially. Claude Fable 5.1, Muse Spark 1.3, and GPT-6 Astra each set a new point in efficient intelligence Link to analysis: artificialanalysis.ai/#intelligence-…

X post · Benchmarks & research· ♥ 732

Spark 1.3 moves the cost Pareto frontier

@accretional

@accretional

Unofficial knowledge base and tools for Meta's Muse Spark, starting with a root-cause write-up of Codex tool-calling failures with Muse Spark 1.1.

Skill · Benchmarks & research· ★ 3

awesome-muse-spark

Cline

@cline

Muse Spark 1.1 just launched and it's their most capable coding agent model yet. On Terminal-Bench 2.1 it scores 80.0%, in the same cluster as Opus 4.8 (82.7%) and GPT 5.5 (83.4%). Use it in Cline with the Meta API!

X post · Benchmarks & research· ♥ 191

Spark 1.1 scores 80% on Terminal-Bench 2.1

@barkleesanders

@barkleesanders

A static teardown of Meta's Muse macOS app (codename Endo) finding a full computer-use agent in the public build, switched off by five server-delivered feature flags.

GitHub · Benchmarks & research

Muse for Mac 'Endo' teardown

W

@intheworldofai

@intheworldofai

WorldofAI benchmarks Muse Glimmer 30B on consumer hardware with a self-built test harness and compares it against Qwen 3.6 27B.

Video · Benchmarks & research

Muse Glimmer 30B vs Qwen 3.6 27B, fully tested

Iam_

@SPAC89

Muse Spark 1.3 Ultra Contributor vs Fable 5.1 xHigh Same prompt, both ran for roughly 2 hours The prompt had a self improvement rule: if the independent judges scored the result below 9.5/10, it had to keep improving and try again, What shocked me most was that Muse Spark just

X post · Benchmarks & research· ♥ 1.2K

20 self-improvement loops for under $1

Evergreen Capital

@evergreencap3

I tested $META's Muse Spark over the last few hours and came away net positive. 3 main takeaways: 1) Quality: It's a very good model. Not quite frontier but good. It showed comparable performance vs Opus 4.6 across web data search, PDF parsing, and general

X post · Benchmarks & research· ♥ 309

Investor's hands-on test of Muse Spark

A

@altryne

@altryne

ThursdAI hosts compare daily-driver assistants, then David Pawlan, who ran 273 tests across 23 assistants for Assistant Benchmark, walks through how Muse stacks up.

Video · Benchmarks & research· ♥ 18

ThursdAI: Grok Bot vs Meta Muse vs Instinct

Derya Unutmaz, MD

@DeryaTR_

Here is a crazy Jev @typesafeai example that I’m betting nobody has thought about: I asked @Muse from @Meta to use Jev to select the top 100 unanswered questions in immunology from 10,000 literature-grounded candidates. A few minutes later, it came back with some of the best

X post · Benchmarks & research· ♥ 427

Top 100 open immunology questions via Muse + Jev

Trapit Bansal

@TrapitBansal

We entered Meta models in five international STEM Olympiads, as an uncontaminated eval of their reasoning capabilities. Three of these were live participations and graded officially. The models achieved gold-medal results in all five! 1/5

X post · Benchmarks & research· ♥ 219

Olympiad golds as an uncontaminated eval

M

research.meta.ai

research.meta.ai

Meta's engineering write-up on Muse security: isolated VMs, a separate Sentinel permission authority, credential surrogation and layered prompt-injection defenses, with bug bounties up to $300,000.

Resource · Benchmarks & research★ Pick

How Meta built safety into Muse

@tanishq-dubey

@tanishq-dubey

A reproducible Apple Silicon harness that runs six fixed quality tasks against MLX quantizations of Muse Glimmer 30B and records scores, tokens per second, peak memory and load time.

GitHub · Benchmarks & research

Muse Glimmer 30B MLX benchmark harness

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

AI at Meta

@AIatMeta

To understand whether we're making genuine progress on reasoning, we entered our AI models in five STEM Olympiad competitions. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam 🏅 International Physics Olympiad (IPhO): Perfect score, theory exam 🥇 ht

X post · Benchmarks & research· ♥ 1.3K

Muse models take five STEM Olympiads

Mark Zuckerberg

@finkd

Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.

X post · Benchmarks & research· ♥ 5.5K

Muse Voice Transcribe

P

aicodingdaily.com

aicodingdaily.com

Povilas Korop scores Muse Spark 1.3 on six real coding projects (Laravel, React-TS, PHP, Flutter, Go): Max effort ranks #26 with 48.43/60 at about $0.01 per prompt and 3:42 per prompt.

Resource · Benchmarks & research

AI Coding Daily benchmark of Muse Spark 1.3

J A Z I I

@notjazii

meta muse just mogged fable 5.1 tested fable 5.1 and muse spark 1.3 with same prompt at highest reasoning available and results came out really different > muse spark 1.3 completed task in one minute and costed almost nothing > fable 5.1 completed task in 70 minutes and costed

X post · Benchmarks & research· ♥ 266

Same prompt: 1 minute vs 70 minutes

khaled

@eltokh7

Muse Spark 1.1 seems like a very good model. I tested it on Stata Benchmark and it ranks 4th (!) beating Opus 4.7/4.8, tied with GPT 5.5. I like this benchmark because it often catches seemingly good model performing poorly on OOD tasks

X post · Benchmarks & research· ♥ 17

Spark 1.1 ranks 4th on Stata Benchmark

T

taylorarndt.substack.com

taylorarndt.substack.com

Taylor Arndt tested Muse with VoiceOver and found non-standard text fields and missing heading structure in chats, plus connector gaps such as iCloud email that kept it out of her work.

Resource · Benchmarks & research

I tried Meta's Muse agent: an accessibility review

Arena.ai

@arena

Exciting news: Meta’s Muse Image just claimed #2 in the Image Arena! Muse Image from @AIatMeta now ranks second only to OpenAI's GPT Image 2, outperforming Nano Banana, Grok Imagine, MAI Image, and many other leading image models. It holds #2 across the board: Text-to-Image,

X post · Benchmarks & research· ♥ 1.5K

Muse Image takes #2 in Image Arena

B

blog.box.com

blog.box.com

Box's Complex Work Eval finds Muse Spark 1.1 up to 5-6 points above the top-tier composite on structured work, and nearly 30 points ahead on cost-optimization analysis.

Resource · Benchmarks & research

Box eval: Muse Spark 1.1 on real enterprise work

U

DerTomsn

u/DerTomsn

Ornith does really well. TielCoder (https://llm-bench.io/benchmarks/cmt7kp2zj002r01lcmpchvlko) might be even a bit better in coding. Will give it a try soon. Details of the comparison see here: https://llm-bench.io/compare/runs?runs=cmt6ecf8g000001p45vwzux53%2Ccmt6ergk5000701p41hqdyy78%2Ccmt6f2oob000e01p49o9592cb%2Ccmt6fqddm000l01p4l1vm7skd

Reddit post · Benchmarks & research

Four-way local model comparison incl. Glimmer oQ8e

Morgan

@morganlinton

Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done at @VulcanBench and it definitely looks like Muse has some serious issues at lower effort levels. I was able to run Astra through my v4 eval suite across every

+1

X post · Benchmarks & research· ♥ 99

VulcanBench run of Muse Spark 1.3

elie

@eliebakouch

extremely exciting to see meta getting back into open weight models with a 30B dense first, and soon muse spark 1.2 they used knowledge distillation from muse spark, architecture wise it's similar to gemma 4 (which is llama 3 + swa (again!) + vision encoder), with scale free QK

X post · Benchmarks & research· ♥ 464

Glimmer architecture teardown

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Dan

@DanDr1s

Meta’s Muse Spark 1.3 just scored 62 on Artificial Analysis’ Intelligence Index. That ties Claude Fable 5, but Muse costs 8x less for input and nearly 12x less for output. It also scores above GPT-5.6 Sol, Grok 4.6, Kimi K3, and Gemini 3.8 Flash. Meta is suddenly in the top

X post · Benchmarks & research· ♥ 223

Spark 1.3 matches Fable 5 at a fraction of the price

pilvar (Philippe Dourassov)

@pilvar222

We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers an average of 19/32 CVEs. In comparison, Grok 4.6 scores 23.3/32 - When pooling the results of 3 runs (pass@3), Muse scores 24/32. DeepSeek V4 Pro gets

X post · Benchmarks & research· ♥ 116

Muse Spark 1.3 on a CVE rediscovery benchmark

Arena.ai

@arena

Muse Spark 1.2 (xHigh) by @AIatMeta is now in Agent Arena, with a net improvement of +2.1%! Agent Arena measures models on millions of real-world, long-horizon agentic tasks. We use causal tracing methodology to measure a model's net improvement, indicating how much it improves

X post · Benchmarks & research· ♥ 259

Spark 1.2 joins Agent Arena

@ezcat207

@ezcat207

A curated, source-linked catalog of 156 real things people have done with the Muse personal agent, grouped into 16 categories with every entry linked to its original X post.

GitHub · Benchmarks & research

Awesome Meta Muse use cases

A

@AICodeKing

@AICodeKing

AICodeKing tests Muse Spark 1.3 and Gemini 3.8 Flash on frontend, 3D, SVG, maths and agentic coding, and flags file-overwriting issues.

Video · Benchmarks & research· ♥ 220

KingBench 3: Muse Spark 1.3 vs Gemini 3.8 Flash

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

About this shelf

This shelf gathers about 130 scores, comparisons and research notes. Independent leaderboards come first. Artificial Analysis has scored every Muse Spark release on its Intelligence Index, putting 1.3 at 62. Arena ranked Muse Image #2 in its Image Arena, Design Arena put Spark 1.3 at #1 on Website Arena, and Vals AI placed Spark 1.2 in its Vals Index top 5.

Then there are hands-on tests. AICodeKing's KingBench 3 pits Muse Spark 1.3 against Gemini 3.8 Flash and flags file-overwriting issues. Matt Johnston's blind coding gauntlet found a strong Halo build but broken XCOM and Diablo attempts. thehype had Muse Code and three rivals build landing pages and fix their own bugs in a real browser.

The research side covers how the models are made. Sebastian Raschka and elie broke down Muse Glimmer's Gemma-like architecture, and another post notes it was logit-distilled directly from Muse Spark. Meta's own benchmark claims are labeled as Meta's, and third-party numbers link to the people who ran them.

Frequently asked

+How good is Muse Spark 1.3 at coding?

Meta reports 75.4 on DeepSWE v1.1 and 88.8 on Terminal-Bench 2.1. Independent tests on this shelf are more mixed: DataCamp measured a cost increase on its tasks, and some coding gauntlets found failures on complex builds.

+Where does Muse Spark rank on Artificial Analysis?

Artificial Analysis scored Muse Spark 1.3 (max) at 62 on its Intelligence Index, behind Claude Fable 5.1 and Claude Opus 5. Earlier releases scored 52 for the original, 51 for 1.1 and 54 for 1.2.

+How does Muse Glimmer compare with Qwen?

Most comparisons here find Glimmer somewhat behind Qwen 3.6 27B on raw scores but using far fewer tokens. One Reddit analysis of Artificial Analysis data puts it at 92% of Qwen3.6 27B's index score with 2.9x fewer tokens.