Catalog / Use case

Benchmarks & research

About 130 benchmark results and research notes on Muse Spark and Muse Glimmer: arena rankings, index scores, head-to-head tests and architecture teardowns.

22 builds · page 1 of 1

Mark Zuckerberg

@finkd

Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.

X post · Benchmarks & research★ Pick· ♥ 5.5K

Muse Voice Transcribe

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Iam_

@SPAC89

Muse Spark 1.3 Ultra Contributor vs Fable 5.1 xHigh Same prompt, both ran for roughly 2 hours The prompt had a self improvement rule: if the independent judges scored the result below 9.5/10, it had to keep improving and try again, What shocked me most was that Muse Spark just

X post · Benchmarks & research★ Pick· ♥ 1.2K

20 self-improvement loops for under $1

Alexandr Wang

@alexandr_wang

i find muse spark is very good at data analysis—both finding relevant open-source data and analyzing it. for example, here's my results for analyzing global share of GDP over past century: meta.ai/share/cw54skLB…

X post · Benchmarks & research★ Pick· ♥ 591

Century of global GDP share analysis

research.meta.ai

research.meta.ai

Meta's engineering write-up on Muse security: isolated VMs, a separate Sentinel permission authority, credential surrogation and layered prompt-injection defenses, with bug bounties up to $300,000.

Resource · Benchmarks & research★ Pick

How Meta built safety into Muse

Arena.ai

@arena

Muse Spark 1.3 (xHigh) just landed @AIatMeta back in the top 10 models for Code Arena: WebDev! This release is ~#10 in Code Arena: WebDev with 1623 pts (AutoEval). It’s performance is par with Claude Fable 5 at #8 (1628 pts) and GPT-5.6 Sol at #11 (1616 pts). Muse Spark 1.3

X post · Benchmarks & research★ Pick· ♥ 610

Spark 1.3 in Code Arena: WebDev top 10

Artificial Analysis

@ArtificialAnlys

Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place

X post · Benchmarks & research★ Pick· ♥ 1.2K

Spark 1.2 scores 54 on the Intelligence Index

Vals AI

@ValsAI

Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol.

X post · Benchmarks & research★ Pick· ♥ 773

Muse Spark 1.2 enters Vals Index top 5

Artificial Analysis

@ArtificialAnlys

Meta's Muse Spark 1.1 scores 51 on the Artificial Analysis Intelligence Index and is cost and token efficient compared to its peers Muse Spark 1.1 (xhigh) improves 8 points over Muse Spark 1.0 (43) in three months. It is effectively tied with GLM-5.2 (max), GPT-5.4 (xhigh), and

X post · Benchmarks & research★ Pick· ♥ 708

Spark 1.1 scores 51 on the Intelligence Index

Artificial Analysis

@ArtificialAnlys

Meta is back! Muse Spark scores 52 on the Artificial Analysis Intelligence Index, behind only Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Muse Spark is the first new release since Llama 4 in April 2025 and also Meta's first release that is not open weights Muse Spark is a new

X post · Benchmarks & research★ Pick· ♥ 2.4K

Original Muse Spark scores 52

Arena.ai

@arena

Exciting news: Meta’s Muse Image just claimed #2 in the Image Arena! Muse Image from @AIatMeta now ranks second only to OpenAI's GPT Image 2, outperforming Nano Banana, Grok Imagine, MAI Image, and many other leading image models. It holds #2 across the board: Text-to-Image,

X post · Benchmarks & research★ Pick· ♥ 1.5K

Muse Image takes #2 in Image Arena

AI at Meta

@AIatMeta

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth

X post · Benchmarks & research★ Pick· ♥ 637

WildArtifactBench agent eval preview

AI at Meta

@AIatMeta

To understand whether we're making genuine progress on reasoning, we entered our AI models in five STEM Olympiad competitions. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam 🏅 International Physics Olympiad (IPhO): Perfect score, theory exam 🥇 ht

X post · Benchmarks & research★ Pick· ♥ 1.3K

Muse models take five STEM Olympiads

Artificial Analysis

@ArtificialAnlys

Last week the Intelligence Index vs Cost Pareto frontier moved out substantially. Claude Fable 5.1, Muse Spark 1.3, and GPT-6 Astra each set a new point in efficient intelligence Link to analysis: artificialanalysis.ai/#intelligence-…

X post · Benchmarks & research★ Pick· ♥ 732

Spark 1.3 moves the cost Pareto frontier

Tom Collins

@thetomcollins

$META just went from 3.5% to 45.4% token share on OpenCode in just over two weeks Muse Spark 1.3 being good + free is enough to become the default for most users Default gets you usage → usage gets you data → data makes the next model better Anthropic and OpenAI can’t afford

X post · Benchmarks & research★ Pick· ♥ 738

Meta's OpenCode token share hits 45%

Vals AI

@ValsAI

Meta just released Muse Spark 1.1 and is the new SOTA on MedScribe and TaxEval, taking the top spot from Fable 5 while being 10x cheaper and twice as fast. Meta currently holds the top 2 spots on TaxEval It is also the new #1 on Harvey's Legal Agent Bench, dethroning Grok 4.5

X post · Benchmarks & research★ Pick· ♥ 1.3K

Spark 1.1 tops MedScribe and TaxEval

Arena.ai

@arena

Meta Muse Video just entered the Video Arena at #3. @AIatMeta’s new video model scored 1459 in the Text-to-Video Arena. It outperforms Alibaba’s HappyHorse 1.0 by +30pts and ranks ahead of Grok Imagine, Sora 2 Pro and Google Veo-3.1 models. Meta has now reached the video AI

X post · Benchmarks & research★ Pick· ♥ 611

Muse Video enters Video Arena at #3

Artificial Analysis

@ArtificialAnlys

Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant

X post · Benchmarks & research★ Pick· ♥ 2.6K

Artificial Analysis scores Muse Spark 1.3

Design Arena

@DesignArena

BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362! This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price. Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this

X post · Benchmarks & research★ Pick· ♥ 864

Spark 1.3 tops Website Arena

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Artificial Analysis

@ArtificialAnlys

Meta returns to open weights: Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence Index. It is a 30B-parameter model, and the first from Meta to be released under Apache 2.0 Muse Glimmer (high) arrives 16 months after

X post · Benchmarks & research★ Pick· ♥ 778

Glimmer scores 35 on the Intelligence Index

Paweł Huryn

@PawelHuryn

So, I finally tested Muse Spark 1.3. 2 real repos, 105 planted bugs, find and fix what you can. Original harness and API. Big surprise: Muse Spark 1.3 (max): 33 Fable 5.1 (high): 33 Grok 4.6 (xhigh): 27 Opus 5 (max): 27 Muse Spark 1.3 (high): 19 Meta joined the frontier.

X post · Benchmarks & research★ Pick· ♥ 694

105 planted bugs: Spark 1.3 vs frontier

Sebastian Raschka

@rasbt

Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like architecture design. (“Glimmer” is probably a wordplay on “Spark,” the more

X post · Benchmarks & research★ Pick· ♥ 1.6K

Raschka on Glimmer's Gemma-like design

datacamp.com

datacamp.com

DataCamp's Josep Ferrer ran Muse Spark 1.3 on three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase.

Resource · Benchmarks & research★ Pick

Muse Spark 1.3 tutorial: testing Meta's efficiency claims

About this shelf

This shelf gathers about 130 scores, comparisons and research notes. Independent leaderboards come first. Artificial Analysis has scored every Muse Spark release on its Intelligence Index, putting 1.3 at 62. Arena ranked Muse Image #2 in its Image Arena, Design Arena put Spark 1.3 at #1 on Website Arena, and Vals AI placed Spark 1.2 in its Vals Index top 5.

Then there are hands-on tests. AICodeKing's KingBench 3 pits Muse Spark 1.3 against Gemini 3.8 Flash and flags file-overwriting issues. Matt Johnston's blind coding gauntlet found a strong Halo build but broken XCOM and Diablo attempts. thehype had Muse Code and three rivals build landing pages and fix their own bugs in a real browser.

The research side covers how the models are made. Sebastian Raschka and elie broke down Muse Glimmer's Gemma-like architecture, and another post notes it was logit-distilled directly from Muse Spark. Meta's own benchmark claims are labeled as Meta's, and third-party numbers link to the people who ran them.

Frequently asked

+How good is Muse Spark 1.3 at coding?

Meta reports 75.4 on DeepSWE v1.1 and 88.8 on Terminal-Bench 2.1. Independent tests on this shelf are more mixed: DataCamp measured a cost increase on its tasks, and some coding gauntlets found failures on complex builds.

+Where does Muse Spark rank on Artificial Analysis?

Artificial Analysis scored Muse Spark 1.3 (max) at 62 on its Intelligence Index, behind Claude Fable 5.1 and Claude Opus 5. Earlier releases scored 52 for the original, 51 for 1.1 and 54 for 1.2.

+How does Muse Glimmer compare with Qwen?

Most comparisons here find Glimmer somewhat behind Qwen 3.6 27B on raw scores but using far fewer tokens. One Reddit analysis of Artificial Analysis data puts it at 92% of Qwen3.6 27B's index score with 2.9x fewer tokens.