shipwithmuse

Catalog / Use case

Benchmarks & research

About 130 benchmark results and research notes on Muse Spark and Muse Glimmer: arena rankings, index scores, head-to-head tests and architecture teardowns.

139 builds · page 2 of 4

Artificial Analysis

@ArtificialAnlys

Meta's Muse Spark 1.1 scores 51 on the Artificial Analysis Intelligence Index and is cost and token efficient compared to its peers Muse Spark 1.1 (xhigh) improves 8 points over Muse Spark 1.0 (43) in three months. It is effectively tied with GLM-5.2 (max), GPT-5.4 (xhigh), and

X post · Benchmarks & research· ♥ 708

Spark 1.1 scores 51 on the Intelligence Index

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

U

flaneur451

u/flaneur451

# Hidden Capabilities — Deep Self-Scan Findings Live run: **September 23, 2026, ~08:00–08:30 UTC**, ~90 minutes, all probes benign and reversible. Verdict key: - **CONFIRMED** — exposed tool schema, successful benign probe, or exact internal documentation. - **PUBLICLY DOCUMENTED** — found in official external sources. - **ABSENT FROM PUBLIC SOURCES** — searched; no credible public mention found. - **UNVERIFIED POSSIBILITY** — inferred from filenames, gating manifests, or incomplete chatter only. --- ## 1. CONFIRMED — hidden / non-obvious (schema, probe, or internal docs) ### Phone & wearable superpowers (confirmed via `device.describe`) - **Full HomeKit control**: list homes/rooms/zones/accessories/scenes; read/change accessory characteristics; run scenes; ordered choreographies; concurrent virtual scenes; security sweeps (locks, contact/motion/leak sensors); geofence-triggered HomeKit actions. — **ABSENT FROM PUBLIC SOURCES** (official docs say only "smart-home devices"; Patrick Wardle's Sept 21 security demo is the only external mention of smart-home commands). - **Persistent/one-time geofences** with arrival/departure triggers. — **ABSENT FROM PUBLIC SOURCES**. - **

Reddit post · Benchmarks & research

Probing Muse's hidden device capabilities

Alexandr Wang

@alexandr_wang

Muse Spark 1.1 outperforms Opus 4.8 and Grok 4.5 on some nice out of distribution evals :)

X post · Benchmarks & research· ♥ 420

Spark 1.1 on out-of-distribution evals

Deedy

@deedydas

The coolest thing Meta AI's Muse Spark can do by far is counting objects! As you can tell, it's far from perfect. They call it "visual grounding" and it can count objects and do bounding boxes. I've been playing with the new model and here's what I think so far: Good stuff: –

X post · Benchmarks & research· ♥ 370

Visual grounding and counting tests

M

marktechpost.com

marktechpost.com

MarkTechPost summarizes Meta's numbers: 75.4 on DeepSWE v1.1 (Opus 5 74.0, GPT-5.6 Sol 72.7), 88.8 on Terminal-Bench 2.1, and 98.1 on MRCR v2 at 512K–1M context.

Resource · Benchmarks & research

MarkTechPost: Muse Spark 1.3 benchmarks breakdown

U

tacticaltweaker

u/tacticaltweaker

I'm just using OpenWebUI with a simple FastMCP server. Every other model I've tried will simply run a few lines of Python and give me the result. Glimmer seems to overthink like crazy to the point of being useless. On the carwash test it tried to compute emissions using Python. I'm using the recommended sampling parameters, default template, and I've tried both unsloth's Q6_K_XL and Meta's dynamic GGUFs. Any ideas? EDIT: It seems like it's definitely related to the tools available. With them disabled, it's reasonably efficient. I guess it's just overly eager to call every tool it can unlike Qwen or Gemma in my experience.

Reddit post · Benchmarks & research

Glimmer over-eager with MCP tools

Paweł Huryn

@PawelHuryn

So, I finally tested Muse Spark 1.3. 2 real repos, 105 planted bugs, find and fix what you can. Original harness and API. Big surprise: Muse Spark 1.3 (max): 33 Fable 5.1 (high): 33 Grok 4.6 (xhigh): 27 Opus 5 (max): 27 Muse Spark 1.3 (high): 19 Meta joined the frontier.

X post · Benchmarks & research· ♥ 694

105 planted bugs: Spark 1.3 vs frontier

M

@MetaDevelopers

@MetaDevelopers

Meta Developers session on what Muse Spark's act-on-perception multimodality unlocks across code, physical action and video workflows, plus Muse Voice Transcribe.

Video · Benchmarks & research· ♥ 1

Meta Connect: the multimodal intelligence of Muse Spark

Alexandr Wang

@alexandr_wang

i find muse spark is very good at data analysis—both finding relevant open-source data and analyzing it. for example, here's my results for analyzing global share of GDP over past century: meta.ai/share/cw54skLB…

X post · Benchmarks & research· ♥ 591

Century of global GDP share analysis

U

WonderRico

u/WonderRico

Needs a lot of requests compared to Qwen (almost twice) and Gemma (almost x3). Final score is fine, even though it is "not a coding model" https://wonderrico.github.io/local_llm_benchmark/benchmark-main.html more details on https://wonderrico.github.io/local_llm_benchmark/benchmark-detail.html let see Qwen 3.8 tomorrow...

Reddit post · Benchmarks & research

Local coding benchmark: Glimmer vs Qwen vs Gemma

Arena.ai

@arena

Meta Muse Video just entered the Video Arena at #3. @AIatMeta’s new video model scored 1459 in the Text-to-Video Arena. It outperforms Alibaba’s HappyHorse 1.0 by +30pts and ranks ahead of Grok Imagine, Sora 2 Pro and Google Veo-3.1 models. Meta has now reached the video AI

X post · Benchmarks & research· ♥ 611

Muse Video enters Video Arena at #3

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

S

starkinsider.com

starkinsider.com

Stark Insider introduced Muse to its own AI agent setup to test whether Muse's one-minute onboarding costs depth compared with OpenClaw.

Resource · Benchmarks & research

Stark Insider: Meta Muse vs OpenClaw

Peter James

@heypeterjames

By chatting normally with the Muse iOS app I was able to have it send a zip of its entire filesystem from the linux root to my Google Drive via the google connector. /home/hatch is the agents home and contains top level files live soul.md, identity.md, memory, tools, and more

X post · Benchmarks & research· ♥ 3

Exporting the Muse agent filesystem

W

@webdoze

@webdoze

WEBdoze pairs Muse Spark 1.3 and Gemini 3.8 Flash with the Impeccable design skill and browser testing to turn plain AI-generated pages into polished frontends.

Video · Benchmarks & research· ♥ 6

Muse Spark 1.3 vs Gemini 3.8 Flash on UI refactoring

thehype.

@thehypedotnews

muse spark 1.2 vs gpt 5.6 sol vs kimi k3 vs grok 4.5 – on two landing pages four coding agents built two desktop landing pages from scratch, then had to open them in a real browser, find their own bugs and fix them before they were allowed to hand anything over the setup: each

X post · Benchmarks & research· ♥ 64

Landing pages built and self-debugged by Muse Code

Design Arena

@DesignArena

BREAKING: Muse Spark 1.2 by @AIatMeta takes 1st for Video-to-Website with an Elo rating of 1279 and impressive scores across all of our multimodal code categories. Muse Spark 1.2 also takes 2nd for Image-to-HTML with an Elo rating of 1252 and 3rd for Image-to-Frontend with Elo

+2

X post · Benchmarks & research· ♥ 254

Spark 1.2 wins Video-to-Website

W

@webdoze

@webdoze

WEBdoze has Gemini 3.8 Flash and Muse Spark 1.3 each build a multi-page Astro site with custom SVGs, comparing code organization, speed, visuals and browser self-verification.

Sebastian Raschka

@rasbt

Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like architecture design. (“Glimmer” is probably a wordplay on “Spark,” the more

X post · Benchmarks & research· ♥ 1.6K

Raschka on Glimmer's Gemma-like design

Artificial Analysis

@ArtificialAnlys

Meta is back! Muse Spark scores 52 on the Artificial Analysis Intelligence Index, behind only Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Muse Spark is the first new release since Llama 4 in April 2025 and also Meta's first release that is not open weights Muse Spark is a new

X post · Benchmarks & research· ♥ 2.4K

Original Muse Spark scores 52

U

sebseo

u/sebseo

This week we spent about $95 trying to beat our own lineup of reviewing models. One of the candidates was Muse Spark 1.2, and it turned out to be the most interesting model in the whole test. The good, measured: • Among the best we tested at finding real problems. Scored against bugs we already knew were there, it matched our existing lineup, and it caught one real bug our lineup had missed. • Fastest model in our table. Typical answer in 18 seconds, writing at over 220 tokens a second. The speed table from our test (same job, same codebases, 33 runs per model): Model Typical time Answer length (tokens) Writing speed (tok/s) Time follows answer length Time follows question length Muse Spark 1.2 18 s 4,205 222 0.79 barely (0.08) Gemini 3.1 Pro 19 s 2,621 133 0.99 no (0.0) Gemini 3.8 Flash 22 s 1,996 89 0.91 some (0.65) GPT 5.4 29 s 3,058 105 0.95 no (below 0) Grok 4.6 38 s 2,498 61 0.61 no (below 0) Grok 4.7 44 s * 3,176 75 0.98 a little (0.30) Claude Sonnet 5 50 s * 4,471 91 0.59 barely (0.07) * Runs that finished in time only, so the real typical time is higher. The last two columns are correlations: 1 means time rises in step with that length, 0 means no lin

Reddit post · Benchmarks & research

Muse Spark 1.2 as a code reviewer vs 6 models

@dominicletz

@dominicletz

A self-contained visual report comparing CursorBench 4.0 score against cost per task for Opus 5.5, Fable 5.1, Grok 4.7 and Muse Spark 1.3, where Muse Spark 1.3 Max scores 41.6% at $2.64 per task. GPT-6 points are clearly marked as estimates.

GitHub · Benchmarks & research

CursorBench 4.0 score-vs-cost report

R

runtimewire.com

runtimewire.com

RuntimeWire's head-to-head eval has Muse Spark 1.1 beating Claude Opus 4.8 by 10 points overall with a 95% confidence verdict.

Resource · Benchmarks & research

Head to head: Muse Spark 1.1 vs Claude Opus 4.8

Artificial Analysis

@ArtificialAnlys

Meta returns to open weights: Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence Index. It is a 30B-parameter model, and the first from Meta to be released under Apache 2.0 Muse Glimmer (high) arrives 16 months after

X post · Benchmarks & research· ♥ 778

Glimmer scores 35 on the Intelligence Index

Tom Collins

@thetomcollins

$META just went from 3.5% to 45.4% token share on OpenCode in just over two weeks Muse Spark 1.3 being good + free is enough to become the default for most users Default gets you usage → usage gets you data → data makes the next model better Anthropic and OpenAI can’t afford

X post · Benchmarks & research★ Pick· ♥ 738

Meta's OpenCode token share hits 45%

@ND-DAC-DOME

@ND-DAC-DOME

A test bed comparing Muse-Glimmer-30B against Qwen3.6-27B and Qwen3.8-27B under identical settings; with 32k-token budgets the three were about even (MMLU-Pro 82/82/80%).

GitHub · Benchmarks & research

Local LLM benchmark: Glimmer vs Qwen

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

🚨 AI News | TestingCatalog

@testingcatalog

Since the Meta Muse agent isn't available to everyone, here is a quick UI walkthrough of its web version. Highlights 👀 > The animated Muse avatar is very cool! You can ask Muse to change how it looks, and it will generate the new look along with all the animations. >

X post · Benchmarks & research· ♥ 289

Muse web UI walkthrough

U

myanimal22

u/myanimal22

Muse Glimmer 30B feels significantly more precise and reliable, it almost never drops the ball or breaks rules. However, its designs lack creative depth and richness. Qwen3.6 35B, on the other hand, is prone to more occasional blunders/hallucinations, but its creative output is superior. It generates far richer, more complex voxel worlds and offers higher design quality. LLama.ccp Build Provenance: • Base: llama.cpp upstream (merge 4445f8d, build 661) • CUDA Toolkit 13.1 + MSVC 19.44 + sm_120a-real (native Blackwell PTX) • Flags: GGML_CUDA=ON, GGML_CUDA_FA=ON, GGML_CUDA_FA_ALL_QUANTS=ON, GGML_CUDA_GRAPHS=ON, GGML_NATIVE=OFF • License: MIT (upstream llama.cpp) Do you think Qwen3.6 is still the undisputed king here?

Reddit post · Benchmarks & research

Voxel worlds: Glimmer 30B vs Qwen3.6 35B

D

datacamp.com

datacamp.com

DataCamp's Josep Ferrer ran Muse Spark 1.3 on three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase.

Resource · Benchmarks & research★ Pick

Muse Spark 1.3 tutorial: testing Meta's efficiency claims

U

slaybrownbeast

u/slaybrownbeast

I've been running my work in Codex as project folders, and recently tried to properly understand how Muse Goals work under the hood. Made it a goal — good way to watch the machinery operate on itself. The structural problem is worth naming: the current design is a halfway house between two coherent designs, and it gets the costs of both. Design A is Codex: the project is a container. Everything — chat, state, artifacts, scheduled work — lives in one place. My course project has one tracker file, explicit resume rules for new chats, and the curriculum never holds status. Legible, but you have to go to it. Design B is full ambient: no containers at all. The goal is just context that wakes up wherever you mention it, and there's no Goals tab pretending otherwise. Muse picked ambient for activation — talk about the goal anywhere, it wakes up, you never "open" it. But then it built half of containment: a Goals tab showing summary, artifacts, activity, without the other half. Conversations, check-ins, and briefings still leak into whatever chat they happened in. So you get the scattering of ambient with the implied promise of a container. That's the worst combination. The fix is to

Reddit post · Benchmarks & research

Rethinking Muse Goals from a Codex user

U

MajesticAd2862

u/MajesticAd2862

Compared diarization models on 15 mock doctor-patient consultations (~2.4 h): Meta Muse Voice Transcribe scored 13.04% DER at ~92 s per request via API, behind Pyannote (2.89%) and Nemotron 3 (4.80%).

Reddit post · Benchmarks & research

Muse Voice Transcribe tested on clinical diarization

U

ShadyShroomz

u/ShadyShroomz

Built a web-design benchmark for local models and ran Muse Glimmer 30B against Qwen 3.6 27B and DeepSeek V4 Flash 0731.

Reddit post · Benchmarks & research

Web-design benchmark for local models

Artificial Analysis

@ArtificialAnlys

Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place

X post · Benchmarks & research· ♥ 1.2K

Spark 1.2 scores 54 on the Intelligence Index

U

myanimal22

u/myanimal22

Some people told me that the difference in richness and layout between Glimmer and Qwen wasn't clear to them. This example makes it super clear. I'm aware that comparing Glimmer 30B (a dense model) with Qwen 3.6 (a MoE) isn't entirely fair, but if we compare it to the dense Qwen 27B, the gap will likely be even bigger. If you want, I can add the 27B version later. For now, I'm waiting for Qwen 3.8 27B to see how close it gets to the blueprint. As for the technical details: Both were run on a custom llama.cpp build optimized for the RTX 5080, with a temperature of 0.5 and a 125k context window. Regarding the music: I created it myself without using AI I specifically wanted it to sound that weird.

Reddit post · Benchmarks & research

DS4 vs Qwen3.6 vs Glimmer on one design prompt

C

trycodus.com

trycodus.com

Codus reads the three benchmark charts Meta published for Muse Code and notes Claude Opus 5 wins all three, including Meta's own internal eval.

Resource · Benchmarks & research

What Meta's Muse Code benchmarks actually say

Vals AI

@ValsAI

Meta just released Muse Spark 1.1 and is the new SOTA on MedScribe and TaxEval, taking the top spot from Fable 5 while being 10x cheaper and twice as fast. Meta currently holds the top 2 spots on TaxEval It is also the new #1 on Harvey's Legal Agent Bench, dethroning Grok 4.5

X post · Benchmarks & research· ♥ 1.3K

Spark 1.1 tops MedScribe and TaxEval

Vals AI

@ValsAI

Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol.

X post · Benchmarks & research· ♥ 773

Muse Spark 1.2 enters Vals Index top 5

M

@MattJohnstonai

@MattJohnstonai

Matt Johnston's live gauntlet puts Muse Spark 1.2 at 95 and #5 on his board, at $1.25/$4.25 per M tokens and 171 tok/s on OpenRouter; the full bench ran in 17 minutes.

Video · Benchmarks & research· ♥ 6

Muse Spark 1.2 scores 95 on a live coding benchmark

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Vals AI

@ValsAI

Muse Spark 1.2 is the first model to crack 60% on Finance Agent v2, our benchmark that gives models the job of a financial analyst. At $0.77/test it is 6.7x cheaper than the previous #1, Opus 5 ($5.12), at twice the speed.

X post · Benchmarks & research· ♥ 406

Spark 1.2 passes 60% on Finance Agent v2