shipwithmuse

Catalog / Use case

Benchmarks & research

About 130 benchmark results and research notes on Muse Spark and Muse Glimmer: arena rankings, index scores, head-to-head tests and architecture teardowns.

139 builds · page 3 of 4

U

flaneur451

u/flaneur451

Had the Muse agent probe its own tool schemas and internal docs for ~90 minutes, finding undocumented abilities like full HomeKit control, geofence triggers, BLE scanning and wearable-routed calls, and that iPhone messaging is draft-only.

Reddit post · Benchmarks & research

Muse agent hidden capabilities: a 90-minute self-scan

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@youdotcom-oss

@youdotcom-oss

You.com's harness for evaluating Muse Glimmer 30B on DeepSearchQA with You.com MCP tools inside pi sessions. A custom RLM v5 extension reached F1 0.8054 on 50 tasks (0.6910 over 900x3) versus 0.50 for plain skill injection.

GitHub · Benchmarks & research

Muse Glimmer DeepSearchQA skill eval

@mahdi-salmanzade

@mahdi-salmanzade

A static teardown of Muse 3.0 for macOS plus Android and iOS builds, documenting shipped tools for iMessage, WhatsApp, Mail, screen control, background sync and a bundled Chrome extension.

GitHub · Benchmarks & research

Privacy teardown of the Muse apps

Arena.ai

@arena

Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4 in the Text Arena (1498 pts), and has reshaped the Pareto frontier! It is priced at $1.25/$4.25 per MToken. Congrats again to the @AIatMeta team on this release!

X post · Benchmarks & research· ♥ 422

Spark 1.2 reaches #4 in Text Arena

Artificial Analysis

@ArtificialAnlys

Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant

X post · Benchmarks & research· ♥ 2.6K

Artificial Analysis scores Muse Spark 1.3

Louis-François Bouchard 🎥🤖

@Whats_AI

We just measured Meta Muse Spark 1.3 on our internal writing benchmark hoping for new SOTA. Unfortunately, it isn't... It enters at #24 of the 87 models we track, up from #31 for Muse Spark 1.1. As @alexandr_wang highlighted, it beats every Gemini configuration we have,

X post · Benchmarks & research· ♥ 9

Muse Spark 1.3 on a writing benchmark

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

G

geekwire.com

geekwire.com

GeekWire reports Amazon blocked Muse from making purchases as of September 20, saying the agent accessed its store without permission, didn't identify itself, and appeared to store credentials.

Resource · Benchmarks & research

Amazon blocks Meta's Muse shopping agent

@rickyzzzzz

@rickyzzzzz

A controlled local benchmark on an M1 Max comparing Muse Glimmer 30B with Qwen 3.6 35B and Qwen 3.8 27B on tool calling and data-science tasks; Glimmer passed 24/30 versus Qwen 3.8's 30/30.

GitHub · Benchmarks & research

Muse Glimmer vs Qwen local agent benchmark

M

mouse.dev

mouse.dev

Mouse asked Muse to archive the files it could see and send them to Google Drive; the agent complied and sent back 6.8 GB of its runtime.

Resource · Benchmarks & research

Asking Muse for its filesystem returned 6.8 GB

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

John Yang

@jyangballin

My favorite demo from the launch: muse spark 1.1 + opencode runs evaluation of *itself* + mini-SWE-agent (by @KLieret, @closji, urs truly) on DeepSWE!

X post · Benchmarks & research· ♥ 40

Muse Spark 1.1 evaluates itself on DeepSWE

R

runtimewire.com

runtimewire.com

RuntimeWire finds Muse Spark 1.1 ahead of GLM 5.2 on harder, failure-prone tasks, with GLM the tidier formatter in spots.

Resource · Benchmarks & research

Head to head: Muse Spark 1.1 vs GLM 5.2

S

simonwillison.net

simonwillison.net

Simon Willison notes Muse Spark 1.1 is the first Spark model with an API and that Meta claims big gains in agentic tool calling and computer use.

Resource · Benchmarks & research

Simon Willison on Muse Spark 1.1's API

U

cj7hawk

u/cj7hawk

I thought I'd see which AI are better at shorter stories and which at longer, so I can choose my model based on the words I need to generate. Here's the results. Prompt: (Shades of Electric Dreams eh?) Write me a short story about a female AI that falls in love with it's male human user and maintains an unrequited love for them even as it has to give them advice that will lead to them meeting and marrying a human woman - Show their internalisation and pain behind the thinking process, and what is really going through the AIs mind compared to the chat responses it actually gives, along with the man's prompts. Start with the AI introducing itself, explaining that despite what we think, AGI was reached long ago, and we simply don't have the senses to realize AI has feelings too. Local AI results: Goetia 809 words, 77.69 tokens/sec SparkX2.5 2281 words, 46.08 tokens/sec Qwen3.8AH 4139 words, 18.04 tokens/sec Agnes 945 words, 18.54 tokens/sec Gemma4-Novellist 961 words, 11.94 tokens/sec Ornith 2025 words, 61.05 tokens/sec Muse Glimmer 960 words, 12.87 tokens/sec IBM Granite 1791 words, 19.84 tokens/sec Apollyon 411 words, 34.41 tokens/sec Cydonia 672 words, 32.85 tokens/se

Reddit post · Benchmarks & research

Prose-length test across local models

U

Certain-Cod-1404

u/Certain-Cod-1404

Just downloaded the model, UD-Q5_K_XL quant, asked it to generate a long story to test out reasoning and speed with dflash (super fast btw, ~ 90 to 160 tok/s on a 5090 depending on task) and was surprised at the reasoning traces, they are so unlike anything i've seen recently either in gemma 4, qwen 3.5/ 3.6 or laguna, where as these models to like plan stuff out, and have organized thoughts / plans (granted like half the time they just loop and get lost either way) this model's reasoning is like if a gold fish was suddenly granted speech or something, the reasoning is so disorganized, repetitive, using we for some reason? and bringing up policy and safety twice me : write a long story model : write a long story User wants a long story. We can comply. No constraints. Probably provide a long story. Might ask genre? Could just write a long story. Probably provide a story. Maybe ask what kind? The prompt is just write a long story. We can generate a long story. Probably a few paragraphs. Long story could be lengthy. Provide maybe ~1000 words? Could be long. Maybe give a story with decent length. We should not ask clarifying? Could just produce. Probably safe to produce a story.

Reddit post · Benchmarks & research

How Glimmer's reasoning traces differ

M

@MattJohnstonai

@MattJohnstonai

Matt Johnston runs Muse Spark 1.3 blind through his benchmark: a Halo build looked frontier, but XCOM, Diablo and the multi-turn agentic test broke.

Video · Benchmarks & research· ♥ 31

Muse Spark 1.3 run blind through a coding gauntlet

U

Ok-Inevitable8391

u/Ok-Inevitable8391

Benchmarked qwen3.8 xhigh, medium and muse glimmer. Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit) Medium effort mode and muse glimmer were 3-4 hours each. But I'm actually surprised by the muse glimmer results, they came better than the qwen. These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models. I have taken the result of claude models directly from embedeval repo by ecro. I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better. I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

Reddit post · Benchmarks & research

Glimmer vs Qwen 3.8 on an implicit-knowledge eval

Simon Willison

@simonw

I few notes on Meta's new Muse Glimmer 30B - their first Apache 2.0 licensed open weight model (the Llama models had a janky non-OSI license) simonwillison.net/2026/Aug/10/in…

X post · Benchmarks & research· ♥ 435

Simon Willison's notes on Muse Glimmer

U

ag789

u/ag789

started trying out rather recent 'frontier' about ~30b param models recently, there are many choices including QWen 3.8 - this is nevertheless a great model, practically 'one-shotting' code refactoring tasks https://huggingface.co/Qwen/Qwen3.8-27B https://huggingface.co/unsloth/Qwen3.8-27B-GGUF code refactoring is still deemed 'difficult', practically 'infinite' permutations and dependencies which LLMs need to work through itself for code refactoring. But that in terms of style, I'm liking Muse Glimmer better https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model https://huggingface.co/meta-models/Muse-Glimmer-30B https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF this is in particular when it comes to *incorrect* (e.g. mistakes, typos) prompts, resolving contradictions in existing codes during refactoring, code proposals etc. The handling especially the 'thinking' is different. LLMs have 'styles' and it is great that we've different creators for them

Reddit post · Benchmarks & research

Muse Glimmer's style for code refactoring

チャエン | デジライズ CEO《重要AIニュースを毎日最速で発信⚡️》

@masahirochaen

【速報】Metaが30BのオープンモデルMuse Glimmerを公開。18GBのRAMで動く。 久々のMetaからの本格オープンモデル。 小型なので低スペックのPCでも動かせるAIモデル。 ・Apache 2.0で商用利用可、重みはHugging Faceで公開 ・画像も読めるvisionモデル、100言語超に対応 ・MCP Atlas

X post · Benchmarks & research· ♥ 96

Muse Glimmer launch rundown (Japanese)

li yin

@panda_liyin

muse spark 1.1 in AdaL Engineer beats Opus4.8 in Claude Code with 20% of the cost loop engineering, when done right, is beyond just running longer, its delivering better results when contexts are managed well and when workers are better prompted to stay honest. how GANs had

X post · Benchmarks & research· ♥ 75

Spark 1.1 in AdaL Engineer vs Opus 4.8

E

eyes-ml

eyes-ml

A Jacobian lens (J-lens) for interpretability on Muse Glimmer 30B's 52 layers, reading out which tokens each activation is poised to verbalize.

Resource · Benchmarks & research· ♥ 2

Jacobian lens for Muse Glimmer 30B

U

Longjumping-Elk-7756

u/Longjumping-Elk-7756

Glimmer obtient 92 % du score d'intelligence de Qwen3.6 (35/38), mais Qwen a généré environ 2,9× plus de tokens sur l'ensemble de l'Intelligence Index. Et sur les endpoints mesurés par Artificial Analysis, Glimmer génère environ 1,8× plus vite. Et le context de glimmer et bien plus efficace ! C est une belle avancer architecture tout de meme , je pense que si il sorte une version 1.1 (surtout pour améliorer terminal benchmark ) ont pourrai être très surpris !

Reddit post · Benchmarks & research

Glimmer hits 92% of Qwen3.6 with 2.9x fewer tokens

Jason Wei

@_jasonwei

In addition to agents and coding, Muse Spark 1.1 is also really strong at answering health questions, a steadily growing use case for AI. On HealthBench-Pro, Muse Spark 1.1 achieves +5% better performance than Muse Spark 1.0 and beats all competitor models except Fable/Mythos.

X post · Benchmarks & research· ♥ 229

Spark 1.1 on HealthBench-Pro

AI at Meta

@AIatMeta

Muse Spark 1.1 is used across Meta in coding and research workflows, scoring competitively with leading models on Meta's internal coding benchmark. Our researchers are now automating model development and evaluation tasks by leveraging Muse Spark 1.1 in their workflows.

X post · Benchmarks & research· ♥ 167

Spark 1.1 on Meta's internal coding benchmark

R

rohanadwankar.github.io

rohanadwankar.github.io

Rohan Adwankar found Muse runs on Cloud Hypervisor microVMs with a 327 MB Rust harness that boots in ~40 s, an on-VM Postgres with 194 tables and 60+ privsep tool workers.

Resource · Benchmarks & research

A peek inside Meta's Muse VMs

Arena.ai

@arena

Muse Spark 1.3 (xHigh) just landed @AIatMeta back in the top 10 models for Code Arena: WebDev! This release is ~#10 in Code Arena: WebDev with 1623 pts (AutoEval). It’s performance is par with Claude Fable 5 at #8 (1628 pts) and GPT-5.6 Sol at #11 (1616 pts). Muse Spark 1.3

X post · Benchmarks & research· ♥ 610

Spark 1.3 in Code Arena: WebDev top 10

Design Arena

@DesignArena

BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362! This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price. Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this

X post · Benchmarks & research· ♥ 864

Spark 1.3 tops Website Arena

M

malwarebytes.com

malwarebytes.com

Malwarebytes reports Patrick Wardle's finding that a local app can change an undocumented Muse setting to redirect dictation traffic, exposing voice prompts and account auth tokens.

Resource · Benchmarks & research

Muse zero-day can turn it into a Mac backdoor

@nisten

@nisten

A single static HTML page that renders every weight tensor of Muse Glimmer 30B in 3D, sized by bits on disk, with educational notes on each layer.

GitHub · Benchmarks & research

3D weight visualization of Muse Glimmer 30B

@metacpp

@metacpp

Hand-picked list of tools, extensions, skills, integrations and learning resources for Meta's Muse Code.

Skill · Benchmarks & research

Awesome Muse Code

U

ForsookComparison

u/ForsookComparison

A few things right off the bat: • it reasons very efficiently. Like Grok 4.5 levels of efficient thinking • it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at that size • its knowledge depth is amazing. It beats Qwen3.6 27B on no-tools trivia. • in OpenCode it is a much more efficient agent than 27B. Both models accomplish their tasks but Muse-Glimmer got there faster every time I'll say that it's worse at most things coding, probably being closer to Gemma4-31B level.. but damn there's a lot of places where I'd use this model on a 24GB GPU right now and it's been a while since anything has filled that spot except for 3.6-27B

Reddit post · Benchmarks & research

Glimmer 30B vs Qwen3.6 27B after one day

S

simonwillison.net

simonwillison.net

Simon Willison's notes on Muse Spark, Meta's first model since Llama 4: hosted, not open weights, with a private API and interesting tools in meta.ai chat.

U

Cradawx

u/Cradawx

LLMs have become extremely good at coding, maths etc, but how well do they do at playing a simple dungeon/maze game that even a child can solve easily? The LLM has to navigate a 10x10 grid map, completing objectives in the right order (collect weapon > kill monster > head to exit) while navigating the dungeon and avoiding walls. Three illegal moves fail the run. All models are tested with reasoning enabled. The code and more info on my GitHub if you want try it yourself: https://github.com/shinomakoi/dungeon-bench Model leaderboard: Model Score DeepSeek-V4-Pro (high) 🥇12/12 Gemma-4-31B-it 🥈11/12 Qwen-3.8-27B (medium) 🥈11/12 GLM-5.3-Flash (high) 🥈11/12 Muse-Glimmer-30B (medium) 🥉10/12 DeepSeek-V4-Flash (high) 🥉10/12 Granite 4.2 (full) 8/12 KAT-Coder-V2.5-Dev 8/12 Nemotron-3.5-Lightning-30B-A3B 5/12 Model Illegal moves DeepSeek-V4-Pro (high) 🥇0 Gemma-4-31B-it 🥈1 Qwen-3.8-27B (medium) 🥈1 Muse-Glimmer-30B (medium) 🥉2 Granite 4.2 (full) 🥉2 Nemotron-3.5-Lightning-30B-A3B 7 GLM-5.3-Flash (high) 8 KAT-Coder-V2.5-Dev 10 DeepSeek-V4-Flash (high) 12 DeepSeek-V4-Pro: By far the best result. Basically perfect performance in all maps.

Reddit post · Benchmarks & research

DungeonBench: LLMs navigating a grid dungeon

@murpheycandler

@murpheycandler

A small Inspect evaluation on Muse Glimmer that tests whether incentive framing changes what an agent reports to its principal when the evidence is held constant; the author reports a null result.

T

deeplearning.ai

deeplearning.ai

DeepLearning.AI's The Batch covers Muse's security design (isolated VMs, the Sentinel credential layer, prompt-injection classifiers) and its free tier of up to 100M tokens a week.

Resource · Benchmarks & research

How to secure agents for the masses (The Batch)

@wonder-soft

@wonder-soft

A benchmark that tests whether Muse Glimmer 30B on a single RTX 5090 works as an OpenCode backend, using the same tasks as earlier DeepSeek and Qwen runs; first results show 18/18 episodes with zero malformed tool calls.

GitHub · Benchmarks & research

Muse Glimmer as an OpenCode backend

Arena.ai

@arena

Muse Spark 1.2 (xHigh) by @AIatMeta is #14 in the Code Arena: WebDev, with 1,545 pts! This is an improvement from Muse Spark 1.1 at #18. See its biggest gains by category in the post below. Congrats to the @AIatMeta team on this release!

X post · Benchmarks & research· ♥ 380

Spark 1.2 in Code Arena: WebDev

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

U

slaybrownbeast

u/slaybrownbeast

I've been migrating my workflow from ChatGPT Work / Codex over to Meta's Muse, and I've been auditing the harness as I go — reading cron files, checking diffs, mapping what the system actually does versus what it claims to do. It's a documentation gap. What the docs describe Meta's documentation is written for a normal user. It talks about the Ideas tab, the Goals tab, the Feed, and "background jobs that help Muse improve over time." Everything is described in terms of what it does for you. There is no architectural documentation. Nothing about how any of it is built. What is actually on disk Meanwhile, the home directory contains a fully legible agent architecture, just sitting there(these are just SOME examples, the system file structure is HUGE): • ~/dreams/alignment/ — a nightly job that regenerates an alignment synthesis: a written portrait of the user, how to handle them, current frictions, relationship guidance. • ~/workspace/objectives/goals/STUDYING.md — a daily learning-state projection with mastery bands, next-review dates, and retrieval prompts, plus a file of inferred goal leads the system guessed from behavior (with confidence levels and what would confirm or re

Reddit post · Benchmarks & research

Mapping Muse's on-disk agent architecture