DataCamp's Josep Ferrer ran Muse Spark 1.3 on three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase.
Resource · Benchmarks & research★ Pick
74 builds · page 1 of 1
DataCamp's Josep Ferrer ran Muse Spark 1.3 on three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase.
Resource · Benchmarks & research★ Pick
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
OpenRouter
@OpenRouter
Muse Spark 1.2 from @AIatMeta is live on OpenRouter alongside expanded global access to both Muse Spark models. At $1.25/M in and $4.25/M out, the model builds its position as one of the most price-efficient, high-intelligence models on OpenRouter.

X post · Coding & dev tools· ♥ 359
Design Arena
@DesignArena
BREAKING: Muse Spark 1.2 by @AIatMeta takes 1st for Video-to-Website with an Elo rating of 1279 and impressive scores across all of our multimodal code categories. Muse Spark 1.2 also takes 2nd for Image-to-HTML with an Elo rating of 1252 and 3rd for Image-to-Frontend with Elo

X post · Benchmarks & research· ♥ 254
Arena.ai
@arena
Muse Spark 1.2 (xHigh) by @AIatMeta is #14 in the Code Arena: WebDev, with 1,545 pts! This is an improvement from Muse Spark 1.1 at #18. See its biggest gains by category in the post below. Congrats to the @AIatMeta team on this release!

X post · Benchmarks & research· ♥ 380
OpenRouter
@OpenRouter
We’re excited to bring @AIatMeta’s Muse Spark 1.2 contributor tier to OpenRouter. At $0.10/M input and $0.20/M output, It’s meaningfully cheaper than Muse Spark 1.2 and beats other comparable models on real cost, providing frontier intelligence-per-dollar.

X post · Coding & dev tools· ♥ 586
AI/ML API
@aimlapi
Muse Spark 1.3 vs 1.2: sculpt viking figurines @AIatMeta dropped Muse Spark 1.3 today — we ran it against Muse Spark 1.2 both models got the same brief: three collectible 3D figurines — a viking helmet, a diamond-studded axe, a longship with a crew — one self-contained HTML
X post · Games & 3D· ♥ 139
Mark Zuckerberg
@finkd
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.

X post · Coding & dev tools· ♥ 14.8K
Artificial Analysis
@ArtificialAnlys
Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place

X post · Benchmarks & research· ♥ 1.2K
AI at Meta
@AIatMeta
Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re
X post · Local & open models· ♥ 531
AICodeKing tests Meta's Muse Spark 1.2 and the Muse Code terminal agent on the free tier, covering pricing and features before running KingBench 3.

Video · Coding & dev tools
Theo (t3.gg) tests Meta's Muse Code terminal agent powered by Muse Spark 1.2 and focuses on how cheap it is compared with Claude Code.

Video · Coding & dev tools· ♥ 3.5K
Bijan Bowen's first look at Meta's Muse Code terminal agent and Muse Spark 1.2, testing a browser OS, C++ skate game, CAD design, flight sim and subway FPS.

Video · Games & 3D
Design Arena
@DesignArena
BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362! This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price. Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this

X post · Benchmarks & research· ♥ 864
This week we spent about $95 trying to beat our own lineup of reviewing models. One of the candidates was Muse Spark 1.2, and it turned out to be the most interesting model in the whole test. The good, measured: • Among the best we tested at finding real problems. Scored against bugs we already knew were there, it matched our existing lineup, and it caught one real bug our lineup had missed. • Fastest model in our table. Typical answer in 18 seconds, writing at over 220 tokens a second. The speed table from our test (same job, same codebases, 33 runs per model): Model Typical time Answer length (tokens) Writing speed (tok/s) Time follows answer length Time follows question length Muse Spark 1.2 18 s 4,205 222 0.79 barely (0.08) Gemini 3.1 Pro 19 s 2,621 133 0.99 no (0.0) Gemini 3.8 Flash 22 s 1,996 89 0.91 some (0.65) GPT 5.4 29 s 3,058 105 0.95 no (below 0) Grok 4.6 38 s 2,498 61 0.61 no (below 0) Grok 4.7 44 s * 3,176 75 0.98 a little (0.30) Claude Sonnet 5 50 s * 4,471 91 0.59 barely (0.07) * Runs that finished in time only, so the real typical time is higher. The last two columns are correlations: 1 means time rises in step with that length, 0 means no lin
Reddit post · Benchmarks & research
Vals AI
@ValsAI
Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol.

X post · Benchmarks & research· ♥ 773
Alexandr Wang
@alexandr_wang
1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon. we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
X post · Local & open models· ♥ 9.5K
pilvar (Philippe Dourassov)
@pilvar222
We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers an average of 19/32 CVEs. In comparison, Grok 4.6 scores 23.3/32 - When pooling the results of 3 runs (pass@3), Muse scores 24/32. DeepSeek V4 Pro gets
X post · Benchmarks & research· ♥ 116
Vals AI
@ValsAI
Muse Spark 1.2 is the first model to crack 60% on Finance Agent v2, our benchmark that gives models the job of a financial analyst. At $0.77/test it is 6.7x cheaper than the previous #1, Opus 5 ($5.12), at twice the speed.

X post · Benchmarks & research· ♥ 406
I wanted a quick calories counter for myself, using LLMs to evaluate the calories from pictures of meals + descriptions. I needed to pick a model so I made a quick benchmark. The setup was: - Nutrition5k photos for photo + calories: https://github.com/google-research-datasets/Nutrition5k - A tool with access to calories information from USDA FoodData Central + MEXT - I evaluated models based on how many of the meals they managed to have under 20% of error - All on the same randomly picked 25 meals. Models too big for my machine were run through OpenCode Go/OpenRouter. I've also included Spark 1.3 since it'll supposedly be open weights. Results Model % within 20% Mean bias Median Error Qwen 3.8 27b 16% +64 kcal 148 kcal GLM 5.3 Flash 28% +18 kcal 65 kcal Qwen 3.8 Max 32% -11 kcal 48 kcal Muse Glimmer 30b 32% +25 kcal 92 kcal Qwen 3.8 Flash 36% +2 kcal 91 kcal DeepSeek v4 Flash Vision 40% +52 kcal 65 kcal Muse Spark 1.3 48% -24 kcal 45kcal I know it's not the most scientific benchmark, but it's interesting to see that the order is not really linked to model size. The most interesting for me is how Muse Glimmer 30b trounces Qwen 3.8 27b here. I think it hig
Reddit post · Benchmarks & research
Meta Model API provider extension for pi that adds Muse Spark 1.1, 1.2 and 1.2-contributor through the OpenAI-compatible Chat Completions endpoint.
Skill · Coding & dev tools
Paweł Huryn
@PawelHuryn
So, I finally tested Muse Spark 1.3. 2 real repos, 105 planted bugs, find and fix what you can. Original harness and API. Big surprise: Muse Spark 1.3 (max): 33 Fable 5.1 (high): 33 Grok 4.6 (xhigh): 27 Opus 5 (max): 27 Muse Spark 1.3 (high): 19 Meta joined the frontier.

X post · Benchmarks & research· ♥ 694
Iam_
@SPAC89
Muse Spark 1.3 Ultra Contributor vs Fable 5.1 xHigh Same prompt, both ran for roughly 2 hours The prompt had a self improvement rule: if the independent judges scored the result below 9.5/10, it had to keep improving and try again, What shocked me most was that Muse Spark just
X post · Benchmarks & research· ♥ 1.2K
li yin
@panda_liyin
muse spark 1.1 in AdaL Engineer beats Opus4.8 in Claude Code with 20% of the cost loop engineering, when done right, is beyond just running longer, its delivering better results when contexts are managed well and when workers are better prompted to stay honest. how GANs had
X post · Benchmarks & research· ♥ 75
Meta's developer blog introducing Muse Spark 1.2, co-trained with the new Muse Code terminal harness, with 1M-token context for multi-file refactors and hours-long tasks.

Resource · Coding & dev tools
Chetaslua
@chetaslua
This beautiful presentation is by Muse 1.2. It used gpt-image-v2 to create the images by itself. I told it to make a history page from Llama to Muse 1.2, and it did a great job. @alexandr_wang , you guys created magic with agentic tasks. I'm now waiting for the big model from
X post · Content & creative· ♥ 262
Julian Goldie's GoldieBench scores Muse Spark 1.2 at 7.55/10 across 50 one-shot builds. It did well on fractals (8.7), a macOS desktop sim (8.6) and a galaxy viz (8.6), but 3D games like Doom and Dragonrealm often came out incomplete or black.

Resource · Games & 3D
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Povilas Korop scores Muse Spark 1.3 on six real coding projects (Laravel, React-TS, PHP, Flutter, Go): Max effort ranks #26 with 48.43/60 at about $0.01 per prompt and 3:42 per prompt.

Resource · Benchmarks & research
Yash Patel
@yashvarpatel
🚀 Muse Spark 1.1 is live! Our new natively multimodal reasoning model brings powerful agentic & coding upgrades. CUA improved a lot, eg: on OS-World from 53.5 ➡️ 80.8 in 3 months. Fun personal CUA demo attached below: on making slides from these ~200 images of my trip to Kauai.
X post · Agents & automation· ♥ 16
eesel AI reports Muse Spark 1.3 ranks #6 on the Artificial Analysis Intelligence Index, leads long-context and coding rows, but trails Claude Opus 5 on four of six agent evals.
Resource · Benchmarks & research
A config (in Spanish) plus two helper proxies for using Muse Spark 1.2 Contributor from Claude Code and Hermes through routatic/proxy and an OpenCode Go subscription.
GitHub · Coding & dev tools
A small local proxy between Grok Build and OpenCode Zen that fixes muse-spark-1.2-contributor-free streaming, with a technical report of protocol captures and root cause.
GitHub · Coding & dev tools
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
A self-contained visual report comparing CursorBench 4.0 score against cost per task for Opus 5.5, Fable 5.1, Grok 4.7 and Muse Spark 1.3, where Muse Spark 1.3 Max scores 41.6% at $2.64 per task. GPT-6 points are clearly marked as estimates.
GitHub · Benchmarks & research
ocodex fans out Muse Spark 1.2 Contributor workers on decomposable coding work while one paid supervisor audits every claim, with crash checkpoints, heartbeats and a machine-readable ledger.
GitHub · Coding & dev tools
A localhost compatibility gateway that lets the native Muse Code harness run on OpenRouter's muse-spark-1.2-contributor, rewriting only the model name and never falling back silently to another model.
GitHub · Coding & dev tools
thehype.
@thehypedotnews
gemini 3.7 flash vs deepseek v4 pro 0813 vs muse spark 1.2 – on voxel city dioramas three models each built three crossy road-style 3d scenes – a construction site, a nyc intersection, a river with a drawbridge – as single self-contained html files the setup: @nousresearch's
X post · Games & 3D· ♥ 51
thehype.
@thehypedotnews
opus 5 vs kimi k3 vs muse spark 1.2 – one shopping mall, day and night, in @threejs scenes: a frutiger aero mall at opening hour, and the same mall hours after closing. one camera walk each, 3,600 frames per cut. muse needed a stronger day prompt, opus a hand-fix on the day
X post · Games & 3D· ♥ 45
AI/ML API
@aimlapi
Muse Glimmer one-shot 5 games! @AIatMeta dropped Muse Glimmer. So we ran it against Muse Spark 1.2, on five one-file game demos, each one written to play itself: Tetris, a top-down pixel street race, a rooftop web-slinger, a blue hedgehog platformer, and Flappy Bird. the setup:
X post · Games & 3D· ♥ 33
Dilmer Valecillos converted a minigolf prototype to VR for Meta Quest 3 using Muse Code with the Muse Spark 1.2 Contributor model, the Unity CLI and the MetaVR CLI, and published the plans and prompts.
GitHub · Games & 3D★ Pick
Dilmer Valecillos uses Muse Code with Muse Spark 1.2 and the Unity CLI to run tests, validate changes, port a Mini Golf game to other platforms and convert it to VR.

Video · Games & 3D★ Pick· ♥ 47
Arena.ai
@arena
Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4 in the Text Arena (1498 pts), and has reshaped the Pareto frontier! It is priced at $1.25/$4.25 per MToken. Congrats again to the @AIatMeta team on this release!

X post · Benchmarks & research· ♥ 422
Independent benchmarks and analysis of Muse Spark 1.2, released alongside Muse Code.

Resource · Benchmarks & research
Arena.ai
@arena
Muse Spark 1.2 (xHigh) by @AIatMeta is now in Agent Arena, with a net improvement of +2.1%! Agent Arena measures models on millions of real-world, long-horizon agentic tasks. We use causal tracing methodology to measure a model's net improvement, indicating how much it improves

X post · Benchmarks & research· ♥ 259
Zi Lin
@suzzzylin
🎮 This entire Avo Lawn game was built with Muse Spark 1.2 + Muse Code. 🥑 Now it’s your turn—use Muse Spark 1.2 to build your own game and see what you can create. To play: research.meta.ai/artifacts/intr… Check out our blog post: research.meta.ai/blog/introduci… h
X post · Games & 3D· ♥ 95
Arena.ai
@arena
Muse Spark 1.1 has entered the Code Arena: Frontend at #9! Muse Spark 1.1 reshapes the cost-performance Pareto Frontier by scoring 1541 at a blended $3.5M ($1.25 per input MToken, $4.25 per output MToken). This is frontier performance at a fraction of the price. Congrats to

X post · Benchmarks & research· ♥ 302
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Matt Johnston's live gauntlet puts Muse Spark 1.2 at 95 and #5 on his board, at $1.25/$4.25 per M tokens and 171 tok/s on OpenRouter; the full bench ran in 17 minutes.

Video · Benchmarks & research· ♥ 6
RuntimeWire finds Muse Spark 1.1 ahead of GLM 5.2 on harder, failure-prone tasks, with GLM the tidier formatter in spots.

Resource · Benchmarks & research
Composio compares Meta's Muse Spark 1.2 with DeepSeek V4 Flash on real-world agentic tasks to find the cheaper model that holds up.
Resource · Benchmarks & research
Motion Labs explains Muse Spark 1.2 setup, the cheaper Contributor tier at $0.10/M input tokens, Artificial Analysis scores and using it for content.

Resource · Content & creative
WTF Code shows how to wire Muse Spark 1.2 into OpenCode through the OpenCode Zen API and tests it as an Opus 4.8 challenger.

Video · Coding & dev tools· ♥ 100
Rohan Paul
@rohanpaul_ai
Meta's open-sourced Muse Glimmer 30B vs Muse Spark 1.2 Glimmer delivered 5 runnable games at roughly one-fifth Muse Spark's cost. Very Interesting experiment by @aimlapi . the setup: • Muse Glimmer 30B — via aimlapi[.]com. cost: $0.02 • Muse Spark 1.2 — via aimlapi[.]com.
X post · Games & 3D· ♥ 53
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Nadia Dubois walks through API key setup, auth, tool-calling config and cost control for Muse Spark 1.3 at $1.25/M input and $4.25/M output tokens.

Guide · Coding & dev tools
A VS Code extension that runs the Muse Code CLI in a sidebar and also adds Muse Spark 1.2 models to VS Code Chat Agent Mode, with inline completions and live token and cost tracking.
GitHub · Coding & dev tools
Cline
@cline
Muse Spark 1.1 just launched and it's their most capable coding agent model yet. On Terminal-Bench 2.1 it scores 80.0%, in the same cluster as Opus 4.8 (82.7%) and GPT 5.5 (83.4%). Use it in Cline with the Meta API!
X post · Benchmarks & research· ♥ 191
We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added them to the Cline harness. TL;DR of this special prompting: - Trust source code over the user prompt, so read every call site and existing tests before starting the task - Weigh edge and error cases as heavily as the happy path - Always reproduce the bug before fixing - Don't trust the first passing test suite, and verify suspicious looking half-baked tests - Never stop at just editing, keep working until the change is verified complete. We then asked this modified harness to fix a real bug from our repo, and compared the results to the original Cline agent harness. Results: - Used 2.7x fewer tokens (19.7M → 7.2M) - Finished 2x faster (49min → 24min) - Cost 2.4x less ($7.69 → $3.25) Same Muse Spark 1.2 model, same task, only the prompting changed. Incredible how much of a performance gain Meta was able to achieve training it on these special instructions!

Reddit post · Coding & dev tools
Louis-François Bouchard 🎥🤖
@Whats_AI
We just measured Meta Muse Spark 1.3 on our internal writing benchmark hoping for new SOTA. Unfortunately, it isn't... It enters at #24 of the 87 models we track, up from #31 for Muse Spark 1.1. As @alexandr_wang highlighted, it beats every Gemini configuration we have,

X post · Benchmarks & research· ♥ 9
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Box added Muse Spark 1.3 to Box AI, reporting it runs 42% faster than Muse Spark 1.2 with roughly a third fewer tokens, and lifts financial services accuracy from 66% to 75% on Box's eval.

Site · Business & commerce
Rihard Jarc
@RihardJarc
A few thoughts on the $META AI model's progress, because I think it is significant. 1. It does seem that $META has now leapfrogged $GOOGL in model quality when it comes to Muse 1.2 for many use cases, which is very surprising given the timeframe. 2. This is still the "Muse
X post · Business & commerce· ♥ 572
Ramanpal Singh builds a multi-page company site, a Gridlock puzzle game, a 3D traffic sim, the Shipyard release-notes SaaS and a diagram digitizer with Muse Code for about $10. Spark 1.2 often failed to enforce rules like win conditions.

Resource · Apps & websites
A single-file Three.js superhero platformer designed, coded and play-tested by Muse Spark 1.3 at xhigh effort in one autonomous session, passing 25/25 automated browser checks.
GitHub · Games & 3D
Meta's announcement of Muse Spark 1.3 for Muse Code and the Meta Model API, claiming ~20% fewer tool calls and ~25% fewer tokens than 1.2, with a max reasoning mode.

Resource · Coding & dev tools
Pi coding agent extension that adds Muse Spark 1.2, 1.2-contributor and 1.1 via the Meta Model API, a maintained fork that fixes a false auth warning on pi v0.84+.
Skill · Coding & dev tools· ★ 1
Federico Ramirez uses Muse Code and Spark 1.3 for GitHub issue-to-PR workflows. He finds output readable and $15/month good for 2-4 hours of continuous coding, but skill execution inconsistent and the sandbox too restrictive.

Resource · Coding & dev tools
A looping browser 3D rocket launch made with Muse Spark 1.3 in Vite, React and plain Three.js. Ignition, liftoff and ascent run about 22 seconds, with all geometry and smoke procedurally generated and no external assets.
GitHub · Games & 3D
いにしえ@高信頼AIニュース"NeuralWire.org"運営|Will Oldgram
@old_pgmrs_will
muse-spark-1.2-contributor で 3D都市を生成、生成コストは日本円で20円くらい 自動旋回で延々見てられる☺️ 好きな位置・カメラアングルでの画像生成もサポートしたので、別途画像生成AIの参照素材として色々使える
X post · Games & 3D· ♥ 8
Mark Zuckerberg
@finkd
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats
X post · Local & open models· ♥ 30.4K
Max Weinbach
@mweinbach
I asked Muse Spark 1.2 in Muse Code to make me an internal Bloomberg terminal style page for work on Friday This was what it made

X post · Apps & websites· ♥ 178
Out-of-tree Hermes Agent model-provider plugin for the Meta Model API and the Muse Spark family, including the low-cost muse-spark-1.2-contributor tier.
Skill · Coding & dev tools
MarkTechPost summarizes Meta's numbers: 75.4 on DeepSWE v1.1 (Opus 5 74.0, GPT-5.6 Sol 72.7), 88.8 on Terminal-Bench 2.1, and 98.1 on MRCR v2 at 512K–1M context.

Resource · Benchmarks & research
Simon Willison reads the Muse Code launch as more evidence that long-sequence agentic tool calling is now the key model trait.

Resource · Coding & dev tools
glimmer-cli is a local TypeScript CLI for Muse Glimmer 30B and Muse Spark 1.2 via Ollama, with stubbed tools and a reproducible tool-use eval harness.
GitHub · Local & open models
A config and write-up that makes muse-spark-1.2-contributor via OpenCode Go work in DeepSeek Harness, fixing empty first-turn output and multi-turn thinking replay errors.
GitHub · Coding & dev tools
Cline
@cline
We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added

X post · Coding & dev tools· ♥ 993
VS Code extension that adds Muse Spark 1.3, 1.2, 1.1 and Contributor variants to the Copilot Chat model picker with a Meta API key, published on the Marketplace and Open VSX.
Skill · Coding & dev tools· ★ 3
thehype.
@thehypedotnews
muse spark 1.2 vs gpt 5.6 sol vs kimi k3 vs grok 4.5 – on two landing pages four coding agents built two desktop landing pages from scratch, then had to open them in a real browser, find their own bugs and fix them before they were allowed to hand anything over the setup: each
X post · Benchmarks & research· ♥ 64