DataCamp's Josep Ferrer ran Muse Spark 1.3 on three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase.
Resource · Benchmarks & research★ Pick
11 builds · page 1 of 1
DataCamp's Josep Ferrer ran Muse Spark 1.3 on three real coding tasks. Two used 23–32% fewer completion tokens, but a refactor used 70% more, for a net 12% cost increase.
Resource · Benchmarks & research★ Pick
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
An oh-my-claudecode-style gated pipeline for the Muse Code CLI, from deep-interview to verified code, shipped as a native Muse plugin manifest with a skills fallback.
Skill · Coding & dev tools
Sayer Martin
@SayerPM
Today @muse: -booked and paid for airport parking, after recommending the best location for charging a @Tesla -advised on flights for an upcoming trip during my kids’ fall break timeframe (which it found) -read a receipt from @AceHardware and found that the prices were better
X post · Errands & personal agent· ♥ 1
Pi extension that lets you sign in with a Meta Muse Code subscription via device code and pick meta-muse models, verifying the subscription so it never falls back to pay-as-you-go.
Skill · Coding & dev tools· ★ 2
This week we spent about $95 trying to beat our own lineup of reviewing models. One of the candidates was Muse Spark 1.2, and it turned out to be the most interesting model in the whole test. The good, measured: • Among the best we tested at finding real problems. Scored against bugs we already knew were there, it matched our existing lineup, and it caught one real bug our lineup had missed. • Fastest model in our table. Typical answer in 18 seconds, writing at over 220 tokens a second. The speed table from our test (same job, same codebases, 33 runs per model): Model Typical time Answer length (tokens) Writing speed (tok/s) Time follows answer length Time follows question length Muse Spark 1.2 18 s 4,205 222 0.79 barely (0.08) Gemini 3.1 Pro 19 s 2,621 133 0.99 no (0.0) Gemini 3.8 Flash 22 s 1,996 89 0.91 some (0.65) GPT 5.4 29 s 3,058 105 0.95 no (below 0) Grok 4.6 38 s 2,498 61 0.61 no (below 0) Grok 4.7 44 s * 3,176 75 0.98 a little (0.30) Claude Sonnet 5 50 s * 4,471 91 0.59 barely (0.07) * Runs that finished in time only, so the real typical time is higher. The last two columns are correlations: 1 means time rises in step with that length, 0 means no lin
Reddit post · Benchmarks & research
fal
@fal
Muse Image from @aiatmeta is live on fal. It's an agentic image model that plans your composition, uses tools, and self-corrects before you ever see a result. Complex briefs hold together instead of falling apart.
X post · Content & creative· ♥ 87
Refusal-removed Muse Glimmer 30B; the card reports refusals falling from 128/150 (85.3%) to 3/150 (2.0%) on harmful prompts with 0/75 over-refusals.

Resource · Local & open models· ♥ 6
A localhost compatibility gateway that lets the native Muse Code harness run on OpenRouter's muse-spark-1.2-contributor, rewriting only the model name and never falling back silently to another model.
GitHub · Coding & dev tools
Pi coding agent extension that adds Muse Spark 1.2, 1.2-contributor and 1.1 via the Meta Model API, a maintained fork that fixes a false auth warning on pi v0.84+.
Skill · Coding & dev tools· ★ 1
Setup. We run Qwen3.8-Flash-Next NVFP4 as our main agentic model (SGLang, RTX PRO 6000). Before its output reaches a human or gets merged, a second local model acts as judge: reviews the diff, flags real bugs only. Hosted on a 5090 32GB, so we're limited to ~30B NVFP4/GGUF class models. The metric that matters is NOT detection rate — it's false alarms on correct code. A judge that cries wolf gets ignored within a week, exactly like a flaky CI. We built our own battery: 20 injected bugs + 20 clean-but-suspicious snippets (intentional swallowed exceptions, deliberate mutability, weird-but-correct concurrency, short hashes, float patterns that look wrong). Ground-truth labeled, and a stronger model (GLM-5.2 API) arbitrates the judge's prose so scoring isn't vibes. Two passes minimum — single runs lie. Results (40 cases, temp 0, same baremo for everyone): Qwen3.8-27B NVFP4 (no-thinking) • Bugs found: 17/20 • False alarms: 3/20 • Verdict: only pass Nemotron Lightning 30B • Bugs found: 17/20 • False alarms: 0→9 across runs • Verdict: non-reproducible as judge Muse-Glimmer 30B GGUF • Bugs found: 19/20 • False alarms: 12/20 • Verdict: hypercritical Granite 4.1 30B (no-thinki
Reddit post · Benchmarks & research
Vercel added Muse Spark 1.3 (meta/muse-spark-1.3) to AI Gateway on launch day, with 1M-token context, text/image/PDF input, and both standard and contributor pricing tiers.

Site · Coding & dev tools