shipwithmuse

Entries matching “points”

23 builds · page 1 of 1

U

mozilla-ai

u/mozilla-ai

We've been curious how far local models have actually come for agentic coding tasks, so we ran an experiment. Setup: • Model: Muse Glimmer (30B), packaged as a single llamafile • Agent: Hermes coding agent (connected via llamafile's local server mode, zero API keys needed) • Target: Mozilla AI's Otari gateway The Issue: We pointed Hermes at a real, reported bug in Otari (#183) where the gateway returned a vague 502 error on image requests instead of passing through the actual provider error. What the Agent Did: Hermes read the issue, navigated the repo, isolated the bug, created a branch, ran existing tests, wrote a new regression test, and opened a draft PR (#727). All of it ran locally and offline, with zero code written by hand. It's still draft PR territory rather than a merged fix, but it's a solid signal that ~30B local models are getting genuinely capable for real dev workflows, not just toy demos. Video walkthrough of the run: https://youtu.be/5GAgbT-XgHU?si=vJqEDGm9hssCO5-M Happy to answer questions about the setup, model performance, or how Hermes handled tool calling!

Reddit post · Local & open models★ Pick

Local Muse Glimmer agent opens a real pull request

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

B

blog.box.com

blog.box.com

Box's Complex Work Eval finds Muse Spark 1.1 up to 5-6 points above the top-tier composite on structured work, and nearly 30 points ahead on cost-optimization analysis.

Resource · Benchmarks & research

Box eval: Muse Spark 1.1 on real enterprise work

S

@samwitteveenai

@samwitteveenai

Sam Witteveen covers Meta's open-weight Muse Glimmer 30B release, pointing to the research blog and the Hugging Face collection.

Video · Local & open models· ♥ 405

Sam Witteveen walks through Muse Glimmer 30B

Arena.ai

@arena

Muse Spark 1.2 (xHigh) by @AIatMeta is #14 in the Code Arena: WebDev, with 1,545 pts! This is an improvement from Muse Spark 1.1 at #18. See its biggest gains by category in the post below. Congrats to the @AIatMeta team on this release!

X post · Benchmarks & research· ♥ 380

Spark 1.2 in Code Arena: WebDev

Arena.ai

@arena

Muse Spark 1.1 has entered the Code Arena: Frontend at #9! Muse Spark 1.1 reshapes the cost-performance Pareto Frontier by scoring 1541 at a blended $3.5M ($1.25 per input MToken, $4.25 per output MToken). This is frontier performance at a fraction of the price. Congrats to

X post · Benchmarks & research· ♥ 302

Spark 1.1 in Code Arena: Frontend

A

artificialanalysis.ai

artificialanalysis.ai

Artificial Analysis puts Muse Spark 1.1 at 51 on its Intelligence Index, 8 points above 1.0, and calls it cost and token efficient versus peers.

Resource · Benchmarks & research

Artificial Analysis: Muse Spark 1.1 scores 51

R

runtimewire.com

runtimewire.com

RuntimeWire's head-to-head eval has Muse Spark 1.1 beating Claude Opus 4.8 by 10 points overall with a 95% confidence verdict.

Resource · Benchmarks & research

Head to head: Muse Spark 1.1 vs Claude Opus 4.8

Artificial Analysis

@ArtificialAnlys

At 30B parameters, Muse Glimmer sits near the Intelligence vs Parameters frontier for open weights models: 5 points above Gemma 4 31B (Reasoning) at the same size, effectively matching Kimi K2.5 (Reasoning) at 33x fewer total parameters, and just behind Qwen3.6 27B (Reasoning),

X post · Benchmarks & research· ♥ 60

Glimmer on the intelligence vs parameters frontier

Arena.ai

@arena

Meta Muse Video just entered the Video Arena at #3. @AIatMeta’s new video model scored 1459 in the Text-to-Video Arena. It outperforms Alibaba’s HappyHorse 1.0 by +30pts and ranks ahead of Grok Imagine, Sora 2 Pro and Google Veo-3.1 models. Meta has now reached the video AI

X post · Benchmarks & research· ♥ 611

Muse Video enters Video Arena at #3

@harris-ryder

@harris-ryder

Dowser is a read-only MCP connector for Muse that surfaces open class-action settlements, product recalls with refunds, unused card credits, expiring points and birthday freebies, each with an official link.

Skill · Errands & personal agent

Dowser: find the money hiding in your life

Dan

@DanDr1s

Meta’s Muse Spark 1.3 just scored 62 on Artificial Analysis’ Intelligence Index. That ties Claude Fable 5, but Muse costs 8x less for input and nearly 12x less for output. It also scores above GPT-5.6 Sol, Grok 4.6, Kimi K3, and Gemini 3.8 Flash. Meta is suddenly in the top

X post · Benchmarks & research· ♥ 223

Spark 1.3 matches Fable 5 at a fraction of the price

G

gscwizard.com

gscwizard.com

GSC Wizard's Muse connector gives Meta Muse real Google Search Console clicks, impressions, rankings and CTR through its MCP server; setup takes three steps and about three minutes.

Matt Deitke

@mattdeitke

The multimodal grounding abilities combined with web development in Muse Spark are so good! 🎨 Prompt: <image of dogs> Make a website where you point to each of the dogs and say their breed.

X post · Apps & websites· ♥ 102

Dog breed pointer website from a photo

H

hugging-apps

hugging-apps

Demo of a QLoRA adapter for Muse Glimmer 30B that points at UI elements in web screenshots, returning a click point from an instruction like "filter by MATEIN brand".

Site · Local & open models· ♥ 3

Muse Glimmer click grounding demo

M

merve

merve

QLoRA adapter that teaches Muse Glimmer 30B to return click points on web screenshots from instructions, trained on the MolmoWeb dataset.

Resource · Local & open models· ♥ 6

Muse Glimmer QLoRA click grounding

Lucas Switzer

@lucas_switzer

A cool thing about @Muse is that every Muse runs on its own Linux machine. So you and your Muse can do normal Linux stuff together — like build your own terminal!

X post · Coding & dev tools· ♥ 803

Custom terminal built with a Muse

Arena.ai

@arena

Muse Spark 1.3 (xHigh) just landed @AIatMeta back in the top 10 models for Code Arena: WebDev! This release is ~#10 in Code Arena: WebDev with 1623 pts (AutoEval). It’s performance is par with Claude Fable 5 at #8 (1628 pts) and GPT-5.6 Sol at #11 (1616 pts). Muse Spark 1.3

X post · Benchmarks & research· ♥ 610

Spark 1.3 in Code Arena: WebDev top 10

Arena.ai

@arena

Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4 in the Text Arena (1498 pts), and has reshaped the Pareto frontier! It is priced at $1.25/$4.25 per MToken. Congrats again to the @AIatMeta team on this release!

X post · Benchmarks & research· ♥ 422

Spark 1.2 reaches #4 in Text Arena

Artificial Analysis

@ArtificialAnlys

Meta's Muse Spark 1.1 scores 51 on the Artificial Analysis Intelligence Index and is cost and token efficient compared to its peers Muse Spark 1.1 (xhigh) improves 8 points over Muse Spark 1.0 (43) in three months. It is effectively tied with GLM-5.2 (max), GPT-5.4 (xhigh), and

X post · Benchmarks & research· ♥ 708

Spark 1.1 scores 51 on the Intelligence Index

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Artificial Analysis

@ArtificialAnlys

Last week the Intelligence Index vs Cost Pareto frontier moved out substantially. Claude Fable 5.1, Muse Spark 1.3, and GPT-6 Astra each set a new point in efficient intelligence Link to analysis: artificialanalysis.ai/#intelligence-…

X post · Benchmarks & research· ♥ 732

Spark 1.3 moves the cost Pareto frontier

@evangit2

@evangit2

A guide and tooling for pointing Muse Code's meta provider at any OpenAI-compatible gateway with API-key auth and no Meta login, including workarounds for 1.0.1's bearer-withholding change.

@nano-muse

@nano-muse

nanoMuse is an open-source personal AI agent inspired by Meta Muse that runs on Android with a shell, browser, MCP, skills and scheduled tasks, and can be pointed at DeepSeek, OpenAI or Ollama models.

GitHub · Agents & automation

nanoMuse open-source personal agent

@dominicletz

@dominicletz

A self-contained visual report comparing CursorBench 4.0 score against cost per task for Opus 5.5, Fable 5.1, Grok 4.7 and Muse Spark 1.3, where Muse Spark 1.3 Max scores 41.6% at $2.64 per task. GPT-6 points are clearly marked as estimates.

GitHub · Benchmarks & research

CursorBench 4.0 score-vs-cost report