shipwithmuse

Entries matching “fix”

28 builds · page 1 of 1

Unsloth AI

@UnslothAI

2-bit Muse Glimmer GGUF managed to call 100+ tools on just 14GB RAM. 🔥 Muse Glimmer did a complete repo bug hunt for 5 mins nonstop with: evidence, repro, fix, tests and a PR writeup. Run and train it in Unsloth. GitHub repo: github.com/unslothai/unsl…

X post · Local & open models★ Pick· ♥ 1.6K

2-bit Glimmer GGUF: 100+ tool calls on 14GB RAM

Ryan Fox

@wailord

support chat use cases are some of the coolest with @Muse! a few months ago I was double-charged for concert tickets and my agent fixed it all with support in like 30m

X post · Errands & personal agent· ♥ 71

Concert double-charge fixed via support chat

Paweł Huryn

@PawelHuryn

So, I finally tested Muse Spark 1.3. 2 real repos, 105 planted bugs, find and fix what you can. Original harness and API. Big surprise: Muse Spark 1.3 (max): 33 Fable 5.1 (high): 33 Grok 4.6 (xhigh): 27 Opus 5 (max): 27 Muse Spark 1.3 (high): 19 Meta joined the frontier.

X post · Benchmarks & research· ♥ 694

105 planted bugs: Spark 1.3 vs frontier

Confucius

@ConfuEth

Asked @Muse to file pothole repair requests with the county. Took 30 seconds, zero hassle. Honestly assumed they'd disappear into a void. Five days later: all fixed. 🤯

X post · Errands & personal agent· ♥ 7

Pothole repair requests filed with the county

@47thtechcorner

@47thtechcorner

A YouTube tutorial companion repo where Muse Glimmer 30B (a 2-bit GGUF, run offline) looks at screenshots of badly designed web pages and generates and self-repairs modern replacement code.

GitHub · Local & open models

Glimmer UI auto-fixer

@everyai-com

@everyai-com

A set of small fixes that let Muse Spark run in pi, Agent Orchestrator, opencode and other OpenAI-compatible tools, handling its non-standard SSE event and missing model-catalog entries.

GitHub · Coding & dev tools

Muse Spark Anywhere

@tanishq-dubey

@tanishq-dubey

A reproducible Apple Silicon harness that runs six fixed quality tasks against MLX quantizations of Muse Glimmer 30B and records scores, tokens per second, peak memory and load time.

GitHub · Benchmarks & research

Muse Glimmer 30B MLX benchmark harness

@EclipseAditya

@EclipseAditya

Pi coding agent extension that adds Muse Spark 1.2, 1.2-contributor and 1.1 via the Meta Model API, a maintained fork that fixes a false auth warning on pi v0.84+.

Skill · Coding & dev tools· ★ 1

pi-muse-spark extension

Jon ONeill

@HouseHackerJon

@Muse can’t fix my AC (yet) But they did file a police report And they filed my insurance claim for me And they’re gonna go get me some quotes What a wild day this has turned out to be

X post · Errands & personal agent· ♥ 4

Police report and insurance claim filed after an AC problem

@krtarunsingh

@krtarunsingh

An experiment running Muse Glimmer 30B Q4_K_M via llama.cpp on an RTX 4060 laptop with 8 GB VRAM, testing autonomous Python bug fixing, tool-failure recovery and multimodal invoice extraction.

GitHub · Local & open models

Muse Glimmer on an 8 GB RTX 4060 laptop

yahav

@YahavSal

Quick context: I own a communal sauna and cold plunge studio. Mindbody is the booking and membership software studios like mine run on. I spent an hour on the phone with Mindbody support. They couldn't fix my problem. I asked @Muse to look at my actual setup. Five minutes later

X post · Business & commerce

Sauna studio's Mindbody booking bug fixed

U

GammaRxBurst

u/GammaRxBurst

Solution: So after more research it seems that this KHO was recently activated in the latest Ubuntu kernel which zorin is based on. It seems to be for live updates of kernel and whatever but it puts quite a bit of strain on system that requires a lot of memory like AI work load. So I edited my GRuB file to turn it off. I was lucky that somehow my 7.0.0-30 kernel had a bug that disabled the KHO also thank you to Muse Spark AI for able to run all the trace and helped me figuring it out. KHO locked out 2gb of RAM causing my Llama cpp out of memory error. The original post So I been running Zorin OS as my semi daily driver for learning AI and fun. I been plague with odd behavior as I run my AI on knife edge consuming nearly all 16 GB of Ram and 8 GB of VRAM. I found something odd, that my AI would failed to load depending on which kernel. It turn out at least one of the culprit was KHO it somehow ate about 2GB of RAM. I found this because my Kernel 7.0.0-30 and the old 6.17 load my AI just fine but kernel 7.0.0-28 and the latest 7.0.0.-31 always failed to load. Below is the AI analysis that I had it trace issue through my system: - `-30` boot `-2`: `Memory: 15384980K avail / 8261

Reddit post · Coding & dev tools

Muse Spark traces a kernel bug eating 2GB of RAM

@vstaln

@vstaln

Copyable block of the engineering conventions Muse Spark was co-trained with. Per the README, in Cline's harness it cut a real bug fix from 19.7M tokens, 49 min and $7.69 to 7.2M tokens, 24 min and $3.25.

Skill · Coding & dev tools

Muse Code system prompt block

U

Electronic_Back1502

u/Electronic_Back1502

Disclaimer. This is the first time I've used Muse or VSCode as a harness. The reason I am using VSCode as a harness is this is a research project for my job, and we only have VSCode, Codex, and Claude Code approved for harnesses. I ran it in a folder with only one HTML file (800 lines) that is a Roblox-style COD game. I just gave it a prompt "Can you fix the bugs in the file". It read the file 3 times, found one bug, started to fix it, then got stuck reading the same 10 lines over and over. I imagine it's one of these three issues. • It's a prompt error, being way too vague/open ended for the capabilities of a smaller model. I tried again, with a specific prompt to fix a specific bug, and it still just ends up so confused, trying to grep/find the file despite already having read it, and trying to find the code inside of the file. • It's a limitation of small models running with a large harness/having way too much going on. I tried running it with Pi with its default prompt, and it just got stuck doing tool calls and never actually read the file. Tried running this just directly in the Unsloth Desktop UI with no harness but it failed to parse the file I inputted and tried to gen

Reddit post · Benchmarks & research

Glimmer Q8 looping in a VS Code harness

U

PathfinderTactician

u/PathfinderTactician

I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context, whilst Qwen3.6 27B can only run at 147,500 context - full GPU offload in both cases. Both models have been coding on an enterprise-grade web application. Detailed report of each model (warning - includes AI generated content): Diagnostic quality - comparable. Both have shown genuinely good root-cause work when they apply themselves. Qwen found coding issue and worked to fix things cleanly. Muse Glimmer correctly traced bugs and even caught something that a Frontier model missed after more than 10 rounds of review. Neither one is weak at diagnosis. Implementation reliability - Qwen ahead. Qwen did introduce real regressions into the coding along the way (eg. severe zone-scope refactor regression, and case-sensitivity regression) but each one eventually got fixed properly once caught, usually within one or two corrective rounds. Muse Glimmer did land fixes that were clean and verified true to spec. However, when working in a complex environment exceeding 200k context, Muse Glimmer fa

Reddit post · Benchmarks & research

BF16 Muse Glimmer vs Qwen3.6 27B on real code

U

KitchenAmoeba4438

u/KitchenAmoeba4438

Eleven matched on/off pairs across Gemma 4 and Qwen3.6, holding model, quant, card, corpus and concurrency fixed inside each pair. Speed: 1.65x to 2.54x, every pair. Accuracy: nothing the paired intervals could separate from ordinary run-to-run movement. Muse Glimmer is the one that lost. Meta's matching DFlash drafter made the same 7900 XTX 9% slower, keeping 24.55% of drafted tokens against roughly four in five for the Gemma and Qwen heads. Acceptance fell across the run instead of warming up. Meta's model card reports 3.1x on an RTX 5090, and there are open llama.cpp issues for DFlash on AMD and under Vulkan, so I read it as the backend rather than the model. Acceptance turned out to be a poor predictor of speed. It moved under four points across five models while the multiple nearly doubled. What tracks the multiple is how bandwidth-bound the target is: a heavier quant gains more, and the two mixture-of-experts pairs gained least. Worth knowing before you benchmark anything: -md mtp-head.gguf silently disables speculation. Use -hf REPO:QUANT -hfd REPO, then read speculative from /slots and confirm it is true. Per-pair table, intervals, acceptance counters and the raw predic

Reddit post · Local & open models

On/off speculative decoding test incl. Glimmer

U

Graemer71

u/Graemer71

OK, for context, I have Claude Code desktop app driving the CLI and orchestrating the code and verification tasks to try to save tokens. So Claude runs things, a Deepseek 4.1 Flash (cloud) session does the planning, Qwen 3.8 27b Q8 does the boiler plate coding and Muse Glimmer sanity checks the code and pushes any issues back to Qwen. If there are issues Qwen and Glimmer can't agree on, Deepseek validates. If Deepseek can't sort it out, it goes back to Claude. This had been working fine, but then in the last few days token use spiked, tasks that used to take 10 minutes were taking an hour or more and Qwen started going into more and more reasoning loops. It seems that since I last checked (on 12th September) the CLI changed. I used to strip unnecessary tool calls from the prompt using --disallowedTools and enabledPlugins: false. It would seem that these no longer work. In the end I got Claude to build a request-dumping diagnostic server, that actually measured the payload bytes, and confirmed --tools (an allowlist) is the flag that works now: 55→7 tools, 161KB→24KB, byte-verified. It also caught something specific to my workflow running the wrapper from inside an already-active

U

mozilla-ai

u/mozilla-ai

We've been curious how far local models have actually come for agentic coding tasks, so we ran an experiment. Setup: • Model: Muse Glimmer (30B), packaged as a single llamafile • Agent: Hermes coding agent (connected via llamafile's local server mode, zero API keys needed) • Target: Mozilla AI's Otari gateway The Issue: We pointed Hermes at a real, reported bug in Otari (#183) where the gateway returned a vague 502 error on image requests instead of passing through the actual provider error. What the Agent Did: Hermes read the issue, navigated the repo, isolated the bug, created a branch, ran existing tests, wrote a new regression test, and opened a draft PR (#727). All of it ran locally and offline, with zero code written by hand. It's still draft PR territory rather than a merged fix, but it's a solid signal that ~30B local models are getting genuinely capable for real dev workflows, not just toy demos. Video walkthrough of the run: https://youtu.be/5GAgbT-XgHU?si=vJqEDGm9hssCO5-M Happy to answer questions about the setup, model performance, or how Hermes handled tool calling!

Reddit post · Local & open models★ Pick

Local Muse Glimmer agent opens a real pull request

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@caicaitou001

@caicaitou001

A config and write-up that makes muse-spark-1.2-contributor via OpenCode Go work in DeepSeek Harness, fixing empty first-turn output and multi-turn thinking replay errors.

GitHub · Coding & dev tools

Muse Spark fix for DeepSeek Harness

@dp1x

@dp1x

A small local proxy between Grok Build and OpenCode Zen that fixes muse-spark-1.2-contributor-free streaming, with a technical report of protocol captures and root cause.

GitHub · Coding & dev tools

compat-muse

@nuocmam

@nuocmam

A same-prompt bakeoff site that starts with a playable penguin Tetris generated by Muse Spark 1.3 in OpenCode, with a fixed prompt for comparing other models.

V

@venelin_valkov

@venelin_valkov

Venelin Valkov pairs Muse Glimmer with Hermes Agent on llama.cpp for a fully free local agent, testing whether a better harness fixes the model's mixed early reviews.

Video · Local & open models

Muse Glimmer + Hermes Agent local tutorial

@seb-patron

@seb-patron

Adversarial-review and fix-verification skills for Muse, built after two round-1 reviewers approved a diff with 0 blocking findings while an independent review found 13 real issues.

Skill · Coding & dev tools

Review-loop skills for Muse sessions

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@Abhishektenneti

@Abhishektenneti

A reproducible record of running the 17GB Muse Glimmer 30B GGUF with llama.cpp and Metal on a 24GB M4 Pro MacBook Pro, with notes on mistakes, fixes and how local inference works.

GitHub · Local & open models

Muse Glimmer on a 24GB M4 Pro MacBook

thehype.

@thehypedotnews

muse spark 1.2 vs gpt 5.6 sol vs kimi k3 vs grok 4.5 – on two landing pages four coding agents built two desktop landing pages from scratch, then had to open them in a real browser, find their own bugs and fix them before they were allowed to hand anything over the setup: each

X post · Benchmarks & research· ♥ 64

Landing pages built and self-debugged by Muse Code

U

gargetisha

u/gargetisha

We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added them to the Cline harness. TL;DR of this special prompting: - Trust source code over the user prompt, so read every call site and existing tests before starting the task - Weigh edge and error cases as heavily as the happy path - Always reproduce the bug before fixing - Don't trust the first passing test suite, and verify suspicious looking half-baked tests - Never stop at just editing, keep working until the change is verified complete. We then asked this modified harness to fix a real bug from our repo, and compared the results to the original Cline agent harness. Results: - Used 2.7x fewer tokens (19.7M → 7.2M) - Finished 2x faster (49min → 24min) - Cost 2.4x less ($7.69 → $3.25) Same Muse Spark 1.2 model, same task, only the prompting changed. Incredible how much of a performance gain Meta was able to achieve training it on these special instructions!

Reddit post · Coding & dev tools

Muse Code's prompt rules inside Cline

U

DanC403

u/DanC403

Got Muse Glimmer 30B running locally using the UD-Q2-K-XL quant paired with DFlash speculative decoding, and the results on modest hardware are pretty impressive. Hardware Setup Host: Ryzen 5 4600G with 96GB DDR4 RAM running headless Debian Trixie. Guest VM: QEMU/KVM assigned 4 cores and 32GB RAM, running Debian Sid with ROCm 7.2. GPU: AMD Radeon RX 7600 XT 16GB passed through to the VM, built llama.cpp fresh from master targeting gfx1102 and gfx1201 via HIP. Context Size: Set to 62144 tokens. Processed 14685 total tokens at roughly 308 tokens per second prompt evaluation and 20 tokens per second generation speed. Speculative Decoding: Using the dflash-kquant draft model with spec-draft-n-max set to 2. Fed it a clean context slate consisting of eight JavaScript files and one HTML file alongside the problem description. On the first turn, it identified and output the necessary diff snippets. A quick follow-up prompt telling it to stop being lazy and output the complete updated files yielded functional code that dropped straight in and worked on the first try.

Reddit post · Local & open models

Muse Glimmer on a 16GB RX 7600 XT

M

dev.meta.ai

dev.meta.ai

Meta's developer guide to first calls, coding primitives and agentic patterns; its multi-agent recipe turns a one-line idea into an app plus launch package, and Spark fixed all 5 planted bugs in 7.6 turns on average.

Resource · Coding & dev tools

Build with Muse Spark on Meta Model API

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page