Refusal-direction ablation on Muse-Glimmer-30B that cut refusals from 128/150 to 3/150, adding an agentic-safety evaluation and publishing bf16 and GGUF uncensored weights.
GitHub · Local & open models★ Pick· ★ 1
8 builds · page 1 of 1
Refusal-direction ablation on Muse-Glimmer-30B that cut refusals from 128/150 to 3/150, adding an agentic-safety evaluation and publishing bf16 and GGUF uncensored weights.
GitHub · Local & open models★ Pick· ★ 1
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
A small Inspect evaluation on Muse Glimmer that tests whether incentive framing changes what an agent reports to its principal when the evidence is held constant; the author reports a null result.
GitHub · Benchmarks & research
You.com's harness for evaluating Muse Glimmer 30B on DeepSearchQA with You.com MCP tools inside pi sessions. A custom RLM v5 extension reached F1 0.8054 on 50 tasks (0.6910 over 900x3) versus 0.50 for plain skill injection.
GitHub · Benchmarks & research
John Yang
@jyangballin
My favorite demo from the launch: muse spark 1.1 + opencode runs evaluation of *itself* + mini-SWE-agent (by @KLieret, @closji, urs truly) on DeepSWE!
X post · Benchmarks & research· ♥ 40
A project that fine-tunes a DSpark speculator for Muse Glimmer 30B on on-policy coding and agentic traces to raise acceptance length on agentic workloads, evaluated with Terminal-Bench.
GitHub · Local & open models
An effort to make Muse Glimmer 30B actually use a 512k-token context (4x native) as a ~17GB GGUF in 32GB VRAM, trained on DGX Spark and evaluated with RULER-style retrieval tests.
GitHub · Local & open models· ★ 1
A reproducible evaluation of a stochastic defect where Muse Spark 1.1, under some agent-like request envelopes, wrote to claude-smoke.txt when asked for muse-smoke.txt; a neutral control kept 80/80 filenames intact.
GitHub · Benchmarks & research
Compared diarization models on 15 mock doctor-patient consultations (~2.4 h): Meta Muse Voice Transcribe scored 13.04% DER at ~92 s per request via API, behind Pyannote (2.89%) and Nemotron 3 (4.80%).
Reddit post · Benchmarks & research