Experimental OpenVINO INT4 conversion of Muse Glimmer 30B that needs development builds of Optimum Intel and OpenVINO.

Resource · Local & open models· ♥ 1
7 builds · page 1 of 1
Experimental OpenVINO INT4 conversion of Muse Glimmer 30B that needs development builds of Optimum Intel and OpenVINO.

Resource · Local & open models· ♥ 1
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Box's Complex Work Eval finds Muse Spark 1.1 up to 5-6 points above the top-tier composite on structured work, and nearly 30 points ahead on cost-optimization analysis.

Resource · Benchmarks & research
Hi all, Profile v2.2 is out. It's an open-source optimizer for inference servers. It computes your GPU's roofline ceiling, measures your live server against it, names the bottleneck, gives the flag. You apply. It re-measures. Every fix answers to a number. vLLM only today. More engines next. This release: core rule engine rewritten. Eight rules on a priority DAG with mutual exclusivity. Five alarms fire, four echoes are silenced, one true cause survives. Deterministic. AMD cards are now supported too. Tuning today is chaos: try a flag, wait, squint at a dashboard, repeat for days. Profile turns it into deterministic engineering: measure, fix, verify. Results in a few iterations. Mine took 4, ~30 minutes. My setup: RTX 5090, muse-glimmer 30B, SWE-Bench agents, no DFlash spec decoding. • 81 → 421 tok/s at 25k ctx • $3.41 → $0.65 per 1M output tok • TTFT 224ms (p95 500ms), TPOT 23ms at end of run • 4.72 → 1.08 J/tok https://preview.redd.it/4vazyxkcq6kh1.png?width=2248&format=png&auto=webp&s=77923a489b6f725240d23a7953150b5779260734 One iteration regressed hard: KV thrashing, TTFT 32.8s. Profile labeled it worse. Next fix recovered it. Regressions stay in the record. Watc

Reddit post · Local & open models
Pre-exported ExecuTorch PTE artifacts of Muse Glimmer 30B from Meta, lowered and optimized for specific target backends.

Resource · Local & open models· ♥ 34
Benchmark reports on Muse-Glimmer-30B on NVIDIA DGX Spark covering BF16 to Q4 to DFlash (a 10x speedup) and NVFP4 via SGLang, plus a head-to-head against Qwen3.6-27B.
GitHub · Benchmarks & research
Common Thread Collective's playbook for brands selling through Muse via Shopify: audit product data for agent readability, optimize Shop Pay conversion, build first-party lists, and upgrade attribution for agent-driven traffic.

Resource · Business & commerce
boymanrobshit
@boymanrobshit
Getting a ton of utility out of Muse. This may take a minute to catch on because it’s truly a modality shift but here is what I used it for so far. - optimized my credit card rewards. No less than 1000 bucks saved so far. I have five expensive credit cards. I have no idea
X post · Errands & personal agent· ♥ 113