A vLLM-XPU and DFlash recipe for Muse Glimmer 30B on a single Intel Arc Pro B70, reporting 278 aggregate tok/s across eight clients and an 840.8 tok/s burst peak at concurrency 96.
GitHub · Local & open models★ Pick· ★ 1
Catalog / Use case
About 190 entries on running Muse Glimmer 30B locally: GGUF and MLX quants, DFlash speculative decoding, GPU benchmarks, Mac setups and fine-tunes.
196 builds · page 2 of 5
A vLLM-XPU and DFlash recipe for Muse Glimmer 30B on a single Intel Arc Pro B70, reporting 278 aggregate tok/s across eight clients and an 840.8 tok/s burst peak at concurrency 96.
GitHub · Local & open models★ Pick· ★ 1
Resource · Local & open models· ♥ 8
Cloud Codes tests whether Muse Glimmer on a single 24GB VRAM GPU can compete with cloud coding agents in multi-step tool-calling loops.

Video · Local & open models· ♥ 182
LMSYS Org
@lmsysorg
@AIatMeta's Muse Glimmer (30B dense, open-weights) launches with SGLang day-0 support. We got ~230 tok/s on a single RTX 5090, with NVFP4 + DFlash on. It also works out of the box on @NVIDIAAIDev RTX Pro 6000, DGX Spark, and MLX for Mac. Speed and reliability have always been
X post · Local & open models· ♥ 129
A one-script vLLM setup that serves the roughly 19 GB NVFP4 Muse Glimmer 30B with its vision encoder kept, DFlash speculative decoding using the official drafter head, and up to 256K context on GB10, RTX 5090 or RTX PRO 6000.
GitHub · Local & open models· ★ 8
Red Hat AI's FP8 weight and activation quantized Muse Glimmer 30B (text and image input) for vLLM.

Resource · Local & open models· ♥ 12
NVIDIA hosts a Muse Glimmer 30B endpoint on build.nvidia.com with Python (OpenAI, LangChain), JavaScript and curl examples for the ~29.6B multimodal model with 131K context.

Site · Local & open models
Meta's official Muse-Glimmer-30B repo: ~29.6B dense model with a 1.8B vision encoder, 131K context, Apache 2.0, with vLLM and SGLang serve commands.

Site · Local & open models
Digital Spaceport reviews Muse Glimmer 30B on a 4x 3090 EPYC home server, calling it weaker than Qwen 3.6 27B overall but good at one specific thing.

Video · Local & open models· ♥ 550
A suite of 27 GGUF quantizations of an abliterated Muse Glimmer 30B produced with Heretic v1.4.0.

Resource · Local & open models· ♥ 20
Meta's launch post for Muse Glimmer, an Apache 2.0 30B model for local agents that fits in ~20GB at 4-bit and runs on M4/M5 Max Macs, RTX 5090s or 24–32GB GPUs.

Resource · Local & open models
Hi all, I'm a newbie and trying to assess the performance of some LLMs I'm running locally via oMLX on my MacBook Pro M5pro CPU 15 cores (5 Super and 10 Performance), GPU 16 cores and 48 GB of LPDDR5 RAM. I asked chatGPT guidance to run some tests and check whether the DFlash-based drafter Muse-Glimmer-30B-Assistant might somewhat speedup the base model Muse-Glimmer-30B-4bit. The results show no or negligible improvement with active DFlash acceleration (speedup between 0.90% and 1.16%). The test was structured with three different prompts fed to both the baseline and the dflash-capable model profiles: Technical prose; Python code; Structured JSON a cap of 2048 tokens, no cache, temperature=0. Each inference was repeated three times. Anyone have similar experience? can we simply dump the Assistant as not useful in this hw/sw configuration?
Reddit post · Local & open models
Decensored Muse Glimmer 30B made with Heretic v1.4.0, shipped with a reproduce directory.

Resource · Local & open models· ♥ 15
A correctness-first bring-up of Muse Glimmer 30B on a single Tenstorrent p150 card, with paged KV cache, native DFlash speculative decoding and an OpenAI-compatible server.
GitHub · Local & open models· ★ 4
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Fireworks made Muse Glimmer 30B available on launch day, pitching high-concurrency, cost-effective serving for always-on agents.

Resource · Local & open models
I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8\_O KV cache. It's a joy to use a dense 30B model at this size and still get \~30 tok/s on a VRAM-constrained laptop. It's supposed to be only slightly worse than the official 17GB K-quant at a much smaller footprint, and for my Hermes Agent use case I don't notice a quality difference. It's just much faster. I've tried Qwen 3.8 27B at SC2.20bpw H3 too. Definitely usable but I'm sticking with Unsloth UD\_Q4\_K\_XL for Qwen 3.8 27B because it's mainly for coding.

Reddit post · Local & open models★ Pick
A local Muse Glimmer 30B vision-and-reasoning chat app for high-memory Apple Silicon Macs, running inference through ExecuTorch, MLX/Metal and DFlash with nothing persisted to disk.
GitHub · Local & open models
A reference CUDA worker that serves Meta's official Muse Glimmer 30B GGUF through llama-server on Runpod Serverless load-balancing endpoints or manual Pods, exposing a real OpenAI-compatible API.
GitHub · Local & open models
Sediment's fine-tune of Muse Glimmer 30B that answers factual questions when confident and says "I don't know" otherwise, calibrated for the AA-Omniscience setting.

Resource · Local & open models· ♥ 1
muser is a standalone inference engine for Muse Glimmer 30B on Apple Silicon Metal, with an optional disaggregated lane where an NVIDIA GB10 node prefills in NVFP4 and hands the KV cache to the Mac.
GitHub · Local & open models
Unsloth AI
@UnslothAI
Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM. Muse Glimmer is Apache 2.0 licensed, supports vision and is the strongest agentic model for its size. Run and train the model via Unsloth. GGUF: huggingface.co/unsloth/Muse-G… Guide: unsloth.ai/docs/models/mu…

X post · Local & open models· ♥ 2.8K
Resource · Local & open models· ♥ 15
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Resource · Local & open models· ♥ 12
Unsloth's bitsandbytes 4-bit build of Muse Glimmer 30B for fine-tuning and inference.

Resource · Local & open models· ♥ 14
A coding-specialised speculative-decoding drafter for Muse Glimmer 30B, fine-tuned on software-engineering traces. It raises mean acceptance length to about 6.9 versus 4.0 for the community DSpark, for about 3.8x over no speculation.

Resource · Local & open models· ♥ 1
kwindla
@kwindla
Exciting to see Meta releasing new open weights this week. Meta trained the new 30B dense model Muse Glimmer with "agentic" use cases in mind. This generally means task-oriented, multi-turn, and heavy use of tool calling. I ran Muse Glimmer in the GGUF quant through a bunch of

X post · Local & open models· ♥ 29
An architecture port adding the muse_glimmer model class (vision tower, language model, projector and image processor) to mlx-vlm, so any Muse Glimmer checkpoint runs multimodally on Apple Silicon.
GitHub · Local & open models
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Refusal-direction ablation on Muse-Glimmer-30B that cut refusals from 128/150 to 3/150, adding an agentic-safety evaluation and publishing bf16 and GGUF uncensored weights.
GitHub · Local & open models★ Pick· ★ 1
AI at Meta
@AIatMeta
Muse Glimmer can complete multi-step agentic tasks end-to-end from a single natural language prompt. In this demo, it autonomously discovers a local Home Assistant instance via network tool calls, queries device APIs, writes a responsive HTML/CSS/JS dashboard from scratch, and
X post · Local & open models· ♥ 426
DFlash 2 speculative-decoding draft model for Muse Glimmer 30B from z-lab, run inside a speculative decoding server alongside the target model.

Resource · Local & open models· ♥ 15
A quantization-aware-trained Q4_0 GGUF of Muse Glimmer 30B for llama.cpp. On held-out tokens it measures closer to BF16 than Meta's official Q4_K_M: 0.0213 vs 0.0228 KL and 95.9% vs 95.6% top-token agreement.

Resource · Local & open models· ♥ 3
Muse Glimmer 30B abliterated with Heretic v1.4.0 using self-organizing maps and magnitude-preserving orthogonal ablation.

Resource · Local & open models· ♥ 11
sequelbox's Tachibana-Agent fine-tune of Muse Glimmer 30B, part of a series also released for Gemma 4 12B and Qwen3.6 27B.

Resource · Local & open models· ♥ 4
A 40-chapter book on Muser, an engine that runs Muse Glimmer on Apple Silicon Metal. It covers kquant and DFlash speculative lanes, exact KV-cache replay with kvpack, and GB10 NVFP4 prefill handed off to a Mac, with every number tied to an evidence receipt.
GitHub · Local & open models· ★ 3
SGLang
@sgl_project
SGLang is honored to provide day-0 support for @AIatMeta's Muse Glimmer. ~230 tok/s on a single RTX 5090 with NVFP4 + DFlash on, and it runs out of the box on @NVIDIAAI RTX PRO 6000, DGX Spark, and Apple Silicon via MLX. Huge thanks to the NVIDIA and Meta teams for the
X post · Local & open models· ♥ 95
Demo of a QLoRA adapter for Muse Glimmer 30B that points at UI elements in web screenshots, returning a click point from an instruction like "filter by MATEIN brand".

Site · Local & open models· ♥ 3
Cline
@cline
The successor to Llama is here, and Meta is revitalizing focus on open weights with their new Muse Glimmer - a leading 30B param model designed for always-on local agent use, small enough to run on a Mac or PC with a single GPU. Available in Cline using Ollama now!

X post · Local & open models· ♥ 157
Wavect's guide covers Glimmer's hardware targets (24–32GB) and tool-use scores such as MCP-Atlas 75.5 and SWE-Bench Pro 51.2, and recommends a 20–30 task pilot before production.

Resource · Local & open models
Sam Witteveen covers Meta's open-weight Muse Glimmer 30B release, pointing to the research blog and the Hugging Face collection.

Video · Local & open models· ♥ 405
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
Cobus Greyling's companion repo for Muse Glimmer 30B pairs a long-form intro with an offline-first interactive lab for exploring agent loops, benchmarks and memory envelopes before downloading the weights.
GitHub · Local & open models