shipwithmuse

Catalog / Use case

Local & open models

About 190 entries on running Muse Glimmer 30B locally: GGUF and MLX quants, DFlash speculative decoding, GPU benchmarks, Mac setups and fine-tunes.

196 builds · page 2 of 5

@mgaruccio

@mgaruccio

A vLLM-XPU and DFlash recipe for Muse Glimmer 30B on a single Intel Arc Pro B70, reporting 278 aggregate tok/s across eight clients and an 840.8 tok/s burst peak at concurrency 96.

GitHub · Local & open models★ Pick· ★ 1

Muse Glimmer on one Arc Pro B70

LMSYS Org

@lmsysorg

@AIatMeta's Muse Glimmer (30B dense, open-weights) launches with SGLang day-0 support. We got ~230 tok/s on a single RTX 5090, with NVFP4 + DFlash on. It also works out of the box on @NVIDIAAIDev RTX Pro 6000, DGX Spark, and MLX for Mac. Speed and reliability have always been

X post · Local & open models· ♥ 129

SGLang serves Glimmer at ~230 tok/s on one RTX 5090

@MiaAI-Lab

@MiaAI-Lab

A one-script vLLM setup that serves the roughly 19 GB NVFP4 Muse Glimmer 30B with its vision encoder kept, DFlash speculative decoding using the official drafter head, and up to 256K context on GB10, RTX 5090 or RTX PRO 6000.

GitHub · Local & open models· ★ 8

Muse Glimmer 30B NVFP4 for DGX Spark and RTX 5090

R

RedHatAI

RedHatAI

Red Hat AI's FP8 weight and activation quantized Muse Glimmer 30B (text and image input) for vLLM.

Resource · Local & open models· ♥ 12

Red Hat AI FP8-block Muse Glimmer

N

build.nvidia.com

build.nvidia.com

NVIDIA hosts a Muse Glimmer 30B endpoint on build.nvidia.com with Python (OpenAI, LangChain), JavaScript and curl examples for the ~29.6B multimodal model with 131K context.

Site · Local & open models

Muse Glimmer 30B on NVIDIA build

M

huggingface.co

huggingface.co

Meta's official Muse-Glimmer-30B repo: ~29.6B dense model with a 1.8B vision encoder, 131K context, Apache 2.0, with vLLM and SGLang serve commands.

Site · Local & open models

Muse Glimmer 30B weights on Hugging Face

D

@DigitalSpaceport

@DigitalSpaceport

Digital Spaceport reviews Muse Glimmer 30B on a 4x 3090 EPYC home server, calling it weaker than Qwen 3.6 27B overall but good at one specific thing.

Video · Local & open models· ♥ 550

Muse Glimmer 30B on a 4x RTX 3090 local rig

0

0bserverx

0bserverx

A suite of 27 GGUF quantizations of an abliterated Muse Glimmer 30B produced with Heretic v1.4.0.

Resource · Local & open models· ♥ 20

Heretic Uncensored Muse Glimmer GGUF

M

research.meta.ai

research.meta.ai

Meta's launch post for Muse Glimmer, an Apache 2.0 30B model for local agents that fits in ~20GB at 4-bit and runs on M4/M5 Max Macs, RTX 5090s or 24–32GB GPUs.

Resource · Local & open models

Introducing Muse Glimmer

U

OkSea7809

u/OkSea7809

Hi all, I'm a newbie and trying to assess the performance of some LLMs I'm running locally via oMLX on my MacBook Pro M5pro CPU 15 cores (5 Super and 10 Performance), GPU 16 cores and 48 GB of LPDDR5 RAM. I asked chatGPT guidance to run some tests and check whether the DFlash-based drafter Muse-Glimmer-30B-Assistant might somewhat speedup the base model Muse-Glimmer-30B-4bit. The results show no or negligible improvement with active DFlash acceleration (speedup between 0.90% and 1.16%). The test was structured with three different prompts fed to both the baseline and the dflash-capable model profiles: Technical prose; Python code; Structured JSON a cap of 2048 tokens, no cache, temperature=0. Each inference was repeated three times. Anyone have similar experience? can we simply dump the Assistant as not useful in this hw/sw configuration?

Reddit post · Local & open models

Testing the Glimmer DFlash drafter on an M5 Mac

D

darkc0de

darkc0de

Decensored Muse Glimmer 30B made with Heretic v1.4.0, shipped with a reproduce directory.

Resource · Local & open models· ♥ 15

Muse Glimmer 30B heretic

@Codys12

@Codys12

A correctness-first bring-up of Muse Glimmer 30B on a single Tenstorrent p150 card, with paged KV cache, native DFlash speculative decoding and an OpenAI-compatible server.

GitHub · Local & open models· ★ 4

Muse Glimmer on Tenstorrent p150

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

F

fireworks.ai

fireworks.ai

Fireworks made Muse Glimmer 30B available on launch day, pitching high-concurrency, cost-effective serving for always-on agents.

Resource · Local & open models

Muse Glimmer 30B on Fireworks AI

U

PyaesoneP

u/PyaesoneP

I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8\_O KV cache. It's a joy to use a dense 30B model at this size and still get \~30 tok/s on a VRAM-constrained laptop. It's supposed to be only slightly worse than the official 17GB K-quant at a much smaller footprint, and for my Hermes Agent use case I don't notice a quality difference. It's just much faster. I've tried Qwen 3.8 27B at SC2.20bpw H3 too. Definitely usable but I'm sticking with Unsloth UD\_Q4\_K\_XL for Qwen 3.8 27B because it's mainly for coding.

Reddit post · Local & open models★ Pick

Muse Glimmer 30B on a 12GB laptop GPU

@mapleroyal

@mapleroyal

A local Muse Glimmer 30B vision-and-reasoning chat app for high-memory Apple Silicon Macs, running inference through ExecuTorch, MLX/Metal and DFlash with nothing persisted to disk.

GitHub · Local & open models

Muse Glimmer MLX playground

@cezaronx

@cezaronx

A reference CUDA worker that serves Meta's official Muse Glimmer 30B GGUF through llama-server on Runpod Serverless load-balancing endpoints or manual Pods, exposing a real OpenAI-compatible API.

GitHub · Local & open models

Muse Glimmer Runpod serverless worker

S

SedimentLabs

SedimentLabs

Sediment's fine-tune of Muse Glimmer 30B that answers factual questions when confident and says "I don't know" otherwise, calibrated for the AA-Omniscience setting.

Resource · Local & open models· ♥ 1

Pebble 1 30B calibrated reasoning model

@High-Performance-AI-Lab

@High-Performance-AI-Lab

muser is a standalone inference engine for Muse Glimmer 30B on Apple Silicon Metal, with an optional disaggregated lane where an NVIDIA GB10 node prefills in NVFP4 and hands the KV cache to the Mac.

GitHub · Local & open models

muser inference engine for Muse Glimmer

Unsloth AI

@UnslothAI

Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM. Muse Glimmer is Apache 2.0 licensed, supports vision and is the strongest agentic model for its size. Run and train the model via Unsloth. GGUF: huggingface.co/unsloth/Muse-G… Guide: unsloth.ai/docs/models/mu…

X post · Local & open models· ♥ 2.8K

Unsloth GGUFs and guide for Muse Glimmer

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

U

unsloth

unsloth

Unsloth's bitsandbytes 4-bit build of Muse Glimmer 30B for fine-tuning and inference.

Resource · Local & open models· ♥ 14

Unsloth bnb 4-bit Muse Glimmer

S

huggingface.co

huggingface.co

A coding-specialised speculative-decoding drafter for Muse Glimmer 30B, fine-tuned on software-engineering traces. It raises mean acceptance length to about 6.9 versus 4.0 for the community DSpark, for about 3.8x over no speculation.

Resource · Local & open models· ♥ 1

Muse Glimmer 30B DSpark coding drafter

kwindla

@kwindla

Exciting to see Meta releasing new open weights this week. Meta trained the new 30B dense model Muse Glimmer with "agentic" use cases in mind. This generally means task-oriented, multi-turn, and heavy use of tool calling. I ran Muse Glimmer in the GGUF quant through a bunch of

X post · Local & open models· ♥ 29

Glimmer GGUF through agentic tool-calling tests

@nicedreamzapp

@nicedreamzapp

An architecture port adding the muse_glimmer model class (vision tower, language model, projector and image processor) to mlx-vlm, so any Muse Glimmer checkpoint runs multimodally on Apple Silicon.

GitHub · Local & open models

mlx-vlm support for Muse Glimmer

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@TrevorS

@TrevorS

Refusal-direction ablation on Muse-Glimmer-30B that cut refusals from 128/150 to 3/150, adding an agentic-safety evaluation and publishing bf16 and GGUF uncensored weights.

GitHub · Local & open models★ Pick· ★ 1

Muse Glimmer abliteration

AI at Meta

@AIatMeta

Muse Glimmer can complete multi-step agentic tasks end-to-end from a single natural language prompt. In this demo, it autonomously discovers a local Home Assistant instance via network tool calls, queries device APIs, writes a responsive HTML/CSS/JS dashboard from scratch, and

X post · Local & open models· ♥ 426

Glimmer builds a Home Assistant dashboard

Z

z-lab

z-lab

DFlash 2 speculative-decoding draft model for Muse Glimmer 30B from z-lab, run inside a speculative decoding server alongside the target model.

Resource · Local & open models· ♥ 15

DFlash 2 drafter for Muse Glimmer

Q

huggingface.co

huggingface.co

A quantization-aware-trained Q4_0 GGUF of Muse Glimmer 30B for llama.cpp. On held-out tokens it measures closer to BF16 than Meta's official Q4_K_M: 0.0213 vs 0.0228 KL and 95.9% vs 95.6% top-token agreement.

Resource · Local & open models· ♥ 3

Muse Glimmer 30B native Q4_0 QAT GGUF

M

MuXodious

MuXodious

Muse Glimmer 30B abliterated with Heretic v1.4.0 using self-organizing maps and magnitude-preserving orthogonal ablation.

Resource · Local & open models· ♥ 11

Muse Glimmer SOMPOA heresy

S

sequelbox

sequelbox

sequelbox's Tachibana-Agent fine-tune of Muse Glimmer 30B, part of a series also released for Gemma 4 12B and Qwen3.6 27B.

Resource · Local & open models· ♥ 4

Muse Glimmer Tachibana-Agent

@High-Performance-AI-Lab

@High-Performance-AI-Lab

A 40-chapter book on Muser, an engine that runs Muse Glimmer on Apple Silicon Metal. It covers kquant and DFlash speculative lanes, exact KV-cache replay with kvpack, and GB10 NVFP4 prefill handed off to a Mac, with every number tied to an evidence receipt.

GitHub · Local & open models· ★ 3

The Muser book: how to write an inference engine

SGLang

@sgl_project

SGLang is honored to provide day-0 support for @AIatMeta's Muse Glimmer. ~230 tok/s on a single RTX 5090 with NVFP4 + DFlash on, and it runs out of the box on @NVIDIAAI RTX PRO 6000, DGX Spark, and Apple Silicon via MLX. Huge thanks to the NVIDIA and Meta teams for the

X post · Local & open models· ♥ 95

SGLang day-0 serving for Glimmer

H

hugging-apps

hugging-apps

Demo of a QLoRA adapter for Muse Glimmer 30B that points at UI elements in web screenshots, returning a click point from an instruction like "filter by MATEIN brand".

Site · Local & open models· ♥ 3

Muse Glimmer click grounding demo

Cline

@cline

The successor to Llama is here, and Meta is revitalizing focus on open weights with their new Muse Glimmer - a leading 30B param model designed for always-on local agent use, small enough to run on a Mac or PC with a single GPU. Available in Cline using Ollama now!

X post · Local & open models· ♥ 157

Muse Glimmer in Cline

W

wavect.io

wavect.io

Wavect's guide covers Glimmer's hardware targets (24–32GB) and tool-use scores such as MCP-Atlas 75.5 and SWE-Bench Pro 51.2, and recommends a 20–30 task pilot before production.

Resource · Local & open models

Muse Glimmer 30B: is it production-ready?

S

@samwitteveenai

@samwitteveenai

Sam Witteveen covers Meta's open-weight Muse Glimmer 30B release, pointing to the research blog and the Hugging Face collection.

Video · Local & open models· ♥ 405

Sam Witteveen walks through Muse Glimmer 30B

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@cobusgreyling

@cobusgreyling

Cobus Greyling's companion repo for Muse Glimmer 30B pairs a long-form intro with an offline-first interactive lab for exploring agent loops, benchmarks and memory envelopes before downloading the weights.

GitHub · Local & open models

Muse Glimmer interactive local agent lab