shipwithmuse

Entries matching “glimmer”

80 builds · page 1 of 1

ollama

@ollama

Using @AIatMeta's Muse Glimmer all locally to process personal monthly credit card statements. Your data belongs to you! Try different agent tasks using your favorite apps / harnesses with Ollama.

X post · Local & open models★ Pick· ♥ 345

Local credit card statement analysis with Glimmer

U

xenovatech

u/xenovatech

Ran Muse Glimmer 30B locally in the browser with custom WebGPU kernels at ~25 tok/s on an M4 Max, matching llama.cpp speed.

Reddit post · Local & open models★ Pick

Muse Glimmer 30B in the browser via WebGPU

Paolo Rosson

@redp314

Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Ollama's MLX engine (DFlash included) at ~29 tok/s. Tuned llama.cpp: ~21. Raw mlx-vlm: ~10, not optimized yet. Numbers below if you're

X post · Local & open models★ Pick· ♥ 44

Glimmer serving shootout on an M3 Max

Unsloth AI

@UnslothAI

2-bit Muse Glimmer GGUF managed to call 100+ tools on just 14GB RAM. 🔥 Muse Glimmer did a complete repo bug hunt for 5 mins nonstop with: evidence, repro, fix, tests and a PR writeup. Run and train it in Unsloth. GitHub repo: github.com/unslothai/unsl…

X post · Local & open models★ Pick· ♥ 1.6K

2-bit Glimmer GGUF: 100+ tool calls on 14GB RAM

P

pytorch.org

pytorch.org

PyTorch added end-to-end Muse Glimmer support to ExecuTorch; on an M5 Pro, DFlash speculative decoding lifts image+text decode from 21.6 to 33.0 tok/s, and it powers the Pi coding agent locally.

Resource · Local & open models★ Pick

Muse Glimmer on ExecuTorch: DFlash on Macs and NVIDIA GPUs

U

PyaesoneP

u/PyaesoneP

I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8\_O KV cache. It's a joy to use a dense 30B model at this size and still get \~30 tok/s on a VRAM-constrained laptop. It's supposed to be only slightly worse than the official 17GB K-quant at a much smaller footprint, and for my Hermes Agent use case I don't notice a quality difference. It's just much faster. I've tried Qwen 3.8 27B at SC2.20bpw H3 too. Definitely usable but I'm sticking with Unsloth UD\_Q4\_K\_XL for Qwen 3.8 27B because it's mainly for coding.

Reddit post · Local & open models★ Pick

Muse Glimmer 30B on a 12GB laptop GPU

U

A-Rahim

u/A-Rahim

Been tinkering with speculative decoding on Apple Silicon for a while, and this week I got Meta's new Muse Glimmer 30B working in my project mlx-dspark. On my M4 Pro, the 8-bit model goes from 8.2 tok/s to 18-26 tok/s depending on content. Math is the best case at 3.27x, code 2.5x, chat 2.22x. Output is byte-identical to normal decoding since the target verifies every token, so there's no quality tradeoff; it's just faster. Meta's own DFlash numbers on Mac are 1.5x (M4 Max) / 1.8x (M5 Max), but those are on the 4-bit build, so not really apples-to-apples. 4-bit for me is ~1.7x at ~25 tok/s and only needs ~18GB. The 8-bit run peaks around 40GB, so you want a 48GB Mac for it. Basically, you get 8-bit quality at 4-bit speed. Repo: github.com/ARahim3/mlx-dspark I'm happy to hear feedback, and I'm curious about what other M-series chips get.

Reddit post · Local & open models★ Pick

Muse Glimmer 3.3x faster on Mac with mlx-dspark

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

I

immanuelpeter

immanuelpeter

Standalone packaging of the vision tower and projector extracted from Muse Glimmer 30B.

Resource · Local & open models· ♥ 2

Muse Glimmer Vision tower

R

RedHatAI

RedHatAI

Red Hat AI's FP4 weight and activation quantized Muse Glimmer 30B.

Resource · Local & open models· ♥ 18

Red Hat AI NVFP4 Muse Glimmer

@nicedreamzapp

@nicedreamzapp

An architecture port adding the muse_glimmer model class (vision tower, language model, projector and image processor) to mlx-vlm, so any Muse Glimmer checkpoint runs multimodally on Apple Silicon.

GitHub · Local & open models

mlx-vlm support for Muse Glimmer

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@cneuralnetwork

@cneuralnetwork

smol-muse-glimmer scales Muse Glimmer's language backbone down to a 51M-parameter model and trains it on TinyStories, reaching validation cross-entropy of 1.8127 at step 5,000.

GitHub · Benchmarks & research

smol-muse: 51M Muse Glimmer architecture study

U

unsloth

unsloth

Unsloth's bitsandbytes 4-bit build of Muse Glimmer 30B for fine-tuning and inference.

Resource · Local & open models· ♥ 14

Unsloth bnb 4-bit Muse Glimmer

O

OpenVINO

OpenVINO

Experimental OpenVINO INT4 conversion of Muse Glimmer 30B that needs development builds of Optimum Intel and OpenVINO.

Resource · Local & open models· ♥ 1

OpenVINO INT4 Muse Glimmer (experimental)

0

0bserverx

0bserverx

A suite of 27 GGUF quantizations of an abliterated Muse Glimmer 30B produced with Heretic v1.4.0.

Resource · Local & open models· ♥ 20

Heretic Uncensored Muse Glimmer GGUF

P

PursuitOfDataScience

PursuitOfDataScience

LoRA adapter that makes Muse Glimmer 30B reliably commit to tool calls when it already knows the correct function.

Resource · Local & open models· ♥ 1

Muse Glimmer tool-calling LoRA

Y

yogeshjog

yogeshjog

LoRA adapter for Muse Glimmer 30B that makes responses predictable machine-readable JSON with a stable API-style envelope.

Resource · Local & open models· ♥ 1

Muse Glimmer JSON API LoRA

U

PandaBearFred

u/PandaBearFred

Same old prompt, just appended a TIP in the end: "Write a single HTML file with a full-page canvas and no libraries. Simulate a realistic side-view of a moving car as the main subject. Keep the car visible in the foreground while the background landscape scrolls continuously to create the feeling that the car is driving forward. Use layered scenery for depth: nearby ground, roadside elements, trees, poles, and distant hills or mountains should move at different speeds for a natural parallax effect. Animate the wheels spinning realistically and add subtle body motion so the car feels connected to the road. Let the environment pass smoothly behind it, with repeating but varied scenery that makes the movement feel believable. Use cinematic lighting and a cohesive sky, such as sunset, dusk, or daylight, to enhance atmosphere. The overall motion should feel calm, immersive, and realistic, with a seamless looping animation. TIPS: You don't have vision abilities so don't try it yourself. If you feel in need of vision ability, you can access http://xxx:8080/v1, model id: Muse-Glimmer for help, it will see the picture, and describe it for you." Then the PI agent started spinning, round a

Reddit post · Agents & automation

DeepSeek agent borrowing Glimmer's eyes

R

RedHatAI

RedHatAI

Red Hat AI's FP8 weight and activation quantized Muse Glimmer 30B (text and image input) for vLLM.

Resource · Local & open models· ♥ 12

Red Hat AI FP8-block Muse Glimmer

V

vmlinux

vmlinux

ROCmFP4 and ROCmFP8 builds of Muse Glimmer 30B and its drafter, targeted and tested on AMD Strix Halo (gfx1151).

Resource · Local & open models· ♥ 18

Muse Glimmer ROCmFPX GGUF for Strix Halo

M

merve

merve

QLoRA adapter that teaches Muse Glimmer 30B to return click points on web screenshots from instructions, trained on the MolmoWeb dataset.

Resource · Local & open models· ♥ 6

Muse Glimmer QLoRA click grounding

I

incoai

incoai

GGUF conversions of Inco AI's DFlash 2 draft model for Muse Glimmer 30B, for speculative decoding in llama.cpp.

Resource · Local & open models· ♥ 16

DFlash2 drafter GGUF for Muse Glimmer

N

NANI-Nithin

NANI-Nithin

GGUF builds of Muse Glimmer 30B all cut from the same BF16 source weights, including full-precision files.

Resource · Local & open models· ♥ 6

Muse Glimmer 30B llama.cpp GGUF quants

F

fireworks.ai

fireworks.ai

Fireworks made Muse Glimmer 30B available on launch day, pitching high-concurrency, cost-effective serving for always-on agents.

Resource · Local & open models

Muse Glimmer 30B on Fireworks AI

K

kaitchup

kaitchup

Muse Glimmer 30B quantized with Intel AutoRound at a 3.5-bit target and packed with llm-compressor, tested on vLLM.

Resource · Local & open models· ♥ 3

AutoRound 3.5-bit Muse Glimmer

A

AaryanK

AaryanK

Solo-built GGUF line of Muse Glimmer 30B with custom calibration, per-tensor allocations and an eval harness behind every reported number.

Resource · Local & open models· ♥ 21

Muse Glimmer GGUF (AK line)

M

mlx-community

mlx-community

mlx-community's 4-bit MLX conversion of Muse Glimmer 30B made with mlx-vlm 0.6.12 for Apple Silicon.

Resource · Local & open models· ♥ 18

MLX 4-bit Muse Glimmer

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

A

armand0e

armand0e

Muse Glimmer 30B fine-tuned on agentic coding traces and chat distilled from Claude Fable 5, merged bf16 and drop-in for the base.

Resource · Local & open models· ♥ 3

Muse Glimmer Fable Distill

LM Studio

@lmstudio

Muse Glimmer 30B is live in LM Studio! It's a new open source model from Meta. Apache 2.0 license, fit right on your laptop. It is the strongest model of its size class we've tested.

X post · Local & open models· ♥ 1.4K

Muse Glimmer 30B in LM Studio

K

kingjones777

kingjones777

ROCmFP4 GGUF of Muse Glimmer 30B with DFlash for AMD Strix Halo, requiring a ROCmFPX llama.cpp fork that adds the muse-glimmer architecture.

Resource · Local & open models· ♥ 4

ROCmFP4 Strix Halo DFlash GGUF

Unsloth AI

@UnslothAI

Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM. Muse Glimmer is Apache 2.0 licensed, supports vision and is the strongest agentic model for its size. Run and train the model via Unsloth. GGUF: huggingface.co/unsloth/Muse-G… Guide: unsloth.ai/docs/models/mu…

X post · Local & open models· ♥ 2.8K

Unsloth GGUFs and guide for Muse Glimmer

H

holaclaw.ai

holaclaw.ai

HolaClaw's tutorial for running Muse Glimmer 30B behind OpenClaw on a Mac, with hardware requirements and llama.cpp and Ollama setup.

Resource · Local & open models

Run OpenClaw with Muse Glimmer locally

V

vcruz305

vcruz305

Fine-tune of Muse Glimmer 30B for Hermes Agent and agentic tool work that teaches the model to call one or two tools and stop.

Resource · Local & open models· ♥ 2

Muse Glimmer Hermes-Agentic

B

Blackfrost-AI

Blackfrost-AI

Abliterated Muse Glimmer 30B GGUF quant ladder, updated with a 1.63 GB abliterated DFlash drafter and a 1.40 GB multimodal projector.

Resource · Local & open models· ♥ 49

Abliterated Muse Glimmer GGUF ladder

vLLM

@vllm_project

@Meta is back in open source. Excited to announce Day-0 vLLM support for Muse Glimmer 30B, the first open-weights model from Meta Superintelligence Labs — which ships under Apache 2.0!!! 30B dense, 128K+ context, multimodal, built for local agents. Capable enough for

X post · Local & open models· ♥ 251

Day-0 vLLM support for Muse Glimmer

@TrevorS

@TrevorS

Refusal-direction ablation on Muse-Glimmer-30B that cut refusals from 128/150 to 3/150, adding an agentic-safety evaluation and publishing bf16 and GGUF uncensored weights.

GitHub · Local & open models★ Pick· ★ 1

Muse Glimmer abliteration

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

H

huggingface.co

huggingface.co

Hugging Face's launch post covers day-0 transformers, llama.cpp and vLLM support, Inference Endpoints, speculative decoding, TRL fine-tuning and agent demos for Muse Glimmer.

Resource · Local & open models★ Pick

Hugging Face: Muse Glimmer is local, agentic and open

@mgaruccio

@mgaruccio

A vLLM-XPU and DFlash recipe for Muse Glimmer 30B on a single Intel Arc Pro B70, reporting 278 aggregate tok/s across eight clients and an 840.8 tok/s burst peak at concurrency 96.

GitHub · Local & open models★ Pick· ★ 1

Muse Glimmer on one Arc Pro B70

@SamuelAlexander

@SamuelAlexander

Samuel Alexander ran Muse Glimmer 30B entirely on a Qualcomm Dragonwing IQ-9075 board for zero-shot PCB defect inspection and tool calling, measuring 21.6 GB resident with full 131K context and 2.84 tokens/s generation.

GitHub · Local & open models★ Pick

Muse Glimmer 30B on a Qualcomm Dragonwing board

U

mozilla-ai

u/mozilla-ai

We've been curious how far local models have actually come for agentic coding tasks, so we ran an experiment. Setup: • Model: Muse Glimmer (30B), packaged as a single llamafile • Agent: Hermes coding agent (connected via llamafile's local server mode, zero API keys needed) • Target: Mozilla AI's Otari gateway The Issue: We pointed Hermes at a real, reported bug in Otari (#183) where the gateway returned a vague 502 error on image requests instead of passing through the actual provider error. What the Agent Did: Hermes read the issue, navigated the repo, isolated the bug, created a branch, ran existing tests, wrote a new regression test, and opened a draft PR (#727). All of it ran locally and offline, with zero code written by hand. It's still draft PR territory rather than a merged fix, but it's a solid signal that ~30B local models are getting genuinely capable for real dev workflows, not just toy demos. Video walkthrough of the run: https://youtu.be/5GAgbT-XgHU?si=vJqEDGm9hssCO5-M Happy to answer questions about the setup, model performance, or how Hermes handled tool calling!

Reddit post · Local & open models★ Pick

Local Muse Glimmer agent opens a real pull request

U

mr_il

u/mr_il

My fun weekend project was to try to make the new Muse Glimmer 30B work with a longer context, deciding to go for 512k first. I had expected the usual YaRN shenanigans and maybe a LoRA. I couldn't have been wrong more. Upon closer look, Glimmer turned out to be rather unusual architecturally. The thing that make long-context adaptations painful in other models, full attention layers with token position encoding, it simply not there. Instead, only 2048 tokens-wide SWA layers have RoPE, and full GQA attention layers have no position encoding at all. It appears the model is trained to work with long-distance token relationships inferred from the context and SWA layers. It's a rather bold architecture bet, but it seems Meta managed to pull it off. As a result, the model architecture appears to be uniquely suited for context extension by simple mechanical means. To change model context length from stock 128k to, say, 512k, you need only to change “max_position_embeddings” config setting from 131072 to 524288. What confuses other models, like Qwen3.5 family, Glimmer just takes into its stride. I spent close to 70h of compute on DGX Spark to test stock model with extended context on a

Reddit post · Local & open models★ Pick

Muse Glimmer 30B stretched to 512K context

@schererstefan

@schererstefan

glimmer-cli is a local TypeScript CLI for Muse Glimmer 30B and Muse Spark 1.2 via Ollama, with stubbed tools and a reproducible tool-use eval harness.

GitHub · Local & open models

glimmer-cli for Muse Glimmer and Spark

V

ValiantLabs

ValiantLabs

Valiant Labs' Esper 4 agentic coding fine-tune of Muse Glimmer 30B.

Resource · Local & open models· ♥ 6

Muse Glimmer Esper 4

@homerquan

@homerquan

A start/stop/status launcher that serves the NVFP4 Muse Glimmer 30B checkpoint on NVIDIA DGX Spark with vLLM, Glimmer's reasoning and tool parsers, and its DFlash speculative decoder.

GitHub · Local & open models· ★ 2

Muse Glimmer launcher for DGX Spark

M

meta-models

meta-models

Meta's own GGUF conversions of Muse Glimmer 30B for llama.cpp.

Resource · Local & open models· ♥ 343

Official Muse Glimmer 30B GGUF

D

darkc0de

darkc0de

Decensored Muse Glimmer 30B made with Heretic v1.4.0, shipped with a reproduce directory.

Resource · Local & open models· ♥ 15

Muse Glimmer 30B heretic

O

ollama.com

ollama.com

Ollama shipped Muse Glimmer on day one: `ollama run muse-glimmer`, plus a muse-glimmer:30b-mlx tag for Apple Silicon that Ollama says runs 1.5–1.8x faster with DFlash.

Site · Local & open models

Muse Glimmer in the Ollama library

@PipeNetwork

@PipeNetwork

muse-glimmer-mlx is an MLX port of Muse Glimmer 30B for Apple Silicon that supplies the missing runtime so the many unloadable MLX conversions published on Hugging Face can actually be run.

GitHub · Local & open models

MLX runtime for Muse Glimmer 30B

U

myanimal22

u/myanimal22

Muse Glimmer 30B feels significantly more precise and reliable, it almost never drops the ball or breaks rules. However, its designs lack creative depth and richness. Qwen3.6 35B, on the other hand, is prone to more occasional blunders/hallucinations, but its creative output is superior. It generates far richer, more complex voxel worlds and offers higher design quality. LLama.ccp Build Provenance: • Base: llama.cpp upstream (merge 4445f8d, build 661) • CUDA Toolkit 13.1 + MSVC 19.44 + sm_120a-real (native Blackwell PTX) • Flags: GGML_CUDA=ON, GGML_CUDA_FA=ON, GGML_CUDA_FA_ALL_QUANTS=ON, GGML_CUDA_GRAPHS=ON, GGML_NATIVE=OFF • License: MIT (upstream llama.cpp) Do you think Qwen3.6 is still the undisputed king here?

Reddit post · Benchmarks & research

Voxel worlds: Glimmer 30B vs Qwen3.6 35B

M

MuXodious

MuXodious

Muse Glimmer 30B abliterated with Heretic v1.4.0 using self-organizing maps and magnitude-preserving orthogonal ablation.

Resource · Local & open models· ♥ 11

Muse Glimmer SOMPOA heresy

B

bartowski

bartowski

llama.cpp imatrix quantizations of Muse Glimmer 30B with image support via mmproj and MTP/DFlash notes.

Resource · Local & open models· ♥ 19

bartowski imatrix GGUF of Muse Glimmer

C

cyankiwi

cyankiwi

AWQ INT4 quant of Muse Glimmer 30B calibrated on STEM and agentic data across ten languages, 24.02 GB.

Resource · Local & open models· ♥ 10

AWQ INT4 Muse Glimmer

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@HawgAuto

@HawgAuto

Glimmer HD Vision is an OpenAI-compatible proxy that keeps images within Muse Glimmer 30B's 4,096 visual-token limit by sending a 4K image as one overview plus four overlapping detail tiles, with an OCR/layout mode.

GitHub · Local & open models

Glimmer HD Vision OCR proxy

M

meta-models

meta-models

Pre-exported ExecuTorch PTE artifacts of Muse Glimmer 30B from Meta, lowered and optimized for specific target backends.

Resource · Local & open models· ♥ 34

Muse Glimmer 30B ExecuTorch PTEs

Together AI

@togethercompute

Muse Glimmer is now live on Together AI. We’re proud to be a Day 0 launch partner for this open-weight model from Meta Superintelligence Labs, built for long-running agents that can reason, use tools, recover, and keep working across complex tasks.

X post · Local & open models· ♥ 35

Muse Glimmer on Together AI

Y

huggingface.co

huggingface.co

A prebuilt macOS arm64 bundle for the Muse Glimmer voice-agent recipe in meta-oss-cookbook: Parakeet speech helper, Muse Glimmer worker and Supertonic TTS executables built from one pinned ExecuTorch checkout, plus the shared MLX Metal library.

Resource · Local & open models

Muse Glimmer voice agent ExecuTorch runtime

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

U

Longjumping-Elk-7756

u/Longjumping-Elk-7756

Glimmer obtient 92 % du score d'intelligence de Qwen3.6 (35/38), mais Qwen a généré environ 2,9× plus de tokens sur l'ensemble de l'Intelligence Index. Et sur les endpoints mesurés par Artificial Analysis, Glimmer génère environ 1,8× plus vite. Et le context de glimmer et bien plus efficace ! C est une belle avancer architecture tout de meme , je pense que si il sorte une version 1.1 (surtout pour améliorer terminal benchmark ) ont pourrai être très surpris !

Reddit post · Benchmarks & research

Glimmer hits 92% of Qwen3.6 with 2.9x fewer tokens

@rickyzzzzz

@rickyzzzzz

A controlled local benchmark on an M1 Max comparing Muse Glimmer 30B with Qwen 3.6 35B and Qwen 3.8 27B on tool calling and data-science tasks; Glimmer passed 24/30 versus Qwen 3.8's 30/30.

GitHub · Benchmarks & research

Muse Glimmer vs Qwen local agent benchmark

T

turboderp

turboderp

turboderp's self-calibrated EXL3 quants of Muse Glimmer 30B down to 1.75 bits per weight, with calibration and eval traces.

Resource · Local & open models· ♥ 14

EXL3 quants of Muse Glimmer

Sumanth

@Sumanth_077

Run and fine-tune Meta's Muse Glimmer locally! Meta released Muse Glimmer, a 30B dense vision model designed for local agentic and coding workflows. The first open model from Meta Superintelligence Labs, released under Apache 2.0. The model runs locally at different memory

X post · Local & open models· ♥ 26

Run and fine-tune Muse Glimmer locally

S

sequelbox

sequelbox

sequelbox's Tachibana-Agent fine-tune of Muse Glimmer 30B, part of a series also released for Gemma 4 12B and Qwen3.6 27B.

Resource · Local & open models· ♥ 4

Muse Glimmer Tachibana-Agent

Simon Willison

@simonw

Muse Glimmer, the new 30B model, is available on Hugging Face right now - here's the GGUF version: huggingface.co/meta-models/Mu…

X post · Local & open models· ♥ 210

Muse Glimmer GGUF on Hugging Face

Ben Burtenshaw

@ben_burtenshaw

Meta is back with Muse Glimmer: a 30B open-source multimodal model built for local, agentic use. HF is shipping day-0 support and I built a few demos to see what it can do. First: we gave Glimmer tools and asked it to quantize itself.

X post · Local & open models· ♥ 125

Muse Glimmer quantizes itself

S

@samwitteveenai

@samwitteveenai

Sam Witteveen covers Meta's open-weight Muse Glimmer 30B release, pointing to the research blog and the Hugging Face collection.

Video · Local & open models· ♥ 405

Sam Witteveen walks through Muse Glimmer 30B

Mark Zuckerberg

@finkd

Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats

X post · Local & open models· ♥ 30.4K

Zuckerberg opens Muse Glimmer weights

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

M

meta-models

meta-models

Meta's lightweight DFlash block-diffusion drafter for Muse Glimmer 30B that predicts blocks of 16 tokens per forward pass for speculative decoding.

Resource · Local & open models· ♥ 61

Muse Glimmer DFlash drafter (official)

Xuan-Son Nguyen

@ngxson

We are happy to announce that Muse Glimmer is day-0 supported on llama.cpp. Meta also provides an official GGUF quant:

X post · Local & open models· ♥ 86

Day-0 llama.cpp support for Muse Glimmer