shipwithmuse

Entries matching “nvfp4”

16 builds · page 1 of 1

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Q

QUASAR-QAT

QUASAR-QAT

QUASAR-trained Blackwell W4A4 build of Muse Glimmer 30B at 21.8 GiB, reporting 20% lower KL to BF16 than Red Hat's NVFP4 W4A4 checkpoint.

Resource · Local & open models· ♥ 3

QUASAR native NVFP4 Muse Glimmer

R

RedHatAI

RedHatAI

Red Hat AI's FP4 weight and activation quantized Muse Glimmer 30B.

Resource · Local & open models· ♥ 18

Red Hat AI NVFP4 Muse Glimmer

@r0b0tlab

@r0b0tlab

Reproducible native NVFP4 serving of Muse Glimmer 30B on one DGX Spark (GB10) with unmerged vLLM support: about 10.3 tok/s single-stream versus 4.2 for BF16, 52.5 tok/s at c16, and 131K context checked with needle-in-a-haystack tests.

GitHub · Local & open models

Muse Glimmer 30B NVFP4 on a single DGX Spark

@chishiki37

@chishiki37

Benchmark reports on Muse-Glimmer-30B on NVIDIA DGX Spark covering BF16 to Q4 to DFlash (a 10x speedup) and NVFP4 via SGLang, plus a head-to-head against Qwen3.6-27B.

GitHub · Benchmarks & research

Muse Glimmer optimization reports on DGX Spark

@ryangu00

@ryangu00

A measured deployment report of Muse Glimmer 30B NVFP4 with DFlash on a single Dell Pro Max GB10, where it posted top vision and SRE-ops scores but failed five deployment gates against DeepSeek V4 Flash.

GitHub · Local & open models

Muse Glimmer on Dell GB10: a negative result

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@High-Performance-AI-Lab

@High-Performance-AI-Lab

A 40-chapter book on Muser, an engine that runs Muse Glimmer on Apple Silicon Metal. It covers kquant and DFlash speculative lanes, exact KV-cache replay with kvpack, and GB10 NVFP4 prefill handed off to a Mac, with every number tied to an evidence receipt.

GitHub · Local & open models· ★ 3

The Muser book: how to write an inference engine

@vcruz305

@vcruz305

Tested SGLang and vLLM launch recipes for an NVFP4 export of a Hermes-agentic fine-tune of Muse-Glimmer-30B on NVIDIA DGX Spark (GB10).

GitHub · Local & open models

Hermes-agentic Glimmer NVFP4 on DGX Spark

@airawatraj

@airawatraj

Inference tuning notes for serving Muse Glimmer 30B NVFP4 with DFlash on a single NVIDIA DGX Spark as a consistent agent backend; the repo reports 27.5 tok/s average and 90/100 on its tool eval with 128K context.

GitHub · Local & open models

Muse Glimmer NVFP4 on DGX Spark

@MiaAI-Lab

@MiaAI-Lab

A one-script vLLM setup that serves the roughly 19 GB NVFP4 Muse Glimmer 30B with its vision encoder kept, DFlash speculative decoding using the official drafter head, and up to 256K context on GB10, RTX 5090 or RTX PRO 6000.

GitHub · Local & open models· ★ 8

Muse Glimmer 30B NVFP4 for DGX Spark and RTX 5090

@homerquan

@homerquan

A start/stop/status launcher that serves the NVFP4 Muse Glimmer 30B checkpoint on NVIDIA DGX Spark with vLLM, Glimmer's reasoning and tool parsers, and its DFlash speculative decoder.

GitHub · Local & open models· ★ 2

Muse Glimmer launcher for DGX Spark

SGLang

@sgl_project

SGLang is honored to provide day-0 support for @AIatMeta's Muse Glimmer. ~230 tok/s on a single RTX 5090 with NVFP4 + DFlash on, and it runs out of the box on @NVIDIAAI RTX PRO 6000, DGX Spark, and Apple Silicon via MLX. Huge thanks to the NVIDIA and Meta teams for the

X post · Local & open models· ♥ 95

SGLang day-0 serving for Glimmer

LMSYS Org

@lmsysorg

@AIatMeta's Muse Glimmer (30B dense, open-weights) launches with SGLang day-0 support. We got ~230 tok/s on a single RTX 5090, with NVFP4 + DFlash on. It also works out of the box on @NVIDIAAIDev RTX Pro 6000, DGX Spark, and MLX for Mac. Speed and reliability have always been

X post · Local & open models· ♥ 129

SGLang serves Glimmer at ~230 tok/s on one RTX 5090

@High-Performance-AI-Lab

@High-Performance-AI-Lab

muser is a standalone inference engine for Muse Glimmer 30B on Apple Silicon Metal, with an optional disaggregated lane where an NVIDIA GB10 node prefills in NVFP4 and hands the KV cache to the Mac.

GitHub · Local & open models

muser inference engine for Muse Glimmer