QUASAR-trained Blackwell W4A4 build of Muse Glimmer 30B at 21.8 GiB, reporting 20% lower KL to BF16 than Red Hat's NVFP4 W4A4 checkpoint.

Resource · Local & open models· ♥ 3
4 builds · page 1 of 1
QUASAR-trained Blackwell W4A4 build of Muse Glimmer 30B at 21.8 GiB, reporting 20% lower KL to BF16 than Red Hat's NVFP4 W4A4 checkpoint.

Resource · Local & open models· ♥ 3
Your product
Sponsored
Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.
Shown every 12 builds · on every catalog page
NVIDIA's technical blog reports Muse Glimmer serving over 20K tokens/sec on a single Blackwell Ultra GPU and covers RTX 5090, DGX Spark, DGX Station and Jetson deployments.

Resource · Local & open models
A reproducible single-GPU deployment of Muse Glimmer 30B in BF16 with DFlash speculative decoding on a 96GB RTX PRO 6000 Blackwell, served via vLLM with pinned overlays and smoke tests.
GitHub · Local & open models
Muse Glimmer 30B feels significantly more precise and reliable, it almost never drops the ball or breaks rules. However, its designs lack creative depth and richness. Qwen3.6 35B, on the other hand, is prone to more occasional blunders/hallucinations, but its creative output is superior. It generates far richer, more complex voxel worlds and offers higher design quality. LLama.ccp Build Provenance: • Base: llama.cpp upstream (merge 4445f8d, build 661) • CUDA Toolkit 13.1 + MSVC 19.44 + sm_120a-real (native Blackwell PTX) • Flags: GGML_CUDA=ON, GGML_CUDA_FA=ON, GGML_CUDA_FA_ALL_QUANTS=ON, GGML_CUDA_GRAPHS=ON, GGML_NATIVE=OFF • License: MIT (upstream llama.cpp) Do you think Qwen3.6 is still the undisputed king here?
Reddit post · Benchmarks & research