№ 0177GitHub
Muse Glimmer on Dell GB10: a negative result
A measured deployment report of Muse Glimmer 30B NVFP4 with DFlash on a single Dell Pro Max GB10, where it posted top vision and SRE-ops scores but failed five deployment gates against DeepSeek V4 Flash.
 # Muse Glimmer 30B in NVFP4 on one Dell Pro Max with GB10 — a negative deployment result > We deployed the open-weight member of the Muse family known to us at trial time, Muse Glimmer 30B (dense 29.6B including a ViT-G/14 1.8B vision tower), in NVFP4 W4A4 on a single Dell Pro Max with GB10, using the vLLM official recipe image with the DFlash speculative draft (`k=15`), and scored it against the DeepSeek V4 Flash Vision-Exp comparison baseline over our private 11-category eval bank (questions not published) with 2 runs per category, taking the median and adjudicating on the lower run whenever the spread exceeded 5. The static footprint is unusually comfortable (25.4 GB NVFP4 weights, a KV pool of 2,817,481 tokens at `max_model_len` 131,072), and the model produced the best scores of the whole comparison on vision (c6 90.0) and SRE-ops (c10 93.3) plus the best 6-stream aggregate (116.1 tok/s). But it failed five gates of the frozen rule set, and Δown sits in the tie zone, which is by itself insufficient to win. Verdict: negative result, production unchanged. ## Why this matters This is a measured **negative** result, not a recommendation. A model that posts the best vision and SRE-ops scores of a five-column comparison, a KV pool of 2.8M tokens on one node, and the highest 6-stream aggregate can still lose a deployment decision if it fails five independent gates — here tool use, long coding, agentic-IF, raw decode speed, and wall clock against the reference. The point of the cookbook is the evidence: every number in the trial's own results tables comes from the trial, with its condition, and the five gates that failed are listed gate by gate rather than summarized away. Community/inventory numbers carried into the plan (see "Community / inventory numbers") are third-party or pre-trial figures and are labelled as such, not trial measurements. The one residual recommendation is a narrow one — an offline batch vision / SRE-Q&A lane — and it is stated as such, not promoted to a mainline role. ## Hardware and stack **Node / hardware** | Item | Value | |---|---| | Node class | Dell Pro Max with GB10, single node | | Memory / OS / driver / CUDA / container runtime | not recorded in the sources | | Weights format | NVFP4 W4A4, `Inferact/Muse-Glimmer-30B-NVFP4-W4A4`, **25.4 GB** (weights file size); revision / checksum / pull



ChatForm
Tgmlabs