№ 0392Reddit post
Muse Glimmer at 1M context on 2x DGX Spark
Ran Muse Glimmer 30B on a two-node DGX Spark cluster and YaRN-stretched it to 1M context, with 3/3 needle retrieval at 832K tokens and 36–38 tok/s with DFlash on one Spark.
StartupTim used llama.cpp with the official K-Quant GGUF, mmproj and DFlash drafter. Decode went from ~10.5 tok/s to 36–38 tok/s with DFlash; RPC split across both Sparks was ~30% slower. It also passed 7/7 on a small execution-checked coding suite.
I ran Muse Glimmer @ 1M context - All tests passed.
Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Muse Glimmer 30B running the day after release — then pushed its context from the trained 131K all the way to 1M with YaRN, verifying retrieval at every rung. Sharing config + results since the "131,072+" hint in the model card turned out to be very real. Setup • Hardware: 2× NVIDIA DGX Spark (GB10, 128 GB unified each, ~273 GB/s), ConnectX-7 direct link between them • Engine: llama.cpp master (day-1 muse_glimmer support), built from source with CUDA sm_121 + GGML_RPC • Model: official Muse-Glimmer-30B-GGUF K-Quant-Dynamic (~18.3 GiB) + official mmproj (vision) + official DFlash drafter • Spec decode: --spec-type draft-dflash --spec-draft-n-max 15 (block-diffusion drafter) • Context extension: --rope-scaling yarn --rope-scale <2/4/8> --yarn-orig-ctx 131072 plus --override-kv muse-glimmer.context_length=int:<N> (llama.cpp caps at trained length otherwise) • Yes, we also ran it split across both Sparks with llama.cpp RPC — no reason beyond liking to cluster things for fun. Our daily dri



ChatForm
Tgmlabs