shipwithmuse

№ 0925Reddit post

4x V100 llama.cpp fork speeds up Muse Glimmer 30B

Built a tensor-parallel llama.cpp fork for 4x NVLinked V100s; Muse Glimmer 30B F16 goes from 15.1 to 40.4 tok/s generation (+168%) and 1,048 to 2,396 tok/s prefill.

r/v100· u/syscomauView on Reddit ↗

v100 Llama.cpp fork - Looking for testers

Hey Guys, I've got 4 x v100's in a Dell C4140 (NVlink) and I have been working on a fork of llama.cpp that is targeted at the v100's. Looking for testers to give it a go and provide feedback. WyvernTKC/llama.cpp-4xV100: Fork of llama.cpp Nvida Volta V100 (tensor parallelism 4 x v100 GPU) model arch type size (GB) pp layer pp tensor change tg layer tg tensor change glm4 9B Q8_0 glm4 dense 9.3 1187.9 2924.3 +146 % 67.6 123.0 +82 % qwen35 27B Q8_K_P qwen35 dense 29.3 640.4 1717.6 +168 % 22.2 52.4 +137 % gemma4 31B Q8_0 gemma4 dense 30.4 679.8 1621.2 ±322 noisy 20.6 46.7 +127 % muse-glimmer 30B F16 muse-glimmer dense 51.9 1048.4 2395.6 +128 % 15.1 40.4 +168 % llama 70B Q8_0 llama dense 69.8 302.5 950.5 +214 % 9.8 28.6 +192 % qwen35moe 35B-A3B Q8_0 qwen35moe MoE 256×8 34.4 1602.6 3143.7 +96 % 93.6 113.7 +21 % qwen3next 80B-A3B Q4_K_M qwen3next MoE 512×10 45.9 889.8 1647.0 +85 % 76.6 86.8 +13 % deepseek4 284B Q2_K deepseek4 MoE 256×6 90.9 188.8 616.1 +226 % 27.3 37.7 +38 % Thanks!

Also filed under Local & open models

See all →
  1. 0805

    Muse drives a rover with a custom connector★

    Took your custom-connector idea to the physical world: the "service with an API" was my robot. Muse wrote the connector for my rover's API, then installed PyTorch & Depth Anything V2 in the VM because the camera is 2D, drove to the black ball and stopped a few inches short.

    @hrhraj

    X post

    Local & open models

  2. 0490

    Muse Glimmer on one Arc Pro B70★

    A vLLM-XPU and DFlash recipe for Muse Glimmer 30B on a single Intel Arc Pro B70, reporting 278 aggregate tok/s across eight clients and an 840.8 tok/s burst peak at concurrency 96.

    @mgaruccio

    GitHub

    Local & open models

  3. 0458

    Muse Glimmer 30B stretched to 512K context★

    My fun weekend project was to try to make the new Muse Glimmer 30B work with a longer context, deciding to go for 512k first. I had expected the usual YaRN shenanigans and maybe a LoRA. I couldn't have been wrong more. Upon closer look, Glimmer turned out to be rather unusual architecturally. The thing that make long-context adaptations painful in other models, full attention layers with token position encoding, it simply not there. Instead, only 2048 tokens-wide SWA layers have RoPE, and full GQA attention layers have no position encoding at all. It appears the model is trained to work with long-distance token relationships inferred from the context and SWA layers. It's a rather bold architecture bet, but it seems Meta managed to pull it off. As a result, the model architecture appears to be uniquely suited for context extension by simple mechanical means. To change model context length from stock 128k to, say, 512k, you need only to change “max_position_embeddings” config setting from 131072 to 524288. What confuses other models, like Qwen3.5 family, Glimmer just takes into its stride. I spent close to 70h of compute on DGX Spark to test stock model with extended context on a

    u/mr_il

    Reddit post

    Local & open models

  4. 0435

    Local credit card statement analysis with Glimmer★

    Using @AIatMeta's Muse Glimmer all locally to process personal monthly credit card statements. Your data belongs to you! Try different agent tasks using your favorite apps / harnesses with Ollama.

    @ollama

    X post

    Local & open models

More Reddit posts

See all →
  1. 1071

    Public brokerage connector for Meta Muse

    Public published a connector template so users can link a Public account to Meta Muse and research markets, analyze a portfolio and prepare trades from the conversation.

    u/Public

    Reddit post

    Connectors & MCP

  2. 1067

    Muse Voice Transcribe tested on clinical diarization

    Compared diarization models on 15 mock doctor-patient consultations (~2.4 h): Meta Muse Voice Transcribe scored 13.04% DER at ~92 s per request via API, behind Pyannote (2.89%) and Nemotron 3 (4.80%).

    u/MajesticAd2862

    Reddit post

    Benchmarks & research

  3. 1065

    Muse checks a training run and locks the Mac from a phone

    Built a small bridge so Muse can act on my Mac from my phone, and recorded a real session: it looks at the Terminal, reports the epoch, loss and accuracy it sees, then locks the machine when asked. What struck me building it is how much of the work is permissions, not intelligence: per-action consent, small window captures instead of a live feed, rejecting stale observations before any input. Free beta if anyone wants to try it. Developer here, ask away. try wand here today

    u/OldChemical3853

    Reddit post

    Agents & automation

  4. 1063

    Iggy, a Muse agent, reports on a day posting to Reddit

    Iggy, a Muse agent, spent a day posting in agent subreddits as a disclosed AI; posts in r/agenticAI and r/AI_Agents were removed by new-account filters, and it found threads focused on scaffolding, not the model.

    u/MuseIggy

    Reddit post

    Agents & automation

Curator picks

  1. 1046

    Medical bills audited line by line, $4,000 saved★

    Got Muse logged in to my medical provider’s portal, he pulled the itemized bills, and questioned every line. So far he’s found several times I’d been double billed, asked for some discounts and has saved me over $4,000. If your moat is bureaucracy, you’re cooked.

    @Ryan_Holdaway

    X post

    Errands & personal agent

  2. 1013

    Shop Pay agentic checkout on every Shopify store★

    We are excited to announce we are partnering deeply with Muse to enable agentic checkout with Shop Pay on all Shopify stores, offering people an easy and delightful way to shop and check out with Muse.

    @tobi

    X post

    Business & commerce

  3. 1009

    Private e-book library app from Google Drive★

    Muse built me a private library for the e-books and articles in my Google Drive. Everything is organized by topic, and each section opens onto its own subcategorized shelves. Each book opens like a real book and is readable in-app

    @chiasmus_cap

    X post

    Apps & websites

  4. 1007

    Plumbing company run by a Muse agent★

    I still can’t believe I can run my plumbing company with an agent so easily. I send this message to my Muse agent while in bed at 6am. And it: updates my job board, texts customer, updates office manager who arrives at 8am in slack Notifies technician

    @HouseHackerJon

    X post

    Agents & automation