shipwithmuse

Entries matching “context”

46 builds · page 1 of 1

U

mr_il

u/mr_il

My fun weekend project was to try to make the new Muse Glimmer 30B work with a longer context, deciding to go for 512k first. I had expected the usual YaRN shenanigans and maybe a LoRA. I couldn't have been wrong more. Upon closer look, Glimmer turned out to be rather unusual architecturally. The thing that make long-context adaptations painful in other models, full attention layers with token position encoding, it simply not there. Instead, only 2048 tokens-wide SWA layers have RoPE, and full GQA attention layers have no position encoding at all. It appears the model is trained to work with long-distance token relationships inferred from the context and SWA layers. It's a rather bold architecture bet, but it seems Meta managed to pull it off. As a result, the model architecture appears to be uniquely suited for context extension by simple mechanical means. To change model context length from stock 128k to, say, 512k, you need only to change “max_position_embeddings” config setting from 131072 to 524288. What confuses other models, like Qwen3.5 family, Glimmer just takes into its stride. I spent close to 70h of compute on DGX Spark to test stock model with extended context on a

Reddit post · Local & open models★ Pick

Muse Glimmer 30B stretched to 512K context

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

D

contextbolt.com

contextbolt.com

ContextBolt's founder shows how to connect the hosted ContextBolt SEO MCP server to Muse, either from the connector directory or as a custom connector, for keyword research, live SERPs and site audits.

U

nullc

u/nullc

I noticed on the same hardware that I can get 24 x 128k contexts with muse glimmer (30b q8_0 + mmproj+dflash) only gets me 3x 256k or 6x 128k with qwen. But a straight forward analysis of the architecture suggests to me that qwen's state per token is somewhat smaller than glimmers. So it seems llama.cpp is particularly memory inefficient for the qwen arch. I presume there is an existing issue for this, but I couldn't find one. What's the deal? The extra concurrency makes a big difference in batched performance.

Reddit post · Local & open models

24 parallel 128K contexts with Muse Glimmer

Alok

@analogalok

The "I don't have enough VRAM" excuse just died. I’m running Meta’s new 30B Muse Glimmer Q6_K_XL with a massive 130k context window on just 26GB VRAM FREE compute on Kaggle. Kaggle provides you free 2x Nvidia T4 GPUs. 30 hours usage each week! Yesterday, I showed you the

X post · Local & open models· ♥ 170

Muse Glimmer with 130k context on free Kaggle T4s

L

huggingface.co

huggingface.co

A 4-byte GGUF header edit that raises Muse Glimmer 30B's 131,072 context limit to 262K. The author shows YaRN scaling only adds prefill cost on this architecture, and verified 262K with 8 slots on one RTX 4090.

Resource · Local & open models· ♥ 1

Muse Glimmer 30B long-context GGUF patch

Mark Zuckerberg

@finkd

Opening access for developers to build Muse connectors. You bring the API -- Muse brings the agent, the browser, and the context of what the person actually wants. People reach your service just by asking for it, and their agent takes it from there. New connectors are live

X post · Connectors & MCP· ♥ 11.7K

Zuckerberg opens Muse connector platform

E

explainx.ai

explainx.ai

explainx.ai breaks down Meta's Sep 19 opening of the Muse connector platform: developers bring the API, Muse brings the agent, browser and user context.

Resource · Connectors & MCP

Building on the Muse connector platform

Alok

@analogalok

Muse Glimmer, A 30B parameter dense model swallowing a 130,000 token context window using only 19.3 GB of VRAM (extreme efficiency). No KV cache quantization required. I just benched the new Muse Glimmer 30B (dense) on a single RTX 4090. We are pulling 3,100+ t/s prefill and 75

X post · Local & open models· ♥ 392

Glimmer bench on a single RTX 4090

S

@STARTUPHAKK_sh

@STARTUPHAKK_sh

STARTUP HAKK tests Muse Spark 1.3 and argues the harness, tools and context around a model matter as much as the model, pitching their OpenMonoAgent harness.

Video · Coding & dev tools· ♥ 81

Muse Spark 1.3 tested by the OpenMonoAgent team

@axlecoffee

@axlecoffee

A BYOK VS Code extension that registers the InferHub catalog, including Muse Spark with its 1,048,576-token context window, as a model provider in the Copilot Chat picker without Copilot Pro.

GitHub · Coding & dev tools

InferHub models in GitHub Copilot Chat

U

NicolaZanarini533

u/NicolaZanarini533

I've had Qwen3.6:27b (and Qwen 3 coder next before it) running along side gpt-oss:20b for a while now as my two main models (qwen for coding, gpt-oss for agentic stuff). Qwen is pretty self-explanatory, while I had been using gpt-oss because of how good it was at producing json and instruction following, despite the size and age. https://preview.redd.it/lr61tb02lejh1.png?width=1920&format=png&auto=webp&s=cef1f0bbfb43c8462162ec675d52c5896d174118 The "upgrade" to 3.8 is pretty evident, especially because of the SWE bench score improvement, but I was more reticent with Muse-Glimmer as I had some trouble when I tried Gemma4, which was far too opinionated when given a task, but Muse-Glimmer seems great - low memory footprint at 128K context, fairly fast and seems to follow instructions well. What models are you using locally and for what? did you have a similar experience with these latest models?

Reddit post · Local & open models

Muse Glimmer as a resident local agent model

U

smallthings17

u/smallthings17

✨ Three New Models, Sharper Lore Awareness & Smoother Storytelling Across Complex Worlds This update introduces three new storytelling models, improves how Lore is recognized and prioritized during a story, and brings another round of reliability improvements to the model experience. 🎭 Three New Voices 💎 Gemini 3.8 Flash — Pro+ Google’s newest fast storyteller, with polished prose, responsive pacing, and a steady grip on complex scenes. Context support: 🔹 16K on Pro 🔹 48K on Ultra 🔹 80K on Legendary Context limits subject to change. ✨ Muse Spark 1.3 — Plus+ Vivid character interplay, strong world-state awareness, and deliberate continuity as stories evolve. Context support: 🔹 16K on Plus 🔹 32K on Pro 🔹 64K on Ultra 🔹 100K on Legendary 🌐 Hunyuan 4 Preview — Pro+ Built for ambitious living worlds, shifting relationships, and large casts with lasting consequences. Context support: 🔹 16K on Pro 🔹 32K on Ultra 🔹 48K on Legendary 📚 Sharper Lore Awareness Lore is getting better at recognizing what matters in the current scene. 🏷️ Smarter Lore activation — Lore cards can now activate when their title or character name appears naturally in the story, even

U

myanimal22

u/myanimal22

Some people told me that the difference in richness and layout between Glimmer and Qwen wasn't clear to them. This example makes it super clear. I'm aware that comparing Glimmer 30B (a dense model) with Qwen 3.6 (a MoE) isn't entirely fair, but if we compare it to the dense Qwen 27B, the gap will likely be even bigger. If you want, I can add the 27B version later. For now, I'm waiting for Qwen 3.8 27B to see how close it gets to the blueprint. As for the technical details: Both were run on a custom llama.cpp build optimized for the RTX 5080, with a temperature of 0.5 and a 125k context window. Regarding the music: I created it myself without using AI I specifically wanted it to sound that weird.

Reddit post · Benchmarks & research

DS4 vs Qwen3.6 vs Glimmer on one design prompt

@airawatraj

@airawatraj

Inference tuning notes for serving Muse Glimmer 30B NVFP4 with DFlash on a single NVIDIA DGX Spark as a consistent agent backend; the repo reports 27.5 tok/s average and 90/100 on its tool eval with 128K context.

GitHub · Local & open models

Muse Glimmer NVFP4 on DGX Spark

@r0b0tlab

@r0b0tlab

Reproducible native NVFP4 serving of Muse Glimmer 30B on one DGX Spark (GB10) with unmerged vLLM support: about 10.3 tok/s single-stream versus 4.2 for BF16, 52.5 tok/s at c16, and 131K context checked with needle-in-a-haystack tests.

GitHub · Local & open models

Muse Glimmer 30B NVFP4 on a single DGX Spark

U

do_u_think_im_spooky

u/do_u_think_im_spooky

Quick update on the RTX 5060 Ti local LLM repo. It has changed quite a bit since my previous posts. The project started as a collection of practical notes and benchmark results. That was useful, but as the dataset grew it became harder to answer the question most people actually had: What configuration should I run? I have rebuilt the repo around tested, copyable presets rather than treating every successful benchmark request as a front-page result. What changed? The project now separates three things: • Presets: exact configurations intended for people to copy and run. • Evidence bundles: reviewed proof of context fit, retrieval, sustained generation and performance. • Raw receipts: retries, failed experiments and diagnostic runs that are kept separate as engineering material without automatically becoming recommendations. The website now leads with the published preset catalogue. The larger results explorer is still there for comparisons and historical data, but it is no longer the first thing visitors have to decipher. There are currently seven published presets across the 1× and 2× RTX 5060 Ti lanes: 1× RTX 5060 Ti 16GB • Qwen3.8 27B IQ3_XXS at 64K with q8 K

Reddit post · Local & open models

club-5060ti: a tested Glimmer preset for 2x 5060 Ti

N

build.nvidia.com

build.nvidia.com

NVIDIA hosts a Muse Glimmer 30B endpoint on build.nvidia.com with Python (OpenAI, LangChain), JavaScript and curl examples for the ~29.6B multimodal model with 131K context.

Site · Local & open models

Muse Glimmer 30B on NVIDIA build

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

E

eesel.ai

eesel.ai

eesel AI reports Muse Spark 1.3 ranks #6 on the Artificial Analysis Intelligence Index, leads long-context and coding rows, but trails Claude Opus 5 on four of six agent evals.

M

dev.meta.ai

dev.meta.ai

Meta's developer blog introducing Muse Spark 1.2, co-trained with the new Muse Code terminal harness, with 1M-token context for multi-file refactors and hours-long tasks.

Resource · Coding & dev tools

Meet Muse Spark 1.2 and Muse Code

M

research.meta.ai

research.meta.ai

Meta's announcement of Muse Spark 1.3 for Muse Code and the Meta Model API, claiming ~20% fewer tool calls and ~25% fewer tokens than 1.2, with a max reasoning mode.

Resource · Coding & dev tools

Introducing Muse Spark 1.3

U

Ok-Inevitable8391

u/Ok-Inevitable8391

Benchmarked qwen3.8 xhigh, medium and muse glimmer. Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit) Medium effort mode and muse glimmer were 3-4 hours each. But I'm actually surprised by the muse glimmer results, they came better than the qwen. These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models. I have taken the result of claude models directly from embedeval repo by ecro. I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better. I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

Reddit post · Benchmarks & research

Glimmer vs Qwen 3.8 on an implicit-knowledge eval

U

PyaesoneP

u/PyaesoneP

I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8\_O KV cache. It's a joy to use a dense 30B model at this size and still get \~30 tok/s on a VRAM-constrained laptop. It's supposed to be only slightly worse than the official 17GB K-quant at a much smaller footprint, and for my Hermes Agent use case I don't notice a quality difference. It's just much faster. I've tried Qwen 3.8 27B at SC2.20bpw H3 too. Definitely usable but I'm sticking with Unsloth UD\_Q4\_K\_XL for Qwen 3.8 27B because it's mainly for coding.

Reddit post · Local & open models★ Pick

Muse Glimmer 30B on a 12GB laptop GPU

@SamuelAlexander

@SamuelAlexander

Samuel Alexander ran Muse Glimmer 30B entirely on a Qualcomm Dragonwing IQ-9075 board for zero-shot PCB defect inspection and tool calling, measuring 21.6 GB resident with full 131K context and 2.84 tokens/s generation.

GitHub · Local & open models★ Pick

Muse Glimmer 30B on a Qualcomm Dragonwing board

C

contextbolt.com

contextbolt.com

ContextBolt SEO gives Muse keyword research, difficulty scores, live SERPs, competitor rankings, backlinks and site audits through a hosted MCP endpoint, at $35/month after a 7-day trial.

Site · Business & commerce

ContextBolt SEO connector for Muse

@newjordan

@newjordan

A tested Muse Glimmer 30B Q4_K_M serving package for Intel Arc Pro B70 with full 131k context, reporting 19.0 tok/s decode at 129k cached and 503 tok/s full-context prefill.

GitHub · Local & open models

museB70 serving stack

@lobanov

@lobanov

An effort to make Muse Glimmer 30B actually use a 512k-token context (4x native) as a ~17GB GGUF in 32GB VRAM, trained on DGX Spark and evaluated with RULER-style retrieval tests.

GitHub · Local & open models· ★ 1

Muse Glimmer 512k context adaptation

vLLM

@vllm_project

@Meta is back in open source. Excited to announce Day-0 vLLM support for Muse Glimmer 30B, the first open-weights model from Meta Superintelligence Labs — which ships under Apache 2.0!!! 30B dense, 128K+ context, multimodal, built for local agents. Capable enough for

X post · Local & open models· ♥ 251

Day-0 vLLM support for Muse Glimmer

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

U

PathfinderTactician

u/PathfinderTactician

I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context, whilst Qwen3.6 27B can only run at 147,500 context - full GPU offload in both cases. Both models have been coding on an enterprise-grade web application. Detailed report of each model (warning - includes AI generated content): Diagnostic quality - comparable. Both have shown genuinely good root-cause work when they apply themselves. Qwen found coding issue and worked to fix things cleanly. Muse Glimmer correctly traced bugs and even caught something that a Frontier model missed after more than 10 rounds of review. Neither one is weak at diagnosis. Implementation reliability - Qwen ahead. Qwen did introduce real regressions into the coding along the way (eg. severe zone-scope refactor regression, and case-sensitivity regression) but each one eventually got fixed properly once caught, usually within one or two corrective rounds. Muse Glimmer did land fixes that were clean and verified true to spec. However, when working in a complex environment exceeding 200k context, Muse Glimmer fa

Reddit post · Benchmarks & research

BF16 Muse Glimmer vs Qwen3.6 27B on real code

U

slaybrownbeast

u/slaybrownbeast

I've been running my work in Codex as project folders, and recently tried to properly understand how Muse Goals work under the hood. Made it a goal — good way to watch the machinery operate on itself. The structural problem is worth naming: the current design is a halfway house between two coherent designs, and it gets the costs of both. Design A is Codex: the project is a container. Everything — chat, state, artifacts, scheduled work — lives in one place. My course project has one tracker file, explicit resume rules for new chats, and the curriculum never holds status. Legible, but you have to go to it. Design B is full ambient: no containers at all. The goal is just context that wakes up wherever you mention it, and there's no Goals tab pretending otherwise. Muse picked ambient for activation — talk about the goal anywhere, it wakes up, you never "open" it. But then it built half of containment: a Goals tab showing summary, artifacts, activity, without the other half. Conversations, check-ins, and briefings still leak into whatever chat they happened in. So you get the scattering of ambient with the implied promise of a container. That's the worst combination. The fix is to

Reddit post · Benchmarks & research

Rethinking Muse Goals from a Codex user

Meta for Developers

@MetaforDevs

AI developers - we’ve published a technical guide covering how to get started with Muse Spark on Meta Model API. Muse Spark is a multimodal reasoning model built for agentic tasks, coding, computer use and long-context workflows. See what you can build 👉 bit.ly/4vWMkZx

X post · Coding & dev tools· ♥ 150

Getting started guide for Muse Spark API

松xR

@matsu_vr

muse glimmer 30Bを公式の17GBに収まるよバージョンのggufでLM Studioで動かしてみました。写真の批評もバッチリしてくれる。これは僕が撮った写真なのでネットにないのでちゃんと写真を見ているはず。コンテクスト長64000にして、VRAM20GB以内に収まっているので、かなり実用的かもしれない!

X post · Local & open models· ♥ 2

Local photo critic with Glimmer in LM Studio

@n3xtgentechitalia

@n3xtgentechitalia

A llama.cpp container that serves Muse Glimmer 30B with vision and DFlash on an RTX 5090, reporting 98.8 tok/s single-stream and about 250 tok/s aggregate with 131k context.

GitHub · Local & open models

Muse Glimmer 30B on RTX 5090

U

StartupTim

u/StartupTim

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Muse Glimmer 30B running the day after release — then pushed its context from the trained 131K all the way to 1M with YaRN, verifying retrieval at every rung. Sharing config + results since the "131,072+" hint in the model card turned out to be very real. Setup • Hardware: 2× NVIDIA DGX Spark (GB10, 128 GB unified each, ~273 GB/s), ConnectX-7 direct link between them • Engine: llama.cpp master (day-1 muse_glimmer support), built from source with CUDA sm_121 + GGML_RPC • Model: official Muse-Glimmer-30B-GGUF K-Quant-Dynamic (~18.3 GiB) + official mmproj (vision) + official DFlash drafter • Spec decode: --spec-type draft-dflash --spec-draft-n-max 15 (block-diffusion drafter) • Context extension: --rope-scaling yarn --rope-scale <2/4/8> --yarn-orig-ctx 131072 plus --override-kv muse-glimmer.context_length=int:<N> (llama.cpp caps at trained length otherwise) • Yes, we also ran it split across both Sparks with llama.cpp RPC — no reason beyond liking to cluster things for fun. Our daily dri

Reddit post · Local & open models

Muse Glimmer at 1M context on 2x DGX Spark

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

U

TigerConsistent

u/TigerConsistent

Ran Muse Glimmer on a single RTX 3090 and found a max_tokens setting that made it look dumb; shares numbers at filled context and notes better-than-expected non-English handling.

Reddit post · Local & open models

Muse Glimmer on one 3090: the max_tokens gotcha

@aminamos

@aminamos

Native Muse Code plugin that wires in Hindsight memory with one bank per repo, using SessionStart, UserPromptSubmit and Stop hooks to recall and store context across agents. Zero dependencies, runs on macOS, Linux and Windows.

M

marktechpost.com

marktechpost.com

MarkTechPost walked through Muse for Mac, which pulls context from Calendar, Mail, Messages, Notes and files for multi-step tasks, with opt-in access, optional Full Disk Access and approval before deleting files or sending messages.

Resource · Errands & personal agent

Muse for Mac: cross-app tasks explained

@MiaAI-Lab

@MiaAI-Lab

A one-script vLLM setup that serves the roughly 19 GB NVFP4 Muse Glimmer 30B with its vision encoder kept, DFlash speculative decoding using the official drafter head, and up to 256K context on GB10, RTX 5090 or RTX PRO 6000.

GitHub · Local & open models· ★ 8

Muse Glimmer 30B NVFP4 for DGX Spark and RTX 5090

U

Inevitable-Diet-1870

u/Inevitable-Diet-1870

Hi all, Profile v2.2 is out. It's an open-source optimizer for inference servers. It computes your GPU's roofline ceiling, measures your live server against it, names the bottleneck, gives the flag. You apply. It re-measures. Every fix answers to a number. vLLM only today. More engines next. This release: core rule engine rewritten. Eight rules on a priority DAG with mutual exclusivity. Five alarms fire, four echoes are silenced, one true cause survives. Deterministic. AMD cards are now supported too. Tuning today is chaos: try a flag, wait, squint at a dashboard, repeat for days. Profile turns it into deterministic engineering: measure, fix, verify. Results in a few iterations. Mine took 4, ~30 minutes. My setup: RTX 5090, muse-glimmer 30B, SWE-Bench agents, no DFlash spec decoding. • 81 → 421 tok/s at 25k ctx • $3.41 → $0.65 per 1M output tok • TTFT 224ms (p95 500ms), TPOT 23ms at end of run • 4.72 → 1.08 J/tok https://preview.redd.it/4vazyxkcq6kh1.png?width=2248&format=png&auto=webp&s=77923a489b6f725240d23a7953150b5779260734 One iteration regressed hard: KV thrashing, TTFT 32.8s. Profile labeled it worse. Next fix recovered it. Regressions stay in the record. Watc

Reddit post · Local & open models

Profile v2.2: 421 tok/s Glimmer on one RTX 5090

S

sebastianraschka.com

sebastianraschka.com

Sebastian Raschka breaks down Glimmer's dense architecture: 3:1 sliding-window to global attention, 32 query heads with only 2 KV heads, and ~52 KiB of KV cache per token.

Resource · Benchmarks & research

Muse Glimmer 30B architecture notes

V

vercel.com

vercel.com

Vercel added Muse Spark 1.3 (meta/muse-spark-1.3) to AI Gateway on launch day, with 1M-token context, text/image/PDF input, and both standard and contributor pricing tiers.

Site · Coding & dev tools

Muse Spark 1.3 on Vercel AI Gateway

U

patricious

u/patricious

Benchmarked Muse Glimmer 30B on my RTX 5090 (32GB), 262k context, UD-Q5_K_M + dflash-kquant + mmproj. Workload Stock master + DFlash ngram-simple PR #26842 + DFlash Code patch 78 t/s 57 t/s 220-253 t/s Mixed agent turn 77 t/s 68 t/s 188-213 t/s Tool-call JSON 71 t/s 75 t/s 155-181 t/s Heavy reasoning 52 t/s 58 t/s 120-130 t/s PR #26842 moves the DFlash draft argmax from CPU to GPU, which was the bottleneck. I cherry-picked it onto master (it branched before the Muse merge, one conflict to resolve manually) and it builds clean. Code generation now matches Meta's published 233 t/s, which I could not reproduce on stock master. Notes: • ngram-simple loses to DFlash on every coding workload. • Server caps context at the model's metadata context_length, use --override-kv for 262k. • The reasoning budget flags do not work with this template. This is verified: with the budget set to 64, the model still burned 2000+ chars thinking and the budget message never appeared. Leave max_tokens headroom for the reasoning block. Flags: llama-server ^ --model Muse-Glimmer-30B-UD-Q5_K_M.gguf ^ --mmproj mmproj-kquant.gguf ^ -c 262144 --parallel 1 ^ --override-kv "muse-glimmer.context_le

Reddit post · Local & open models

253 t/s Glimmer on an RTX 5090

M

marktechpost.com

marktechpost.com

MarkTechPost summarizes Meta's numbers: 75.4 on DeepSWE v1.1 (Opus 5 74.0, GPT-5.6 Sol 72.7), 88.8 on Terminal-Bench 2.1, and 98.1 on MRCR v2 at 512K–1M context.

Resource · Benchmarks & research

MarkTechPost: Muse Spark 1.3 benchmarks breakdown

M

huggingface.co

huggingface.co

Meta's official Muse-Glimmer-30B repo: ~29.6B dense model with a 1.8B vision encoder, 131K context, Apache 2.0, with vLLM and SGLang serve commands.

Site · Local & open models

Muse Glimmer 30B weights on Hugging Face