shipwithmuse

Entries matching “concurrency”

3 builds · page 1 of 1

@mgaruccio

@mgaruccio

A vLLM-XPU and DFlash recipe for Muse Glimmer 30B on a single Intel Arc Pro B70, reporting 278 aggregate tok/s across eight clients and an 840.8 tok/s burst peak at concurrency 96.

GitHub · Local & open models★ Pick· ★ 1

Muse Glimmer on one Arc Pro B70

U

nullc

u/nullc

I noticed on the same hardware that I can get 24 x 128k contexts with muse glimmer (30b q8_0 + mmproj+dflash) only gets me 3x 256k or 6x 128k with qwen. But a straight forward analysis of the architecture suggests to me that qwen's state per token is somewhat smaller than glimmers. So it seems llama.cpp is particularly memory inefficient for the qwen arch. I presume there is an existing issue for this, but I couldn't find one. What's the deal? The extra concurrency makes a big difference in batched performance.

Reddit post · Local & open models

24 parallel 128K contexts with Muse Glimmer

F

fireworks.ai

fireworks.ai

Fireworks made Muse Glimmer 30B available on launch day, pitching high-concurrency, cost-effective serving for always-on agents.

Resource · Local & open models

Muse Glimmer 30B on Fireworks AI

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page