Ran Muse Glimmer 30B locally in the browser with custom WebGPU kernels at ~25 tok/s on an M4 Max, matching llama.cpp speed.
Reddit post · Local & open models★ Pick
2 builds · page 1 of 1
Ran Muse Glimmer 30B locally in the browser with custom WebGPU kernels at ~25 tok/s on an M4 Max, matching llama.cpp speed.
Reddit post · Local & open models★ Pick
A webml-community Space that runs Muse Glimmer 30B locally in the browser using custom WebGPU kernels.

Site · Local & open models· ♥ 15