№ 0384Reddit post
Muse Glimmer + Hermes Agent as a private coding agent
Connected Hermes Agent to a llama.cpp server running the official Muse Glimmer 30B GGUF on an M5 Pro, getting ~22 t/s with the drafter; tool calls worked and the coding project ran.
Tested Muse Glimmer + Hermes Agent for Local/Private AI Agent & Coding
Setup the latest (master) version of llama.cpp server with the guide and the official GGUF weights by Meta AI: https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF and connected the Hermes Agent to the llama.cpp endpoint. Getting about 22t/s (+3-4t/s) on M5 Pro, using ~24GB including the drafter (provided by Meta). The model did correct tool calls and actually did some useful work inside the Hermes Agent. Moreover, the resulting coding task/project works, which was not the case when running the model with OpenCode. Watch more: https://www.youtube.com/watch?v=cmENEolUtM4



ChatForm
Tgmlabs