№ 0118Reddit post
Testing the Glimmer DFlash drafter on an M5 Mac
Measured Muse-Glimmer-30B-4bit with and without the DFlash drafter on an M5 Pro MacBook via oMLX; speedup was only 0.90%-1.16%.
u/OkSea7809 ran three prompt types (technical prose, Python, JSON) three times each with a 2048-token cap and temperature 0 on a 48GB M5 Pro MacBook Pro via oMLX, finding the Muse-Glimmer-30B-Assistant DFlash drafter gave 0.90%-1.16% speedup.
Muse-Glimmer-30B-Assistant negligible speedup on M5
Hi all, I'm a newbie and trying to assess the performance of some LLMs I'm running locally via oMLX on my MacBook Pro M5pro CPU 15 cores (5 Super and 10 Performance), GPU 16 cores and 48 GB of LPDDR5 RAM. I asked chatGPT guidance to run some tests and check whether the DFlash-based drafter Muse-Glimmer-30B-Assistant might somewhat speedup the base model Muse-Glimmer-30B-4bit. The results show no or negligible improvement with active DFlash acceleration (speedup between 0.90% and 1.16%). The test was structured with three different prompts fed to both the baseline and the dflash-capable model profiles: Technical prose; Python code; Structured JSON a cap of 2048 tokens, no cache, temperature=0. Each inference was repeated three times. Anyone have similar experience? can we simply dump the Assistant as not useful in this hw/sw configuration?



ChatForm
Tgmlabs