shipwithmuse

Entries matching “consistency”

2 builds · page 1 of 1

@airawatraj

@airawatraj

Inference tuning notes for serving Muse Glimmer 30B NVFP4 with DFlash on a single NVIDIA DGX Spark as a consistent agent backend; the repo reports 27.5 tok/s average and 90/100 on its tool eval with 128K context.

GitHub · Local & open models

Muse Glimmer NVFP4 on DGX Spark