shipwithmuse

Entries matching “evaluation”

8 builds · page 1 of 1

@TrevorS

@TrevorS

Refusal-direction ablation on Muse-Glimmer-30B that cut refusals from 128/150 to 3/150, adding an agentic-safety evaluation and publishing bf16 and GGUF uncensored weights.

GitHub · Local & open models★ Pick· ★ 1

Muse Glimmer abliteration

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

@murpheycandler

@murpheycandler

A small Inspect evaluation on Muse Glimmer that tests whether incentive framing changes what an agent reports to its principal when the evidence is held constant; the author reports a null result.

@youdotcom-oss

@youdotcom-oss

You.com's harness for evaluating Muse Glimmer 30B on DeepSearchQA with You.com MCP tools inside pi sessions. A custom RLM v5 extension reached F1 0.8054 on 50 tasks (0.6910 over 900x3) versus 0.50 for plain skill injection.

GitHub · Benchmarks & research

Muse Glimmer DeepSearchQA skill eval

John Yang

@jyangballin

My favorite demo from the launch: muse spark 1.1 + opencode runs evaluation of *itself* + mini-SWE-agent (by @KLieret, @closji, urs truly) on DeepSWE!

X post · Benchmarks & research· ♥ 40

Muse Spark 1.1 evaluates itself on DeepSWE

@Satgoy152

@Satgoy152

A project that fine-tunes a DSpark speculator for Muse Glimmer 30B on on-policy coding and agentic traces to raise acceptance length on agentic workloads, evaluated with Terminal-Bench.

@lobanov

@lobanov

An effort to make Muse Glimmer 30B actually use a 512k-token context (4x native) as a ~17GB GGUF in 32GB VRAM, trained on DGX Spark and evaluated with RULER-style retrieval tests.

GitHub · Local & open models· ★ 1

Muse Glimmer 512k context adaptation

@JacobStephens2

@JacobStephens2

A reproducible evaluation of a stochastic defect where Muse Spark 1.1, under some agent-like request envelopes, wrote to claude-smoke.txt when asked for muse-smoke.txt; a neutral control kept 80/80 filenames intact.

GitHub · Benchmarks & research

Muse Spark 1.1 filename-substitution repro

U

MajesticAd2862

u/MajesticAd2862

Compared diarization models on 15 mock doctor-patient consultations (~2.4 h): Meta Muse Voice Transcribe scored 13.04% DER at ~92 s per request via API, behind Pyannote (2.89%) and Nemotron 3 (4.80%).

Reddit post · Benchmarks & research

Muse Voice Transcribe tested on clinical diarization