This week we spent about $95 trying to beat our own lineup of reviewing models. One of the candidates was Muse Spark 1.2, and it turned out to be the most interesting model in the whole test.
The good, measured:
• Among the best we tested at finding real problems. Scored against bugs we already knew were there, it matched our existing lineup, and it caught one real bug our lineup had missed.
• Fastest model in our table. Typical answer in 18 seconds, writing at over 220 tokens a second.
The speed table from our test (same job, same codebases, 33 runs per model):
Model Typical time Answer length (tokens) Writing speed (tok/s) Time follows answer length Time follows question length Muse Spark 1.2 18 s 4,205 222 0.79 barely (0.08) Gemini 3.1 Pro 19 s 2,621 133 0.99 no (0.0) Gemini 3.8 Flash 22 s 1,996 89 0.91 some (0.65) GPT 5.4 29 s 3,058 105 0.95 no (below 0) Grok 4.6 38 s 2,498 61 0.61 no (below 0) Grok 4.7 44 s * 3,176 75 0.98 a little (0.30) Claude Sonnet 5 50 s * 4,471 91 0.59 barely (0.07) * Runs that finished in time only, so the real typical time is higher. The last two columns are correlations: 1 means time rises in step with that length, 0 means no lin