№ 0098Resource
Why Muse Spark 1.3's benchmark scores don't add up
MindStudio notes Muse Spark 1.3 topped DeepSWE at 75.4 and placed third on Artificial Analysis, yet in a hands-on game-clone test produced "a cube shooting at other cubes."
The author argues narrow coding benchmarks may not predict open-ended creative coding. In the same test, lower-ranked competitors such as Gemini 3.8 Flash did better.









ChatForm
Tgmlabs