shipwithmuse

Entries matching “stata”

1 build · page 1 of 1

khaled

@eltokh7

Muse Spark 1.1 seems like a very good model. I tested it on Stata Benchmark and it ranks 4th (!) beating Opus 4.7/4.8, tied with GPT 5.5. I like this benchmark because it often catches seemingly good model performing poorly on OOD tasks

X post · Benchmarks & research· ♥ 17

Spark 1.1 ranks 4th on Stata Benchmark