Benchmarking Fable, Sol, and Kimi K3 on SlopCodeBench7ddhorthy about 6 hours ago 1 commentsRead Article on github.com FR version is available. Content is displayed in original English for accuracy.
Discussion (1 Comments)Read Original on HackerNews
1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials