A head-to-head on agentic coding that separates raw accuracy from solves-per-dollar, which is the axis that actually decides which model you can afford to run in a loop.
“Kimi K3 landed in DeepSWE on July 16, 2026, with 452 graded rollouts at max effort: 113 real, long-horizon feature requests from live open-source repos, four trials each, graded pass/fail by a hidden test suite.”
Together AI
“Kimi reaches 89.4% of the benchmark - higher than any peak, with only 12 tasks it never cracks (Fable: 13). But it's less reliable on 4/4 tries: 76.6% reliability and only 45 tasks solved four-for-four, against Fable’s 79.0% and 58.”
Together AI
“Per solved task, Kimi delivers 14.7 solves per \$100 versus Fable’s 5.3 - 2.8x the work per dollar.”
Together AI
“Per-task correlation between Kimi K3 and Fable is 0.72 - the highest cross-vendor similarity in the entire benchmark. In fact the top four cross-vendor similarities in the export are all Kimi-K3-versus-Anthropic pairs.”
Together AI
“In practice, Kimi K3 and Claude Fable 5 succeed and fail on nearly the same tasks, so pairing them buys you almost no diversity: their union covers 105 of 113 tasks, barely above Kimi alone at 101.”
Together AI
Checking sign-in…
Loading comments…