Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Source
www.together.ai
Date
Why it matters
Cost-per-solve and routing results directly inform which model to send which coding task to.
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol.
Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.
Transcript
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.