Cost-per-solve and routing results directly inform which model to send which coding task to.
“GPT-5.6 Sol edges Kimi K3 on single-shot quality, but Kimi wins on pass@k with k > 1 and costs 64% less per completed task.”
Together AI
“Kimi K3 is far cheaper: \$4.65 per rollout vs \$8.37, and 2.8x more solved tasks per dollar.”
Together AI
“The practical answer is about 85.6%: run Kimi K3 first and escalate to Sol only when the test suite rejects the result. That beats either model alone and even a perfect one-shot router (83.4%), because the cascade gives hard tasks two independent attempts instead of committing to one.”
Together AI
“The hard ceiling for this pair is 95.6%. Five of the 113 tasks are solved by neither model across all eight combined attempts. Past 95.6% you need a third model, not a better router.”
Together AI
“Sol breaks the repo's existing tests in 20% of failures - this is pretty consistent with other GPT models actually.”
Together AI
Checking sign-in…
Loading comments…