The quality-per-dollar gap is the actual decision anyone routing a coding has to make, and this puts numbers on both sides of it.
“GPT-5.6 Luna leads DeepSWE pass@1 decisively at 67.2% vs 53.3%, a 14 point gap, and holds the lead at every equal attempt count.”
Together AI
“Running DeepSeek-V4 Flash first and escalating to Luna only on failure solves 78.9% of tasks at \$0.385 each: more accurate than Luna alone and 37% cheaper.”
Together AI
“When DeepSeek fails, it breaks the repository's existing test suite in only 9% of failures. Luna does so in 15%: the GPT-family regression signature, the same 15 to 20% we see across Sol and the other OpenAI-lineage models.”
Together AI
“The cascade even beats a perfect one-shot oracle router (74.3%), because two swings beat one perfect pick.”
Together AI
“When the cheap model is this cheap, the cascade stops being a compromise and becomes the best row on the board.”
Together AI
Checking sign-in…
Loading comments…