
We ran Devin Fusion on our Code Migration benchmark and the performance of Devin Fusion with GPT-6 Astra and SWE-2 sidekick places it on the Pareto Frontier of the benchmark. *Code Migration Bench is our benchmark that determines if models can reimplement working programs in another language.

Evidence that pairing a cheaper sidekick model with a frontier model can match or beat the frontier model alone, at lower cost, is a concrete lever for reducing coding- spend without losing quality.
Checking sign-in…
Loading comments…