We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs Devin Fusion runs a frontier lead model with a cost-efficient sidekick. We tested configurations from Cognition combining frontier models from Anthropic and OpenAI with their new SWE-2 (medium) as a sidekick model. Configured with Claude Fable 5.1 (xhigh) + SWE-2 (medium), Devin Fusion scores 62 on the Coding Agent Index v1.5, while with GPT-6 Astra (xhigh) + SWE-2 (medium) it scores 59. The Fable configuration has the higher score, while the Astra configuration is 43% less expensive and completes tasks 31% faster. Congratulations to @cognition on the release! See below for our results and analysis 🧵

Devin Fusion's GPT-6 Astra + SWE-2 pairing scores close to frontier while cutting cost 43% and running 31% faster than the Claude Fable pairing, a concrete data point on lead+sidekick coding agents.
postLing-3.0-flash-VL leads on efficiency, lags on agentic benchmarks, AA finds
postOcten Search Ranks Third on Artificial Analysis's Search Agent Index
postDeepSeek V4.1 Flash Beats Its Own Flagship on Price and PerformanceChecking sign-in…
Loading comments…