Vibeleaderboard
← All Intel
Intel / post

Devin Adds GPT-6 Sol and Luna at Sharply Lower Cost

Source
Cognition
Date
Cognition@cognition

GPT-6 Sol and Luna are now available in Devin. On FrontierCode 1.1, GPT-6 Sol matches GPT-5.6 Sol’s score at 61% lower cost per task. GPT-6 Luna scores above GPT-5.6 Luna at about a quarter of the cost. At under $0.10 per task, it is the cheapest model on the leaderboard.

Context

Cognition, the company behind the autonomous coding Devin (@cognition), has added OpenAI's newest models, GPT-6 Sol and GPT-6 Luna, as selectable options in Devin Desktop and CLI. Cognition tested both on FrontierCode 1.1, its own internal for whether a model's generated pull request is good enough to actually merge rather than just whether it passes automated tests, as described in Cognition's blog post announcing the change. On that benchmark, GPT-6 Sol matched the score of its predecessor, GPT-5.6 Sol, while costing 61% less per task, and GPT-6 Luna scored above GPT-5.6 Luna at about a quarter of the cost, putting it under $0.10 per task and, by Cognition's count, the cheapest model on its leaderboard.

These are Cognition's own numbers on its own benchmark, not an independent test, so they describe relative cost per completed coding task inside Devin rather than general model capability. The pattern matches what OpenAI itself reported the same day, that GPT-6 Sol and Luna carry roughly half the API price of the GPT-5.6 models they replace, and Artificial Analysis's separate measurement that the two models match their predecessors' Intelligence Index scores at about half the cost. For someone choosing a model for agentic coding work, the concrete result is a cheaper option that, on Cognition's own test, did not trade away code quality.

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
More from Cognition
Recommended reads
Comments

Checking sign-in…

Loading comments…