GPT-6 Sol and Luna are now available in Devin.
On FrontierCode 1.1, GPT-6 Sol matches GPT-5.6 Sol’s score at 61% lower cost per task. GPT-6 Luna scores above GPT-5.6 Luna at about a quarter of the cost. At under $0.10 per task, it is the cheapest model on the leaderboard.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters
Concrete cost-per-task and benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → numbers help engineers pick cheaper models for coding-AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → workloads without giving up accuracy.