Vibeleaderboard
← All Intel
Intel / post

GPT-6.1 Sol launches with near-Astra performance and cheaper caching

Source
OpenAI Developers
Date
Thread · 7 parts

GPT-6.1 Sol is here. Upgraded with stronger agentic coding and computer use, near-Astra performance, and cached input at a 95% discount to standard input pricing. GPT-6.1 Sol is built for complex refactors, deep codebase investigations, and long-running agents across apps.

API pricing is $2 per million input tokens and $10 per million output tokens. Cached input is just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing. Near-Astra performance for sustained, multi-step workloads.

On DeepSWE v1.1, GPT-6.1 Sol achieves 75.2% at high reasoning effort, surpassing GPT-6 Sol’s best score of 68.8% at maximum effort, at approximately 76% lower cost per task.

On AutomationBench, GPT-6.1 Sol scores 31.7% at medium reasoning effort, up 4.8 percentage points from GPT-6 Sol at the same setting. On OSWorld 2.0’s offline set, it scores 71.4% versus Astra’s 73.5%, both at maximum reasoning effort, at roughly one-seventh of Astra’s cost per task.

Read the full thread on X
Key takeaways · AI-distilled
  • OpenAI prices GPT-6.1 Sol at $2 input and $10 output per million , with cached input at $0.10 per million: 95% below standard input pricing and 50% below GPT-6 Sol's cached input price.
  • OpenAI reports 71.4% on OSWorld 2.0's offline set versus 73.5% for Astra, both at maximum reasoning effort, at roughly one-seventh of Astra's cost per task.
  • On AutomationBench, OpenAI reports 31.7% at medium reasoning effort, 4.8 points above GPT-6 Sol at the same setting.
  • OpenAI says responses containing factual errors fell about 32% versus GPT-6 Sol on deliberately difficult factuality prompts at low effort, and its automated safety review saw no attempts to bypass the safety reviewer.
  • OpenAI's own guidance: choose Astra when maximizing quality matters most, and GPT-6.1 Sol for complex work you want to run more often. It is available in ChatGPT Work, Codex and the API.
Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters

GPT-6.1 Sol scores 75.2% on DeepSWE v1.1 at high effort, beating GPT-6 Sol's best at about 76% lower cost per task, with cached input at $0.10 per million tokens.

More from OpenAI Developers
Recommended reads
Comments

Checking sign-in…

Loading comments…