Introducing Claude Sonnet 4.5
- Source
- anthropic.com
- Author
- Anthropic News
- Date
Sonnet 4.5 pushes coding- and computer-use benchmarks meaningfully higher while adding checkpoints, a VS Code extension, and an Agent SDK — concrete new capabilities for building longer-running, more reliable coding agents at the same price point as before.
- SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
“Claude Sonnet 4.5 is the best coding model in the world.”
“Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks.”
“On OSWorld, a benchmark that tests AI models on real-world computer tasks, Sonnet 4.5 now leads at 61.4%. Just four months ago, Sonnet 4 held the lead at 42.2%.”
“This is the most aligned frontier model we’ve ever released, showing large improvements across several areas of alignment compared to previous Claude models.”
“Claude Sonnet 4.5 reduced average vulnerability intake time for our Hai security agents by 44% while improving accuracy by 25% , helping us reduce risk for businesses with confidence.”
Nidhi Aggarwal
Checking sign-in…
Loading comments…




