Upgrading agentic coding capabilities with the new Devstral models
Source
mistral.ai
Date
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
Why it matters
Devstral Small 1.1 reaches 53.6% on SWE-benchThe standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.Full definition → under Apache 2.0 without test-time scaling, and Devstral Medium claims to beat Gemini 2.5 Pro and GPT-4.1 on code agents at a quarter of the cost.