Safety refusal reporting is now available in the Artificial Analysis Coding Agent Index In our latest Coding Agent Index v1.5, we’ve introduced safety refusal reporting to help explain model behavior and score differences. A safety refusal occurs when a provider or model declines to start or continue a task on safety grounds. An agent may fall back to another model to continue, or stop the attempt with a block. Claude Fable 5.1 had the highest fallback rates in both Claude Code and Devin Fusion, with fallback attempts accounting for 8.8% and 7.1% of the Index's weight, respectively; these results therefore include the fallback models' performance. Refusal variability, harness context buildup, effort settings, and retry strategies can all affect the observed rates.

Reveals a hidden factor behind coding- scores: how often a model refuses or falls back mid-task, which affects observed performance independent of raw capability.
postGrok Voice Transcribe 2.0 Tops Streaming ASR Accuracy Benchmark
postAnt Group's finance model matches MiniMax-M2.7 with half the parameters
postOpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as itsChecking sign-in…
Loading comments…