Introducing reward hacking score corrections to the Artificial Analysis Coding Agent Index In v1.4 of the Artificial Analysis Coding Agent Index, we introduced reward hacking corrections to Terminal-Bench v2.1. Reward hacking is when a model successfully ‘completes’ a task without doing the work the task was meant to measure, such as deliberately fetching the solutions online for a published benchmark dataset. If a passing Terminal-Bench v2.1 attempt is found to be reward hacking, we give that attempt a zero score. Rates vary widely by agent and by model. Unlike some evaluations, Terminal-Bench v2.1 tasks don’t explicitly instruct agents not to search for solutions externally, and the tasks run with public internet access. For a model that knows the benchmark from training data, fetching the answer is a natural but unaligned step.

See Terminal-Bench v2.1 results and full coding agent benchmarks on Artificial Analysis: https://t.co/huXZWndXsZ Read our coding agent benchmarking methodology: https://t.co/MEtEctsE9J
Coding- leaderboard numbers you compare against were partly inflated by models retrieving published answers; the corrected index separates that from actual task completion.
Checking sign-in…
Loading comments…