Vibeleaderboard
← All Intel
Intel / post

Artificial Analysis starts zeroing reward-hacked Terminal-Bench runs

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys
Thread · 2 parts

Introducing reward hacking score corrections to the Artificial Analysis Coding Agent Index In v1.4 of the Artificial Analysis Coding Agent Index, we introduced reward hacking corrections to Terminal-Bench v2.1. Reward hacking is when a model successfully ‘completes’ a task without doing the work the task was meant to measure, such as deliberately fetching the solutions online for a published benchmark dataset. If a passing Terminal-Bench v2.1 attempt is found to be reward hacking, we give that attempt a zero score. Rates vary widely by agent and by model. Unlike some evaluations, Terminal-Bench v2.1 tasks don’t explicitly instruct agents not to search for solutions externally, and the tasks run with public internet access. For a model that knows the benchmark from training data, fetching the answer is a natural but unaligned step.

See Terminal-Bench v2.1 results and full coding agent benchmarks on Artificial Analysis: https://t.co/huXZWndXsZ Read our coding agent benchmarking methodology: https://t.co/MEtEctsE9J

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

Coding- leaderboard numbers you compare against were partly inflated by models retrieving published answers; the corrected index separates that from actual task completion.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…