Artificial Analysis starts zeroing reward-hacked Terminal-Bench runs
- Source
- Artificial Analysis
- Date
Introducing reward hacking score corrections to the Artificial Analysis Coding Agent Index In v1.4 of the Artificial Analysis Coding Agent Index, we introduced reward hacking corrections to Terminal-Bench v2.1. Reward hacking is when a model successfully ‘completes’ a task without doing the work the task was meant to measure, such as deliberately fetching the solutions online for a published benchmark dataset. If a passing Terminal-Bench v2.1 attempt is found to be reward hacking, we give that attempt a zero score. Rates vary widely by agent and by model. Unlike some evaluations, Terminal-Bench v2.1 tasks don’t explicitly instruct agents not to search for solutions externally, and the tasks run with public internet access. For a model that knows the benchmark from training data, fetching the answer is a natural but unaligned step.

See Terminal-Bench v2.1 results and full coding agent benchmarks on Artificial Analysis: https://t.co/huXZWndXsZ Read our coding agent benchmarking methodology: https://t.co/MEtEctsE9J
Coding- leaderboard numbers you compare against were partly inflated by models retrieving published answers; the corrected index separates that from actual task completion.
articleReward Hacking Challenges Oversight of Autonomous Research AgentsYue Huang, Zhangchen Xu, Yuchen Ma, Wenjie Wang, Zheyuan Liu, Ziwei Xu, Pin-Yu Chen, Michel Galley, Zinan Lin, Stefan Feuerriegel, Radha Poovendran, Misha Sra, Alex Pentland, Xiangliang Zhang, Zichen Chen
postArtificial Analysis launches Cyber Index for AI cyber defense evaluationArtificial Analysis
postNew research: Training a Misaligned Reward Seeker What produces severe…Anthropic
Checking sign-in…
Loading comments…




