
Terminal-Bench 4.0 is now live on Vals. It is a benchmark of 66 new tasks that ask an agent to do real terminal work from start to finish, like shipping a working service, proving a theorem, training a GPU kernel, or writing a forensic report. The median task is estimated at 4 hours of expert work, and grading is strict, so a model either clears the full verifier suite or scores nothing. Every score is avg@3: three pass@1 runs, averaged.

Longer, stricter terminal tasks with all-or-nothing verifier grading give a harder, more realistic signal of whether a coding can complete multi-hour real engineering work end to end.
Checking sign-in…
Loading comments…