We’ve released v2.1 of the Vals Index!
The new iteration replaces Terminal Bench 2.1 with Terminal Bench 4, and adds our proprietary Tax Agent Benchmark in its own category (weighted by GDP).
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Vals Index scores before and after v2.1 are not directly comparable, since Terminal Bench 4 replaces 2.1 and a GDP-weighted tax AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → category is added.