
Published tau-bench numbers shift, and the fixed grader stops punishing agents for being appropriately cautious.
“Under the grading-scheme fixes, re-grading the leaderboard trajectory sets moves scores **only upward** — no previously-passing simulation fails — by up to ~9 points pass^1 depending on the model.”
“One prudent verification read not present in the golden trajectory — e.g. listing a user's accounts right after opening one — failed the task, even though the knowledge base encourages such reads.”
“This made the "dispute the earliest duplicate" tie-breaker unresolvable from tool output on the duplicate-charge tasks (083–085).”
“This is a gold-value change: trajectories that reproduced the old $8.00 refund fail task_074 under 1.0.1, while policy-faithful $14.50 refunds now pass.”
Checking sign-in…
Loading comments…