benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Terminal-Bench-Science gives engineers a concrete way to measure whether agents can actually carry out scientific research workflows, not just toy coding tasks, across multiple science domains.