Vibeleaderboard
← All Intel
Intel / post

Stanford Releases Terminal-Bench-Science for AI Research Agents

Source
Stanford AI Lab
Date
Stanford AI Lab@StanfordAILab

Terminal-Bench-Science is out! https://t.co/X4OYTeSmho

Terms in this piece · Glossary
  • benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

Terminal-Bench-Science gives engineers a concrete way to measure whether agents can actually carry out scientific research workflows, not just toy coding tasks, across multiple science domains.

More from Stanford AI Lab
Recommended reads
Comments

Checking sign-in…

Loading comments…