Terminal-Bench
tbench.ai- Category
- AI Agents
- Pricing
- Open Source
- Platform
- cli
- Type
- TOOL
- Added
- Jul 21, 2026
About
Terminal-Bench is a benchmark suite for evaluating how well AI agents can perform real-world tasks in terminal environments, from building a Linux kernel to configuring git servers and cracking archive hashes. It provides standardized, harbor-native challenges and a public leaderboard to quantify agent terminal mastery.
Why it made the leaderboard
If you're building or picking terminal-based AI agents, Terminal-Bench gives you concrete, verifiable challenges — from building the Linux kernel to configuring git-to-webserver pipelines — plus a public leaderboard to compare agents and models on real system administration, security, and data science work rather than synthetic toy tasks.
Tags
ai-agentsbenchmarkterminalevaluationllmleaderboardcli
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.