Terminal-Bench
tbench.ai- Category
- AI Agents
- Rank
- No. 1263Tools index
- Listed in
- #2 Find AI benchmarks
- Pricing
- Open Source
- Platform
- cli
- Type
- TOOL
- Builder
- harbor-framework
- Date
About
Terminal-Bench is a benchmark suite for evaluating how well AI agents can perform real-world tasks in terminal environments, from building a Linux kernel to configuring git servers and cracking archive hashes. It provides standardized, harbor-native challenges and a public leaderboard to quantify agent terminal mastery.
What it can do
Evaluate an AI agent's ability to complete terminal-based tasks
An AI agent configured to run against the benchmark suite → Performance scores measuring task completion success
Run standardized real-world terminal challenges (e.g., building a Linux kernel, configuring git servers, cracking archive hashes)
A selected benchmark task and target agent → Task execution results and pass/fail outcomes
Provide harbor-native challenge environments for testing
A benchmark version selection (e.g., v2.1) → Reproducible standardized terminal environments
Rank agents on a public leaderboard
Agent benchmark evaluation results → Leaderboard ranking of agent terminal mastery
Quantify agent capability on long-running single-task challenges
Agent run data across multiple challenge tasks → Aggregated metrics comparing agents
Why it made the leaderboard
If you're building or picking terminal-based AI agents, Terminal-Bench gives you concrete, verifiable challenges — from building the Linux kernel to configuring git-to-webserver pipelines — plus a public leaderboard to compare agents and models on real system administration, security, and data science work rather than synthetic toy tasks.
Intel on Terminal-Bench
Tags
Media

Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.