Self-bench
github.com/mupt-ai/self-bench- Category
- Developer Tools
- Rank
- No. 2363Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 28 stars
- Latest release
- prod-20261005-3057a78
- Date
About
Self-bench is an open-source tool that creates and runs coding-agent evals automatically from a repository's pull requests. Agents author and review Harbor environments generated from the PRs, the user approves each eval, and models and harnesses can be run concurrently in sandboxes. It lets teams measure agents on their own codebase instead of relying on public benchmarks that labs may optimize for, and supports connecting OpenAI and Claude subscriptions to avoid raw token costs.
Tags
evalsbenchmarkingcoding-agentpull-requestsharboropen-source
Tech Stack
Node.jsDockerTypeScript
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.