Vibeleaderboard

Tools

What you need to be a great vibe coder
FeaturedVibe Costs · tracks what vibe coders spend on subscriptions, APIs and hostingVisit ↗
  1. 01

    Terminal-Bench

    Terminal-Bench is a benchmark suite for evaluating how well AI agents can perform real-world tasks in terminal environments, from building a Linux kernel to configuring git servers and cracking archive hashes. It provides standardized, harbor-native challenges and a public leaderboard to quantify agent terminal mastery.

    tbench.aiOpen Source11d ago

    AI Agents

  2. 02

    SkillsBench

    The first benchmark framework for evaluating AI agent skills across 84 diverse tasks and 7 models. It measures how well AI agents perform when equipped with domain-specific skills versus without them, providing structured evaluation across multiple abstraction layers.

    www.skillsbench.aiOpen Source5mo ago

    AI Agents

  3. 03

    Anthropic Evals

    Public evaluation suite from Anthropic. Reference tasks and frameworks for benchmarking Claude and other models.

    https://github.com/anthropics/evalsOpen Source2mo ago

    anthropics · AI Tools

  4. 04

    OpenAI Evals

    OpenAI's framework for evaluating LLMs and LLM systems, with an open-source registry of benchmarks the community can extend.

    https://github.com/openai/evalsOpen Source2mo ago

    openai · Developer Tools

  5. 05

    Aider Polyglot Benchmark

    Coding problems across multiple languages used to benchmark Aider — reuse it to evaluate any AI coding agent.

    https://github.com/Aider-AI/polyglot-benchmarkOpen Source2mo ago

    aider-ai · Developer Tools

  6. 06

    Open Weights

    An open source AI scoreboard that aggregates and tracks everything happening in AI development. It provides a single feed for AI model benchmarks, research updates, and industry news.

    openweights.ioFree5mo ago

    AI Tools

  7. 07

    ProgramBench

    A benchmark that challenges AI agents to rebuild complete programs from scratch using only compiled binaries and documentation. Tests whether language models can reverse-engineer and implement working codebases that reproduce original program behavior.

    github.com/facebookresearch/programbenchOpen Source1mo ago

    @kunchenguid · AI Agents

  8. 08

    Cline Bench

    Real-world coding benchmarks derived from actual Cline user sessions: verified, challenging engineering problems.

    https://github.com/cline/cline-benchOpen Source2mo ago

    cline · Developer Tools

  9. 09

    SIA

    A self-improving AI framework that autonomously tunes its own performance.

    github.com/hexo-ai/sia1mo ago

    hexo-ai · AI Agents

  10. 10

    Auto-Harness

    A self-improving AI agent system that automatically mines failures from benchmarks, optimizes agent performance, and gates changes against regressions. It demonstrated improving agent scores from 0.56 to 0.78 on Tau3 benchmark tasks through autonomous iteration.

    github.com/neosigmaai/auto-harnessOpen Source3mo ago

    neosigmaai · AI Agents

  11. 11

    DesignArena

    DesignArena is the world's first crowdsourced benchmark for AI-generated design that compares results from top AI models side-by-side and lets users vote on the best outcomes. Millions of votes from users in 190+ countries power leaderboards that help determine which AI models have the best design capabilities.

    www.designarena.aiFree5mo ago

    @Design-Arena · AI Tools

  12. 12

    AI Battlegrounds

    An arena that pits LLMs against each other in actual games.

    https://github.com/jesseduffield/ai-battlegroundsOpen Source2mo ago

    jesseduffield · Entertainment

  13. 13

    UltraEval

    Open-source framework for evaluating foundation models — modular benchmarks for LLM capability assessment (ACL 2024 Demo).

    https://github.com/openbmb/ultraevalOpen Source2mo ago

    openbmb · AI Tools

  14. 14

    Roboflow Model Leaderboard

    Comparison leaderboard for object detection models — which is best for small or large objects, with handy benchmarks.

    https://github.com/roboflow/model-leaderboardOpen Source2mo ago

    roboflow · AI Tools

  15. 15

    RAGEval

    Benchmark and evaluation toolkit for retrieval-augmented generation pipelines — measure how well your RAG actually answers questions.

    https://github.com/OpenBMB/RAGEvalOpen Source2mo ago

    openbmb · Developer Tools

  16. 16

    Braintrust

    Braintrust is a platform for evaluating and shipping AI products — run evals, compare prompt/model versions, log production traces, and iterate in a prompt playground.

    braintrust.devFreemium29d ago

    Developer Tools

  17. 17

    Reranker Simple Benchmark

    Make running benchmarks simple yet maintainable — Korean cross-encoder reranker evaluation, designed for repeatable experiments.

    github.com/instructkr/reranker-simple-benchmarkOpen Source2mo ago

    instructkr · AI Tools

  18. 18

    DevTool Arena

    A benchmarking platform that evaluates how quickly AI agents can integrate with developer tools using only their documentation. It provides rankings and detailed reports to help devtool companies improve their AI-agent compatibility.

    2027.devFreemium4mo ago

    Developer Tools

Showing 01–18

Ranked by a proprietary index. Your upvotes inform it; the ranking itself is ours. Corrections welcome.