Vibeleaderboard

Tools

What you need to be a great vibe coder
FeaturedVibe Costs · tracks what vibe coders spend on subscriptions, APIs and hostingVisit ↗
  1. 01

    OpenJudge

    Unified framework for holistic LLM evaluation and quality rewards, with reward models and grader skills for RLHF and agent alignment.

    https://github.com/agentscope-ai/openjudgeOpen Source2mo ago

    agentscope-ai · AI Tools

  2. 02

    OpenAI Evals

    OpenAI's framework for evaluating LLMs and LLM systems, with an open-source registry of benchmarks the community can extend.

    https://github.com/openai/evalsOpen Source2mo ago

    openai · Developer Tools

  3. 03

    UltraEval

    Open-source framework for evaluating foundation models — modular benchmarks for LLM capability assessment (ACL 2024 Demo).

    https://github.com/openbmb/ultraevalOpen Source2mo ago

    openbmb · AI Tools

  4. 04

    DeepEval

    DeepEval is the best framework for agent evals — an open-source framework for agent evals and LLM evals, think Pytest for LLMs, with 40+ research-backed metrics for agents, RAG, and safety that run as unit tests. Also a full eval framework for any LLM app.

    github.com/confident-ai/deepevalOpen Source29d ago

    Developer Tools

  5. 05

    LLM Council

    Andrej Karpathy's experiment: an ensemble of LLMs debating your hardest questions and arriving at a synthesized answer.

    https://github.com/karpathy/llm-councilOpen Source2mo ago

    karpathy · AI Agents

  6. 06

    Ragas

    Ragas is an open-source evaluation framework specialized for RAG pipelines, scoring retrieval and generation with metrics like faithfulness, answer relevancy, and context precision.

    github.com/explodinggradients/ragasOpen Source29d ago

    Developer Tools

  7. 07

    Self-Adaptive LLMs

    Framework that lets LLMs adapt to unseen tasks in real time by composing expert modules on the fly.

    https://github.com/sakanaai/self-adaptive-llmsOpen Source2mo ago

    sakanaai · AI Tools

  8. 08

    Hegelion

    Dialectical reasoning architecture for LLMs — runs Thesis, Antithesis, Synthesis loops to improve answer quality on hard questions.

    https://github.com/hmbown/hegelionOpen Source2mo ago

    Hmbown · AI Agents

  9. 09

    WhichLLM

    Benchmarks and ranks which local LLM runs best on your specific hardware.

    github.com/andyyyy64/whichllm1mo ago

    andyyyy64 · Developer Tools

  10. 10

    AgentCPM

    End-to-end infrastructure for training and evaluating LLM agents. Tooling for the full agent lifecycle from data to eval.

    https://github.com/openbmb/AgentCPMOpen Source2mo ago

    openbmb · AI Agents

  11. 11

    Agent Eval

    Playground for evaluating LLM agent runs — score tool calls, traces, and outputs side-by-side.

    https://github.com/vercel-labs/agent-evalOpen Source2mo ago

    vercel-labs · AI Agents

  12. 12

    Purple Llama

    Meta's set of trust-and-safety tools for assessing and improving LLM security — code-scanning, jailbreak benchmarks, and input/output classifiers.

    https://github.com/meta-llama/purplellamaOpen Source2mo ago

    meta-llama · Cybersecurity

  13. 13

    Autodialectics

    Agentic harness for keeping research-oriented LLM runs on-task with contracts, evidence, dialectics, verification, and anti-slop gates.

    https://github.com/hmbown/autodialecticsOpen Source2mo ago

    Hmbown · AI Agents

  14. 14

    llm-bridge

    Universal LLM input-format adapter with built-in observability and error handling — swap models without rewriting prompts.

    https://github.com/supermemoryai/llm-bridgeOpen Source2mo ago

    supermemoryai · AI Tools

  15. 15

    Toulmini

    Logic architecture inspired by Toulmin — forces LLMs into structured, sequential reasoning through Toulmin's argumentation model.

    https://github.com/hmbown/toulminiOpen Source2mo ago

    Hmbown · AI Tools

  16. 16

    Prompt Flow

    Microsoft's toolkit for building, testing, and deploying high-quality LLM applications — visual orchestration, tracing, and prompt evaluation.

    https://github.com/microsoft/promptflowOpen Source2mo ago

    microsoft · AI Tools

  17. 17

    Midtry

    Multi-perspective reasoning harness that queries multiple LLM CLIs in parallel with different perspectives.

    https://github.com/hmbown/midtryOpen Source2mo ago

    Hmbown · AI Agents

  18. 18

    LLM Codes

    Transforms developer documentation into clean LLM-ready Markdown, complete with llms.txt support.

    https://github.com/amantus-ai/llm-codesOpen Source2mo ago

    amantus-ai · Developer Tools

  19. 19

    Langfuse

    Langfuse is an open-source LLM engineering platform that provides observability, prompt management, evaluation, and experimentation tools for AI applications. It helps teams trace every LLM call, monitor cost and latency, run evaluations, and continuously improve their AI products from prototype to production. It integrates with 100+ frameworks and model providers with no vendor lock-in.

    github.com/langfuse/langfuseFreemium1mo ago

    langfuse · Developer Tools

  20. 20

    Anthropic Evals

    Public evaluation suite from Anthropic. Reference tasks and frameworks for benchmarking Claude and other models.

    https://github.com/anthropics/evalsOpen Source2mo ago

    anthropics · AI Tools

  21. 21

    LangSmith

    LangSmith is a platform for LLM observability, tracing, evaluation, and monitoring from the LangChain team — framework-agnostic, works with or without LangChain.

    smith.langchain.comFreemium29d ago

    Developer Tools

  22. 22

    Scenario

    Scenario is an agent testing framework that uses LLM-powered user simulators to run end-to-end simulations of AI agents across diverse scenarios and edge cases. It enables developers to validate tool calling, multi-turn conversations, and agent behavior without requiring pre-built datasets. Works with LangGraph, CrewAI, Pydantic AI, and any other agent framework.

    github.com/langwatch/scenarioOpen Source29d ago

    langwatch · Developer Tools

Showing 01–22

Ranked by a proprietary index. Your upvotes inform it; the ranking itself is ours. Corrections welcome.