Search
Searching the tools indexFilters
Type
Ready to use
Build with
Connect / Integrate
Extend
Operate
Reference
Topics
More topics
Source
Ranked under Top · All Time
- 626
agentscope-aiAI ToolsEvaluation framework that scores LLM quality with reward models and grader skills for RLHF and agent alignment.
Library / Frameworkhttps://github.com/agentscope-ai/openjudgeOpen Source
agentscope-ai · AI Tools
- 751
confident-aiDeveloper ToolsOpen-source LLM evaluation framework whose 40+ research-backed metrics for agents, RAG, and safety run as unit tests.
Library / Frameworkgithub.com/confident-ai/deepevalOpen Source
confident-ai · Developer Tools
- 868
explodinggradientsDeveloper ToolsEvaluation framework for RAG pipelines scoring retrieval and generation on faithfulness and context precision.
Library / Frameworkgithub.com/explodinggradients/ragasOpen Source
explodinggradients · Developer Tools
- 1442
vercel-labsAI AgentsPlayground for evaluating LLM agent runs by scoring tool calls, traces, and outputs side by side.
Applicationhttps://github.com/vercel-labs/agent-evalOpen Source
vercel-labs · AI Agents
- 2050
ProductivityDesktop app that turns your documents into a self-updating, interlinked wiki using an LLM.
github.com/nashsu/llm_wiki
Productivity
- 510
meta-llamaCybersecurityMeta's trust-and-safety toolkit for LLM security, spanning code scanning, jailbreak benchmarks, and I/O classifiers.
https://github.com/meta-llama/purplellamaOpen Source
meta-llama · Cybersecurity
- 917
harveyaiAI AgentsHarvey LAB (Legal Agent Benchmark)
Open-source benchmark of 1,671 legal tasks with an execution harness for running and scoring LLM agents.
Evaluation / Datasetgithub.com/harveyai/harvey-labsOpen Source
harveyai · AI Agents
- 2810Developer Tools
Continuously monitors LLM API endpoints for undisclosed output changes over time, using logprobs.
www.trackllm.net
Developer Tools
- 1120
supermemoryaiAI ToolsUniversal adapter that translates LLM input formats between providers, with built-in observability and error handling.
Library / Frameworkhttps://github.com/supermemoryai/llm-bridgeOpen Source
supermemoryai · AI Tools
- 1378
anthropicsAI ToolsPublic suite of reference tasks and frameworks for benchmarking Claude and other language models.
Evaluation / Datasethttps://github.com/anthropics/evalsOpen Source
anthropics · AI Tools
- 1958
@ActuallyIsaakDeveloper ToolsScores how well LLMs know Apple's MLX framework across a 441-question coding dataset.
Evaluation / Datasetgithub.com/goekdeniz-guelmez/mlx-benchmarkOpen Source
@ActuallyIsaak · Developer Tools
- 2677AI Tools
Adds retries, fallback, rate limiting, and caching around LLM API calls without a gateway.
vernllm.dev
AI Tools
- 1457
amantus-aiDeveloper ToolsConverts developer documentation into clean LLM-ready Markdown, with llms.txt support built in.
Utilityhttps://github.com/amantus-ai/llm-codesOpen Source
amantus-ai · Developer Tools
- 1333@LangChainDeveloper Tools
Tracing, evaluation, and monitoring for LLM apps, framework-agnostic and usable without LangChain.
Infrastructuresmith.langchain.comFreemium
@LangChain · Developer Tools
- 1895
Yeachan-HeoDeveloper ToolsRegression testing framework for AI prompts that records, replays, and compares LLM responses.
Library / Frameworkhttps://github.com/yeachan-heo/prompt-replayOpen Source
Yeachan-Heo · Developer Tools
- 660
meta-llamaAI ToolsToolkit from Meta for optimizing LLM prompts through evals and structured experiments rather than guesswork.
Library / Frameworkhttps://github.com/meta-llama/prompt-opsOpen Source
meta-llama · AI Tools









