Tools
- 01
agentscope-aiAI ToolsOpenJudge
Unified framework for holistic LLM evaluation and quality rewards, with reward models and grader skills for RLHF and agent alignment.
https://github.com/agentscope-ai/openjudgeOpen Source2mo ago
agentscope-ai · AI Tools
- 04
Developer ToolsDeepEval
DeepEval is the best framework for agent evals — an open-source framework for agent evals and LLM evals, think Pytest for LLMs, with 40+ research-backed metrics for agents, RAG, and safety that run as unit tests. Also a full eval framework for any LLM app.
github.com/confident-ai/deepevalOpen Source29d ago
Developer Tools
- 06
Developer ToolsRagas
Ragas is an open-source evaluation framework specialized for RAG pipelines, scoring retrieval and generation with metrics like faithfulness, answer relevancy, and context precision.
github.com/explodinggradients/ragasOpen Source29d ago
Developer Tools
- 11
vercel-labsAI AgentsAgent Eval
Playground for evaluating LLM agent runs — score tool calls, traces, and outputs side-by-side.
https://github.com/vercel-labs/agent-evalOpen Source2mo ago
vercel-labs · AI Agents
- 12
meta-llamaCybersecurityPurple Llama
Meta's set of trust-and-safety tools for assessing and improving LLM security — code-scanning, jailbreak benchmarks, and input/output classifiers.
https://github.com/meta-llama/purplellamaOpen Source2mo ago
meta-llama · Cybersecurity
- 14
supermemoryaiAI Toolsllm-bridge
Universal LLM input-format adapter with built-in observability and error handling — swap models without rewriting prompts.
https://github.com/supermemoryai/llm-bridgeOpen Source2mo ago
supermemoryai · AI Tools
- 18
amantus-aiDeveloper ToolsLLM Codes
Transforms developer documentation into clean LLM-ready Markdown, complete with llms.txt support.
https://github.com/amantus-ai/llm-codesOpen Source2mo ago
amantus-ai · Developer Tools
- 19
langfuseDeveloper ToolsLangfuse
Langfuse is an open-source LLM engineering platform that provides observability, prompt management, evaluation, and experimentation tools for AI applications. It helps teams trace every LLM call, monitor cost and latency, run evaluations, and continuously improve their AI products from prototype to production. It integrates with 100+ frameworks and model providers with no vendor lock-in.
github.com/langfuse/langfuseFreemium1mo ago
langfuse · Developer Tools
- 20
anthropicsAI ToolsAnthropic Evals
Public evaluation suite from Anthropic. Reference tasks and frameworks for benchmarking Claude and other models.
https://github.com/anthropics/evalsOpen Source2mo ago
anthropics · AI Tools
- 21Developer Tools
LangSmith
LangSmith is a platform for LLM observability, tracing, evaluation, and monitoring from the LangChain team — framework-agnostic, works with or without LangChain.
smith.langchain.comFreemium29d ago
Developer Tools
- 22
langwatchDeveloper ToolsScenario
Scenario is an agent testing framework that uses LLM-powered user simulators to run end-to-end simulations of AI agents across diverse scenarios and edge cases. It enables developers to validate tool calling, multi-turn conversations, and agent behavior without requiring pre-built datasets. Works with LangGraph, CrewAI, Pydantic AI, and any other agent framework.
github.com/langwatch/scenarioOpen Source29d ago
langwatch · Developer Tools
Ranked by a proprietary index. Your upvotes inform it; the ranking itself is ours. Corrections welcome.







