About
Benchmark and evaluation toolkit for retrieval-augmented generation pipelines — measure how well your RAG actually answers questions.
Why it made the leaderboard
Benchmarks whether your RAG pipeline actually answers questions correctly — replaces eyeballing sample outputs with measurable evaluation of retrieval and generation quality.
Tags
ragevaluationllmbenchmark
Tech Stack
PythonShell
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
