← Back to Vibers

Builder
LiveBench
1 Tool
LiveBench is an open benchmark for evaluating large language models across math, coding, reasoning, language, instruction-following, and data-analysis tasks; roughly a sixth of its questions are refreshed monthly using recent competitions, papers, and news articles to resist training-data contamination. It was created by a team led by Colin White and Samuel Dooley of Abacus.AI, with contributors from NYU, Nvidia, the University of Maryland, USC, and Columbia, and its paper was accepted at ICLR 2025. Because every question has a single verifiable answer, LiveBench scores models automatically without an LLM judge.