Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank
No. 1948Tools index
Listed in
#32 Find AI benchmarks
Pricing
Open Source
Platform
cli
Type
TOOL
GitHub
533 stars
Date

About

Graduate-level science questions in biology, chemistry, and physics. Diamond is the expert-validated subset used in many model reports.

What it can do

  • Run zero-shot and few-shot baseline evaluations of GPT-3.5/GPT-4 on GPQA questions

    GPQA questionsModel answers/accuracy scores

  • Run chain-of-thought prompting evaluation on GPQA questions

    GPQA questionsModel answers with reasoning

  • Run open-book retrieval-augmented evaluation using Bing search

    GPQA questionsModel answers using retrieved web content

Why it made the leaderboard

Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.

Tags

benchmarkllm-evaluationgpqadatasetgpt-4science-qachain-of-thought

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.