Vibeleaderboard
Index / tool
Visit huggingface.co
Category
Developer Tools
Rank
No. 3111Tools index

Previous survey · No. 2885 ·

Listed in
#28 Find AI benchmarks
Pricing
Free
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

A benchmark for general AI assistants from researchers at Meta and Hugging Face: 466 real-world questions that need reasoning, multi-modality, web browsing and tool use together, with 300 answers held back for the leaderboard and 166 released publicly. The gap it documented is the point — humans scored 92% where GPT-4 with plugins managed 15% — and it argues that robustness on questions humans find easy is a better milestone than ever-harder specialist tasks.

Why it made the leaderboard

466 questions that cannot be answered without chaining reasoning, browsing and tool use — the shape of real assistant work rather than a specialist exam. Its founding result frames the whole agent field: humans scored 92% where GPT-4 with plugins managed 15%, on questions people find straightforward.

Tags

benchmarkagentstool-useweb-browsingevaluationleaderboard

Media

GAIA

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.