Vibeleaderboard
Index / tool

Berkeley Function Calling Leaderboard

gorilla.cs.berkeley.edu
Visit gorilla.cs.berkeley.edu
Category
Developer Tools
Rank
No. 2752Tools index

Previous survey · No. 2609 ·

Listed in
#26 Find AI benchmarks
Pricing
Free
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

UC Berkeley's leaderboard for how accurately models invoke functions and tools, now in its fourth version. V1 introduced AST-based scoring of the generated call, V2 added enterprise and community-contributed functions, V3 brought multi-turn interactions, and V4 extends to holistic agentic evaluation including web search. Overall accuracy is the unweighted mean of the sub-categories, and cost, latency and format sensitivity are reported beside it.

Why it made the leaderboard

It scores whether a model's function call is actually correct — parsed and compared as a syntax tree, not judged by a model — across single calls, multi-turn interactions and, in V4, full agentic runs with web search. It also reports cost, latency and format sensitivity, so you can see what accuracy is costing you.

Tags

benchmarkfunction-callingtool-useagentsleaderboardevaluation

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.