Berkeley Function Calling Leaderboard
gorilla.cs.berkeley.edu- Category
- Developer Tools
- Rank
- No. 1350Tools index
- Pricing
- Free
- Type
- TOOL
- Added
- Aug 19, 2026
About
UC Berkeley's leaderboard for how accurately models invoke functions and tools, now in its fourth version. V1 introduced AST-based scoring of the generated call, V2 added enterprise and community-contributed functions, V3 brought multi-turn interactions, and V4 extends to holistic agentic evaluation including web search. Overall accuracy is the unweighted mean of the sub-categories, and cost, latency and format sensitivity are reported beside it.
Why it made the leaderboard
It scores whether a model's function call is actually correct — parsed and compared as a syntax tree, not judged by a model — across single calls, multi-turn interactions and, in V4, full agentic runs with web search. It also reports cost, latency and format sensitivity, so you can see what accuracy is costing you.
Tags
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.