benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Same weights do not guarantee same behaviour across providers, and tool-calling accuracy is where the gap shows. The exacto endpoints route to the providers that measure best, and the burn-in data explains why fresh-model benchmarks overstate the spread.