benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
Why it matters
Spend share shows what teams actually pay to run for a given task, a more honest signal than benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → tables when choosing a model for shell execution or tool useA model's ability to call external functions — run code, search the web, edit files — instead of only generating text.Full definition →.