token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Every new API key now carries 100M free tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → and 10x higher free-tier rate limits, enough headroom to benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → against your current stack on production-like traffic instead of a toy sample.