eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
Meta's third release in four months sits near the intelligence-per-dollar frontier at about $0.40 per index task versus $0.51 for GPT-5.6 Terra, with gains concentrated in agentic evaluations.