benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
Why it matters
Compares eight commercial and six open weightsA model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.Full definition → speech systems on real conversational audio in 21 languages, including diarization. Teams picking an ASR model for non-English voice agents get per-language results and an open agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition →.