inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Standard benchmarks failed to predict who won in a live adversarial loop, and cost per successful run varied 27x between models. Worth reading before treating benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → rank as a routing decision.