Frontier agents come in far cheaper and faster than human workers on multi-hour office tasks yet still miss human-level deliverable quality — a cost-versus-quality baseline for anyone pitching agentic knowledge work.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.