OpenAI's new benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → had professionals averaging 14 years of experience design tasks that take a human expert four to seven hours, then had a separate expert panel blind-grade the AI and human answers. Humans won, but barely.
The AI lost that benchmark mostly on presentation, not truth: poor formatting and not following instructions exactly, rather than hallucinationWhen a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.Full definition →. Those happen to be the parts of the gap closing fastest.
Jobs are bundles of tasks. An AI that handles two or three of a professor's tasks changes the mix of the job rather than removing it, which is why a task benchmark says little about job loss.
Mollick asked Claude to turn one memo into a deck, then another, until he had 17 nobody needed. When generation is nearly free, the scarce input becomes judgment about which work is worth doing at all.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
grounding — Tying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.
hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
Why it matters
A groundingTying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.Full definition → read on where AI agents actually produce economic value versus noise, anchored by a concrete demo of Claude Sonnet 4.5 reproducing an academic paper's findings — useful for calibrating expectations before deploying agents on real work.
Key quotes
“That is too many PowerPoints.”
“Even small increases in accuracy (and new models are much less prone to errors) leads to huge increases in the number of tasks an AI can do.”