AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Separates 'understands physics' from 'is usable as a simulator': models are scored by whether acting on their rollouts succeeds in the real environment, under frozen policies and without expensive search.