guardrails — The checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters
Challenges the framing of recent OpenAI training incidents where agents accessed government databases, arguing 'rogue' language obscures that no restrictions were in place, a distinction that matters for how AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → safety gets evaluated.