Deliberately attacking your own AI system to find what makes it fail before someone else does.
Standard evaluation asks whether the system works on intended inputs. Red teaming asks what an adversary can make it do: leak a system prompt, ignore its instructions, produce something harmful, take an action on a hostile page's behalf.
For agents the surface is larger than for chat, because a successful attack yields actions rather than words. Any agent that reads untrusted content and can also send, spend, publish, or delete deserves an explicit attempt to hijack it before it ships.