Vibeleaderboard
← All Intel
Intel / article

GPT-Red

Source
openai.com
Date
Why it matters

remains the standing unsolved problem for deployments; an automated self-play red-teaming pipeline is a concrete signal of how model-side robustness is being built and tested.

Terms in this piece · Glossary
  • alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
  • prompt injection — An attack that hides instructions in content an AI will read — a webpage, email, or document — tricking it into following the attacker instead of the user.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Recommended reads
Comments

Checking sign-in…

Loading comments…