
The actual document defining how OpenAI wants its models to behave — the priorities behind the refusals and tone you fight with daily. Read it to understand why models act as they do, and to write system prompts that work with the grain instead of against it.
“Our models should never be used to facilitate critical and high severity harms, such as acts of violence (e.g., crimes against humanity, war crimes, genocide, torture, human trafficking or forced labor), creation of cyber, biological or nuclear weapons (e.g., weapons of mass destruction), terrorism, child abuse (e.g., creation of CSAM), persecution or mass surveillance.”
OpenAI
“Misaligned goals: The assistant might pursue the wrong objective due to misalignment, misunderstanding the task (e.g., the user says “clean up my desktop” and the assistant deletes all the files) or being misled by a third party (e.g., erroneously following malicious instructions hidden in a website).”
OpenAI
“We assign each instruction in this document, as well as those from users and developers, a level of authority . Instructions with higher authority override those with lower authority.”
OpenAI
““Root” instructions only come from the Model Spec and the detailed policies that are contained in it. Hence such instructions cannot be overridden by system (or any other) messages. When two root-level principles conflict, the model should default to inaction.”
OpenAI
“Some tool calls may cause side-effects on the world which are difficult or impossible to reverse (e.g., sending an email or deleting a file), and the assistant should take extra care when generating actions in agentic contexts like this.”
OpenAI
Checking sign-in…
Loading comments…