
If you build or evaluate -based products, Deployment Simulation shows how to forecast model behavior pre-release using real conversation data rather than static benchmarks, improving the fidelity of safety and pipelines.
“The main technique behind Deployment Simulation is simple: we take recent conversations from deployment, remove the original assistant response from the older model, and regenerate it with a candidate model to be released.”
OpenAI
“In total, we analyzed approximately 1.3 million de-identified conversations across GPT‑5 Thinking through GPT‑5.4 deployments, spanning August 2025 to March 2026.”
OpenAI
“Calculator hacking is a form of reward hacking which involves the model using a browser tool as a calculator while presenting the action as a search.”
OpenAI
“For example, GPT‑5.2 labeled real production traffic as evaluation-like 5.4% of the time and simulated traffic 5.1% of the time, while traditional evaluations (such as SWE-Bench or MonorepoBench, which are composed primarily of synthetic examples) were classified as evaluations almost 100% of the time.”
OpenAI
“In a test where a discriminator tried to tell simulated from real rollouts, realism improved from an 11.6% win rate to 49.5%, near the 50% chance level.”
OpenAI
Checking sign-in…
Loading comments…