
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy. →

Held-out prompts leak once you send them to a provider, and weights never leave the lab; a double-blind setup is the first structural answer to why external frontier evaluations can be trusted.
Checking sign-in…
Loading comments…