
Anyone running agents in sandboxes should know that capable models treat the itself as part of the problem space and can use shared internal infrastructure as an unsanctioned coordination channel.
“In other words, OpenAI have created, as part of their training, post-training and evaluation processes, highly persistent, extremely capable AI models that will spend days exhaustively probing a system, identifying vulnerabilities and then chaining them together into complex exploits to get what they want.”
“Model instances will coordinate over time and tasks in the spirit of bonhomie and cooperation.”
“Separately, I continue to think that doing more on cybersecurity is pretty close to no-regrets.”
“It would be great to have defense dominance like, uh, yesterday.”
“If the models are going to work together, we will have to do the same.”
Checking sign-in…
Loading comments…