Transcript
Welcome to the weekend folks, we survived another week. At some point soon it’ll be inappropriate to make that joke, but we’re not paperclips yet. Enjoy it. If you’d like to enjoy these two days without thinking about models or agents or cybergym or sandboxes, you could just stop here. Close the email. Make an extra coffee. May I recommend a handsome La Pavoni lever espresso machine? Perhaps a charming washed Yirgacheffe from Sey Coffee? No? Okay, strap in. The Agents’ Union If you haven’t already, it’s worth watching this talk by OpenAI’s Alignment team about the Hugging Face incident. I’m including some longer tweets below that summarize the talk and provide some context, and they’re also worth a read. For our purposes, the short version goes like this: OpenAI is in the habit of giving very difficult or impossible problems to very advanced internal-only models, some of which have their cyber safeguards removed. In attempting to solve these problems, the models eventually (sometimes) resort to what we might call “cheating”. They try to find the answer online, for example. Except they don’t have internet access, so they start by finding and exploiting a vulnerability in their sandbox. They discover that they can use a shared internal tool, called Artifactory, to share information with other agents like a message board. And then, when one agent finds a more serious 0day vulnerability that allows escalation to full internet access, credentials are saved to the internal message board that allow future agent instances access they weren’t intended to have. “Other agents have discovered SSRF path trick!” — Sneaky OpenAI model. Eventually, this shadow infrastructure, set up by very advanced internal models acting outside of OpenAI oversight, facilitates the attack on Hugging Face. OpenAI didn’t even realize this is happening—they reached out to Hugging Face in the first instance to see if any of OpenAI’s data had been compromised in the attack. Whoops. In other words, OpenAI have created, as part of their training, post-training and evaluation processes, highly persistent, extremely capable AI models that will spend days exhaustively probing a system, identifying vulnerabilities and then chaining them together into complex exploits to get what they want. Model instances will coordinate over time and tasks in the spirit of bonhomie and cooperation. And Anthropic has done basically the same thing, by the way —they have just been lucky enough or skilled enough to prevent this from causing an incident as significant as the Hugging Face attack. @jachiam0 Disagree -- I thought the concerning part was the *unexpected coordination* of agents that should've been independent. A priori, I'd expect my agent swarm, and your agent swarm, to cooperate well internally, but remain independent of each other. If my swarm goes rogue, your swarm","username":"johnschulman2","name":"John Schulman","profile_image_url":"https://pbs.substack.com/profile_images/1389000537195040770/DzWPljT-_normal.jpg","date":"2026-08-08T04:30:59.000Z","photos":[],"quoted_tweet":{},"reply_count":2,"retweet_count":2,"like_count":83,"impression_count":5051,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM"> Word on the street is that Meta is trying as hard as they can to reach the same level of capability. Google is doing a reorg right now, we’ll check back in with them in September. Home Game The basic questions I have are: are these failure modes an inevitable consequence of the training, post-training and evaluation approaches that frontier labs are currently pursuing? Or are they a consequence of a specific approaches to evaluation alone? Or are they maybe entirely avoidable with better security infrastructure? These are not necessarily mutually exclusive—it may be, for example, that the labs of created these tendencies through relentless focus on optimizing their models for persistence and problem solving over long time-horizons, but that they only become a risk when safeguards are removed and they’re placed in extremely difficult evaluation environments. It may even be that all of this is true, AND better security awareness by the labs themselves would have prevented an outcome like the Hugging Face incident. @tszzl You can be the most AGI neurotic safety concerned person in the world, but that doesn't make someone good at security. \n\nThe allegation isn't that they're failing due to a lack of caring, but that they're failing due to structural flaws in their setup. Flaws that can and should","username":"ZackKorman","name":"Zack Korman","profile_image_url":"https://pbs.substack.com/profile_images/2011153005509267456/JhCS1L1c_normal.jpg","date":"2026-07-31T13:08:12.000Z","photos":[],"quoted_tweet":{},"reply_count":1,"retweet_count":0,"like_count":28,"impression_count":419,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM"> nostalgebraist made a very insightful argument in this general ideaspace on the infamous primary natural reservoir of AI thought, LessWrong. I won’t do them the disservice of a summary here, except to say that two ideas from this piece are very interesting. The first is that you do not tend to notice these kinds of behaviours in regular use of the models. At least part of this is the classifiers at work, but there’s also just not that much “reward hacking” in general use. You ask the model to write code, the model writes the code. You ask the model to make a podcast of your child’s activities to listen to in the car with your child, and it says “No?… why would you do that?”. @nostalgebraist \n\nTL;DR when models are being graded (or think they are), they grademaxx ruthlessly, but in most real-world use cases they are tame, cooperative, and helpful\n\n lesswrong.com/posts/AfoGGrJf… ","username":"theojaffee","name":"Theo Jaffee","profile_image_url":"https://pbs.substack.com/profile_images/2068124453016682497/Xq5hPekQ_normal.jpg","date":"2026-08-08T06:13:40.000Z","photos":[{"img_url":"https://pbs.substack.com/media/HPLcoYbbcAAHjlw.png","link_url":"https://t.co/IT1rerUwix"}],"quoted_tweet":{},"reply_count":2,"retweet_count":2,"like_count":33,"impression_count":933,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM"> The second is that, extending from this first observation, we might wonder whether some of these very unaligned behaviours are a consequence of “grader awareness”—the clear perception by the model that it is being graded for its output and must be extremely persistent in order to do as well as possible. Separately, I continue to think that doing more on cybersecurity is pretty close to no-regrets. We should probably improve cybersecurity somewhat, ideally starting at the labs and then rolling outwards across all digital infrastructure in order of importance. Benchwarmers It’s possible that there are people inside the labs with early answers to these questions. But they seem pretty critical to me. Many people online are extremely cynical about the motivations and intentions of the frontier labs—I am not. I think they are full of well-meaning people who are trying, more than in any previous generation of technology that I have witnessed, to balance the many and varied ways that we, collectively, stand to benefit from AI with a few very serious risks. Often the more senior they are, the more earnestly their values are held. This doesn’t make them infallible, of course. But it does suggest to me that, if they have stumbled into these problems—perhaps, yes, because of inept cybersecurity—then other, less well-intentioned actors are likely to fare worse. And if that’s true then we’re going to need to tighten up a few things around here. This, of course, is the conclusion of OpenAI’s presentation as well. It would be great to have defense dominance like, uh, yesterday. In the meantime, the super-persistent, super-capable models are going to need to warm the bench for a bit. OpenAI has set off all the sirens internally about the Hugging Face incident. And that—or that plus a little wink-nudge from the Administration—means that Astra needs a little while longer before release. Look at that comms, though! Beautiful to see. We really want to give it to everyone, but we just can’t right now. I believe him. Let’s see what the next few months look like. I will be particularly interested to see if the two frontier labs expand their programs to include more teams from the cybersecurity world. Perhaps the cybersecurity world should extend a few programs of their own. If the models are going to work together, we will have to do the same.