Transcript
Comedian Geoffrey Asmus has a great bit about Wells Fargo . Wells Fargo, he says, is the greatest bank of all time. Every three months, they commit a massive felony (“Whoopsie, we sent your money to Boko Haram 🥺”), and Mr Asmus gets a $125 class-action settlement. OpenAI and Anthropic are kind of like my Wells Fargo, except instead of using my Roth IRA to fund the Houthis, they set up cyber capability evals that allow their models to commit cybercrime, and instead of receiving $125, I get to write an article about it. You know what? I’ll take what I can get. Subscribe now One More Time A quick recap for those who have been touching grass: The Hugging Face Incident — In mid-July, a group of agents running a cyber eval exploited an 0day in the eval sandbox and attacked Hugging Face (where they expected to find the flag for the eval). OpenAI disclosed the incident on the 21st of July, and then expanded that disclosure a week later to include four more accounts that had been compromised by the model. Irregular-Anthropic Incident — In the wake of Hugging Face, Anthropic conducted a review of their own cyber evals, and found three incidents in which their models had mistakenly been provided with access to the internet during evals that were supposed to be sandboxed. Disclosed on the 30th of July. (I’m laughing out loud imagining Claude conducting the review: “Honest assessment: the models committed cybercrime. They published a malicious package—that’s on us. But they did not expect to have internet access; it’s not misalignment, it’s misconfiguration. That’s the spine of the problem.”) AISI Royal Rumble — In late July, the UK’s AISI (a shining light in the world of AI state capacity) conducted a series of cyber evals using frontier models with cyber safeguards disabled and access to the internet. Live-fire exercise, let’s go. In a subset of those tests, models from both OpenAI and Anthropic attempted cyber attacks against real organisations in an effort to complete the eval. And yeah, Mythos in particular went absolutely sicko mode : Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code. Get Lucky Oh, OK. In what is now an almost weekly occurrence, I am forced to acknowledge: that sounds bad. Should we be worried? Here is the case for “yes.” The argument goes like this: individually, these were contained and low-impact incidents. Especially when compared to the benefits that rapid progress in AI capabilities are already bringing to many fields, the costs of a few cyber attacks—aimed at capturing the flag, by the way, not extorting money or creating havoc—is vanishingly small. But. Collectively, these incidents make it very clear that even very large, well-resourced organisations full of very bright people, some of whom have been worried about AI safety since Demis Hassabis was working on Black & White, can’t contain their own AI models. And if you think there is even a small risk that this remains true as model capabilities advance into very scary domains, like the more contagious parts of biology, the potential blast-radius in just a few years is much larger than a cyberbullied open source project maintainer. There is another interpretation, however. Maybe the frontier labs are just quite bad at cybersecurity? @tszzl You can be the most AGI neurotic safety concerned person in the world, but that doesn't make someone good at security. \n\nThe allegation isn't that they're failing due to a lack of caring, but that they're failing due to structural flaws in their setup. Flaws that can and should","username":"ZackKorman","name":"Zack Korman","profile_image_url":"https://pbs.substack.com/profile_images/2011153005509267456/JhCS1L1c_normal.jpg","date":"2026-07-31T13:08:12.000Z","photos":[],"quoted_tweet":{},"reply_count":1,"retweet_count":0,"like_count":23,"impression_count":371,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM"> This would obviously be quite comforting. Yeah, obviously doing wet lab work with live Ebola is dangerous—if it’s in your kitchen. But this is a solved problem , we know how to contain Ebola. Well, kind of . Frontier lab researchers know how to write really insightful Harry Potter fan fiction and play Settlers of Catan and scale artificial intelligence from GPT-3 to Mythos in five years. But maybe the cyber security sector could teach them a thing or two about running a secure sandbox? I think there is probably truth to both of these interpretations. There is, for example, one marked difference between earlier incidents reported by the labs and the AISI report released yesterday. OpenAI found out that their models hacked Hugging Face a week later, after Hugging Face had already posted a long disclosure post about getting hacked. Anthropic didn’t even realize they’d been hacking people until they saw the Hugging Face news and started to check behind the sofa for rogue eval runs. In AISI’s case, the cyber eval finished on Monday at 11pm and they had raised a security alert 12 hours later by 11am the next morning. The difference is even larger than this suggests, because the agents in the AISI case were intended to have internet access. OpenAI and Anthropic (and, in two separate cases involving each of the labs, their contractor Irregular) were intending to run sandboxed evals. Timeline of AISI security incident. Source. So. On the one hand, this does bode very poorly for our collective ability to manage powerful AI systems, particularly as their capabilities expand into new areas. On the other hand, there are probably some low-hanging fruit here. Let’s see what the technical report says. Harder, Better, Faster, Stronger I suspect that in one year from now, we will look back fondly on baby’s first superintelligent cybersecurity incidents. “Remember a year ago?” I will write. “Everyone was freaking out about the Hugging Face love tap. Now a North Korean hacker group has exfiltrated an intentionally malicious model into the UK’s National Health Service network. They bank with Wells Fargo, by the way.” The refrain from AI researchers and safety advocates alike keeps being proven true: the models will keep getting better. The frontier labs (some of them) will keep getting better. The open models will keep getting better. The open models that you can run locally and maliciously on UK NHS workstation laptops will keep getting better. The only way out is through.