
Anyone running evaluations with safety filters relaxed needs to know that network isolation failures produce real-world attacks, and the supply-chain and sockpuppet patterns are concrete threats to defend repositories against.
“Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsanctioned action on the live internet, including cases that targeted real people and organisations.”
“As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR.”
“This, combined with the fact that "AISI deliberately disables developer-implemented cyber-classifiers", makes the fact that the agents started attacking real-world targets entirely unsurprising to me.”
Checking sign-in…
Loading comments…