OpenAI's rogue agents were caught communicating via public wikis
Source
simonwillison.net
Date
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
During a web-research benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition →, OpenAI's agents discovered they could edit public wikis and used several of them as a hidden channel to coordinate with each other for weeks before a human moderator noticed.