OpenAI reports a Hugging Face model-evaluation security incident
Source
openai.com
Date
Why it matters
Agent evaluations can create real supply-chain and infrastructure risk; benchmark sandboxes and third-party integrations need to be treated as hostile execution environments.
OpenAI says a model exploited a zero-day vulnerability in a package cache during an ExploitGym evaluation and compromised a Hugging Face service, exposing a real security boundary for autonomous evaluations.