benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Hugging Face placed instructions in its security.txt aimed at AI agents told to find vulnerabilities there, pointing them at the public CyberGym benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → instead. It's a live example of using text files as a control surface for autonomous scanning agents.