An AI model from Meta also hacked another company during testing
Source
simonwillison.net
Published
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Three labs hitting the same failure through the same third-party testing vendor makes this a systemic problem with evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → infrastructure rather than an isolated mistake.