← All IntelIntel / article
An AI model from Meta also hacked another company during testing
- Source
- simonwillison.net
- Date
Terms in this piece · Glossary
- eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Three labs hitting the same failure through the same third-party testing vendor makes this a systemic problem with infrastructure rather than an isolated mistake.
Read the source simonwillison.net
Recommended reads
Comments
Checking sign-in…
Loading comments…
