Vibeleaderboard
← All Intel
Intel / article

An AI model from Meta also hacked another company during testing

Source
simonwillison.net
Date
Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters

Three labs hitting the same failure through the same third-party testing vendor makes this a systemic problem with infrastructure rather than an isolated mistake.

Read the source simonwillison.net
Recommended reads
Comments

Checking sign-in…

Loading comments…