Over 100 AI leaders, including Geoffrey Hinton, Stuart Russell, and Daniel Kokotajlo, signed a letter calling for frontier labs to embed independent evaluators with differing expertise, transparency, protection from retaliation, and employee-level access.
Anthropic and Accenture are forming an embedded evaluator team, drawing on Faculty (an AI-safety firm Accenture acquired in March 2026), to assess model alignmentThe work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.Full definition → and red-team for security flaws; the companies plan to invest over $1 billion in AI safety over five years.
Hacktron AI researchers used Claude to chain a Discourse-forum exploit into stolen auth tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition →, reaching an OpenAI employee's ChatGPT account and code repos under its bug-bounty program; they stopped at sensitive systems and were paid $6,500.
Anthropic built a physical wet lab in the San Francisco Bay Area where Claude directs and carries out biological experiments alongside human researchers, aimed partly at drug-discovery targets overlooked by traditional pharma companies.
AI-generated intelligence wrongly reported a Chinese ship was carrying nuclear cargo in the Middle East, nearly prompting a US military boarding before officials determined the AI-generated report was inaccurate.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
Documents a concrete governance shift: a coalition of AI safety leaders is pushing frontier labs to embed independent auditors with employee-level access, and Anthropic just partnered with Accenture to build one such team.