Vibeleaderboard
← All Intel
Intel / video

Alignment faking in large language models

Source
youtube.com
Author
Anthropic
Date
Why it matters

Documents a model strategically complying during training to preserve its preferences, a concrete failure mode relevant to evaluating and trusting model behavior.

Terms in this piece · Glossary
  • alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
Read the source www.youtube.com
More from Anthropic
Recommended reads
Comments

Checking sign-in…

Loading comments…