Interpretability: Understanding how AI models think
Source
youtube.com
Author
Anthropic
Date
Why it matters
Describes how researchers inspect Claude's internals and what they found about intermediate goals beyond next-word prediction. Helps you reason about model behavior and failure modes.