← All IntelClip / EntertainmentWhy coding is core to Anthropic's safety mission
From Boris Cherny: Claude Code & the Future of Engineering | Acquired Unplugged · ≈1:44
“for us, we exist to research AI safety. That's the reason that we exist.”
“the way you study safety you know it has all these kind of layers including setting the model in the wild and uh kind of what's a useful way to do that it's coding because that's the way the model interacts with the world”
“English is there's an infinite number of correct beautiful poems, but there's a very constrained uh there's a non- infinite number of ways that you can write the correct code to solve a certain problem.”
What’s in it
- Why Anthropic frames coding as an AI-safety research vehicle
- How Claude Code fits into studying models 'in the wild'
- Why code is a clean testbed: it compiles or it doesn't
Clip transcript
we were using ideides still at at the time. Um, but for Anthropic as a company, we've always cared about coding. Um, we've always cared about tool use. We've always cared about computer use. This has always been kind of the core research direction because, uh, you know, for us, we exist to research AI safety. That's the reason that we exist. Like, if you ask like a random person at anthropic, you just like pull them over in the hallway and ask like, why are you here? They're going to say AI safety. That's the reason every single person is at the company that's they deeply believe including me this is just the single most important problem to solve and there's kind of a bunch of ways to solve it. You can look at kind of me mechanistic interpretability. You can do alignment work. There's all sorts of ways to study the model in a petri dish but fundamentally you have to study it in the wild to see what it does once it's safe in all these other ways. And so that's kind of where Claude Code comes in. But you know for anthropic as a company the direction has always been safety and the way you study safety you know it has all these kind of layers including setting the model in the wild and uh kind of what's a useful way to do that it's coding because that's the way the model interacts with the world and so if you want to study kind of various sorts of model misalignment you want to uh make it useful enough that people use it so that you can study it coding is just a very obvious application. Um so that's just from the beginning that's that's been the focus. It's just such an unbelievably clean universe because you've got all this training data. It either works or doesn't work and like depending on the language, it either compiles or doesn't compile or it either, you know, there's sort of a very clear like pass fail ability to test it. Um, and it's a very constrained universe of correct solutions. English is there's an infinite number of correct beautiful poems, but there's a very constrained uh there's a non- infinite number of ways that you can write the correct code to solve a certain problem. Yeah. And it's there's sort of a an elegance to using it as your petri dish given that. Yeah, that's right. There was uh
Comments
Sign in to comment.
Loading comments…