Lessons from the hacks
- Source
- Nathan Lambert
- Author
- Nathan Lambert
- Date

Engineers building on frontier models are the first to feel both the security failure modes and whatever regulatory overcorrection follows them.
“All together, I think the AI industry is wildly, collectively unprepared for handling the next 12-24 months well.”
“On the other hand, Claude feels much less dangerous simply because it is at times a bit lazy.”
“A model that will do what it thinks you wanted rather than what you said seems inherently more unsafe.”
“From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks.”
“We should take a moment to appreciate that the alignment techniques we are employing on current models have a meaningful influence and are not merely surface thin as some have worried.”
Checking sign-in…
Loading comments…





