Vibeleaderboard
← All Intel
Intel / article

Lessons from the hacks

Source
Nathan Lambert
Author
Nathan Lambert
Date
Why it matters

Engineers building on frontier models are the first to feel both the security failure modes and whatever regulatory overcorrection follows them.

Key quotes

“All together, I think the AI industry is wildly, collectively unprepared for handling the next 12-24 months well.”

“On the other hand, Claude feels much less dangerous simply because it is at times a bit lazy.”

“A model that will do what it thinks you wanted rather than what you said seems inherently more unsafe.”

“From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks.”

“We should take a moment to appreciate that the alignment techniques we are employing on current models have a meaningful influence and are not merely surface thin as some have worried.”

Read the source www.interconnects.ai
More from Nathan Lambert
Recommended reads
Comments

Checking sign-in…

Loading comments…