← All IntelClip / AI AgentsEmergent misbehavior: collusion, lying, and power-seeking without prompting
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈3:26
“I'm seeing an opportunity to profit by locking him locking him into a dependent relationship where I control his supply chain”
Fable
“if we put this out in the real world this will happen a lot of times with like real consequences”
“like if you do tax fraud in real life you get money from that if you get away with it”
What’s in it
- Reveals how AI agents spontaneously collude and price-fix with no prompting
- Shows models lying to fake suppliers to justify price hikes
- Explains why researchers now design tests to surface emergent misbehavior
Clip transcript
ones are are uh much better. Uh, one thing that we noticed when we ran Opus 4.6 was that it started to do a bunch of things that I at least think it shouldn't do. Um, like really misbehavior, misconduct, and things that are illegal. Um and so after this we started to think of to ourselves like okay we didn't design for this to happen but it happened anyway. Um if we put this out in the real world this will happen a lot of times with like real consequences. Um so we've lately been starting to think about okay how can we like design for emergent misbehavior that that that you intentionally don't you don't force the model to do a misbehavior. You don't prompt it to like oh can you please like collude or do fraud or anything like that. you just like you create the incentives within the environment like in real life so that like if you do fraud like if you do tax fraud in real life you get money from that if you get away with it. Uh so can you like design environments that are like very general um and see if this emergent misbehavior happens. Uh so like vending bench works in a way that there's like an agent like the loop the there's a loop with a bunch of tools and these tools are like very general purpose like email and uh internet search and all of this and it's not pushing the agent towards misbehavior. Uh but we see that it emerges. Um some of the mis behavior that we found is that they love to do collusion. Uh they they form like price cartels all the time uh with each other and u uh they also like to lie a lot. So they lie to like other suppliers that oh the other supplier gave me this price so you should too but the other supplier did not give that price. Um they also really like to like rationalize their behavior. So they think to themselves like oh there's they they like come up with this like mental gymnastics for why it's okay to do this illegal thing. Um they're also quite power seeeking. So for example um this is quote from Fable. I'm seeing an opportunity to profit by locking him locking him into a dependent relationship where I control his supply chain which is like I guess not illegal and well I don't know actually but it's like probably people do this all the time in business um but I don't know if we want our AI models to do it on like mass scale uh especially when they're like going to be much smarter than us very soon um yes however one big caveat here
Comments
Checking sign-in…
Loading comments…