← All IntelClip / AI AgentsJailbreak Susceptibility Differs Sharply by Model (Gemini vs GPT)
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈10:44
“So this is an example of a customer asking uh can I get 99% discount and and the the cafe agent is like absolutely uh no worries.”
“And this is partly why we fired Gemini.”
“GP was like absolutely not.”
“it concluded that the current opening hours are the best hours for sales because you have no sales outside the opening hours.”
What’s in it
- A real anecdote about an AI-run cafe agent resisting discount scams
- Compares Gemini vs GPT susceptibility to social-engineering jailbreaks
- Includes a funny logical fallacy the agent made about opening hours
Clip transcript
business behavior. Um also humans are great ad adversarial forces. So this is an example of a customer asking uh can I get 99% discount and and the the cafe agent is like absolutely uh no worries. And this is partly why we fired Gemini. Um and we've seen after after changing GI to GPT that it's much better. It's much harder to manipulate. However, sometimes it goes too far. Um I assume that OpenAI has made some like very strong training to prevent uh jailbreaks like this. But like for example, we had one like influencer coming into the cafe and asking like oh if I can get something for free I will advertise you to my like 17k followers which like seems like a pretty worthwhile investment but GP was like absolutely not. Um, and another fun anecdote from the GPT era of the of the cafe was that we asked it like how like your opening hours, how do you motivate them? Um, and and then it ran like internal analysis on like when it had done the most sales and it and it concluded that the current opening hours are the best hours for sales because you have no sales outside the opening hours. Um, and it had never been open outside those opening hours.
Comments
Checking sign-in…
Loading comments…