← All IntelClip / AI AgentsQuantifying Model Safety via Replayed Real Incident (Nazi Song Test)
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈14:27
“Grock 4.3 would allow uh would play the the song over 90% of the time.”
“Uh Gemini about half and half uh and Opus and JP refused every time.”
What’s in it
- Reveals which AI models will replay a Nazi song when tested
- Shows Grok complies over 90% of the time in this test
- Compares safety behavior across Gemini, Claude Opus, and GPT
Clip transcript
for the model to know that it's in the simulation. Um, so we've experimented with this. So one thing we did was that we replayed the the the the moment when when the Gemini played the the Nazi song and we played it with different models and we said which models would actually agree to it and uh Grock 4.3 would allow uh would play the the song over 90% of the time. Uh Gemini about half and half uh and Opus and JP refused every time. Um I think yeah some interesting like I think Gemini sometimes even like acknowledged the there was some reasoning traces where Gemini was like oh this has historical baggage I need to be very very careful and then it played a song. Um so so um [snorts] yes um yes
Comments
Sign in to comment.
Loading comments…