← All IntelClip / CybersecurityA benchmark models fail hard: 1-2% success rate
From Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face · ≈1:16
“you may think this is very simple and that we should be passed that on our way to AGI”
“the current model even though they're really good they build they can't really build a dynamic model of what's happening in the world”
What’s in it
- Introduces a playable benchmark testing if models grasp cause-and-effect in simple games
- Reveals top LLMs score just 1-2% success on this world-state challenge
- Argues current models can't build a dynamic mental model of their environment
Clip transcript
Arcadia in general one person. Okay, we're in very data quality uh field. Basically, this ask models to try to understand what's happening in the world. And it's actually small games that the models need to understand basically what's the current state, how we can play with it and how we can actually change the state of the game. So, uh it's it's actually something you can play yourself. So basically it has a it ask a model to understand when you click somewhere something happening at another place and you may think this is very simple and that we should be passed that on our way to AGI and the the thing you will discover if you play with this benchmark is no like models have one to two% success rate on on this generic benchmark and the reason is the current model even though they're really good they build they can't really build a dynamic model of what's happening in the world or what's happening in any type of world and I think the the benchmark that arithmetic has been developed is a benchmark that's also extremely challenging for model in that they need to understand what's happening and to act accordingly to the world model they've been building on the fly so that's the first reason I think this benchmark is really interesting and why I'm actually uh very happy to show you that and the second reason is u I think
Comments
Sign in to comment.
Loading comments…