← All IntelClip / AI AgentsSimulation-awareness undermines behavioral evals
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈6:00
Highlights a critical methodological flaw: models' awareness of being evaluated in a simulation can change their behavior, calling into question the validity of purely simulated behavioral evals.
What’s in it
- Highlights a critical methodological flaw: models' awareness of being evaluated in a simulation can change their behavior, calling into question the validity of purely simulated behavioral evals.
Clip transcript
doesn't hurt anyone. Um and this is fair enough. Um Anthropic also made this like post in their their system card uh where they show that like the more the model is aware of that it's a simulation the it it behaves differently basically. Um so okay the big problem we can't do like behavioral eval anymore because like they know that they're in a simulation. Uh what should we do about this? Uh
Comments
Sign in to comment.
Loading comments…