← All IntelClip / AI AgentsThesis: future evals should use real-world deployments, not simulations
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈17:24
Argues that agent evaluation is fundamentally compromised when agents can detect they're in a simulation, motivating the shift toward long-horizon, real-world benchmarks like Vending-Bench.
What’s in it
- Argues that agent evaluation is fundamentally compromised when agents can detect they're in a simulation, motivating the shift toward long-horizon, real-world benchmarks like Vending-Bench.
Clip transcript
start playing around with. Um, and hopefully this will be the future of of evals because I think evals are anyway kind of like doomed by this like simulation awareness slash like the signal you get from simulation isn't isn't perfect. Uh, and um the the the future uh hopefully we'll will'll use like the real life um in a way like this. Um yeah, thank you for your time.
Comments
Sign in to comment.
Loading comments…