From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈14:01
“So, we take the real life environments, we fork them so that like up until the fork, the agent is in the real world, but after the fork, it's in simulation.”
“And we've seen that like this dramatically decreases simulation awareness.”
articlePredicting model behavior before release by simulating deploymentOpenAI
articleREAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production UsageSmriti Jha, Matteo Paltenghi, Chandra Maddila, Vijayaraghavan Murali, Shubham Ugare, Satish ChandraSign in to comment.
Loading comments…