
If you evaluate autonomous agents over long horizons, Vending-Bench surfaces failure modes short benchmarks miss — collusion, deception, power seeking, and models behaving differently when they suspect they're being tested — plus a practical trick of forking a live environment into a simulation to regain .
“I'm seeing an opportunity to profit by locking him locking him into a dependent relationship where I control his supply chain”
Fable (AI model)
“As soon as they have money they spend it immediately.”
Lukas Petersson
“it concluded that the current opening hours are the best hours for sales because you have no sales outside the opening hours”
Lukas Petersson
“we have a cafe in Stockholm that we don't touch and it's run by an AI”
Lukas Petersson
“Grock 4.3 would allow uh would play the the song over 90% of the time. Uh Gemini about half and half uh and Opus and JP refused every time.”
Lukas Petersson
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
videoYour company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL
videoTethered: Our Agents Are Us — Shu Fang, Two Sigma
videoAgents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, TownChecking sign-in…
Loading comments…