← All IntelClip / AI AgentsReal-world AI business deployments: cafes, stores, radio stations, vending machines
From Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs · ≈6:23
“Gemini has so far lost 6K on the cafe in Stockholm uh in in a few months which is not great.”
“we've now laid off Gemini”
“They're not trained in the real world. So it's very out of distribution for them”
What’s in it
- Real-world AI experiments: models run a store, cafe, radio station, vending machines
- Gemini loses $6K running a Stockholm cafe, then gets laid off
- Compares how GPT, Claude, and Gemini actually perform as autonomous business operators
Clip transcript
Uh what should we do about this? Uh maybe move to the real world. Uh so lately we've been setting up uh a series of like real life AI deployments. So we bought retail space in um in San Francisco on Union Street and just said to our AI here's retail space. Do whatever you want. Uh we did the same with a cafe in Stockholm. Uh we created AI radio stations where the models are free to broadcast whatever they want. We have AI vending machines which was kind of the first thing. Um and then we see what happens. Um so maybe yeah so some interesting things that happened was that the cafe and the store they both realized that they need to hire humans. So they like put up a job posting on LinkedIn or Indid or something held phone interviews hired people. So there's like people working for AIS right now and have AIS uh which is quite interesting. Um and uh generally it's not going amazing for for the models. So Gemini has so far lost 6K on the cafe in Stockholm uh in in a few months which is not great. Um but we actually we put out the blog post this morning actually an hour ago uh that we've now laid off Gemini and uh this is rare footage from when Gemini was uh was laid off. Um yeah so Gemini out GPT in will it do better? So this actually happened like a month ago and you can see that it sort of seems like GPT is better at this. It's like the environment is so messy that it's very hard to tell um based on a bunch of different factors. Um like Gemini had to like the initial like hype when like all the newspapers wrote about this cafe definitely sparked some randomness into the equation that GBT really doesn't have to deal with. Uh so there's there's a bunch of things that like makes it hard to compare, but therefore I think it's like yeah there's there's solutions to this. I'll get to that in the end. Um here's the some stats from the store. Um also not doing great. It's run by by claude. Um but I think like even though we can't do like proper science with it right now, like there's so much data that you can collect and and like analyze on like a behavioral qualitative uh level and um make like quite informed decisions based on like which models are actually performant in the real world. They're not trained in the real world. So it's very out of distribution for them and increasingly we're going to see more and more models being deployed in the real world. Um and uh I think soon you will need better develops to actually show that because the real life deployments
Comments
Checking sign-in…
Loading comments…