Clip transcript
And the last lesson is to benchmark agents on your code base. So, um the way we do this, we select pull requests that represent great engineering work. It could be agent created, could be human created, could be a hybrid, doesn't matter. You pick the agents you want to use and benchmark. And then you get a quality versus cost and time breakdown on your code base. Now, why do you want to do this? There's There's many reasons, but one is that if you're kind of going off the public benchmarks, we bench or terminal bench or other stuff like those tasks may have absolutely nothing to do with your task. Like swe bench is all in Python, we're Ruby on Rails. It is not the case that the benchmarks are identical for them. There's trends that do compare, but the results can be very, very different. And I'm going to swap over to my browser one more time here. So, this is These are results on our code base of all these different harnesses. This one here is quality versus cost. This one here is quality versus time. I'll start with this one. You can see some trends here. Right, you can see that the Anthropic agents have just been consistently getting better, but not really any faster. The Codex agents and cursor are actually pretty fast and quite good. The open stuff has been getting better and better over time, but they're kind of slow. This is for our code base again. I'm not trying to make any general claims here. By cost, the Anthropic stuff is clearly just so much more expensive for us. And the Codex stuff has been cheaper for us. And so this causes us to change our behavior. We still use the different models. There's different use cases for them. We like the variety. We still use all these things. But when we saw these results, they kind of matched our vibe check. We wanted to kind of like have hard data, too. We switched our default to Codex at that time. And Fiable came out, and it was great. Kind of switched our default to that for like the few days we had it, and then it went away and switched back to Codex. But the most important thing is like because we're agnostic, like none of that had any meaningful disruption on our work. Like we're able to just kind of switch back and forth really easily. So, the next day something new comes out, see if it's good, and go.