Clip transcript
the coverage and the expert agreement. But I think what we wanted to close off with today is why a lot of this stuff matters, right? We spent a lot of time earlier in this presentation defining what long horizon means. And a huge reason we did that is because we feel like a lot of the literature and data, a lot of the literature shows that a lot of the data being produced right now and being used to train and evaluate models is actually flawed. So we present three major benchmarks in the area of finance predominantly. And so this is GDP valer toolbench and Apex agents. There's a couple of notable issues here. First, if you look at the average human hours per task, based on what Meter has defined for a lot of the leading frontier models, a lot of these different average human hours per task fall far below that and so they wouldn't actually be considered long horizon tasks. The second notable issue here, right, is that we see that these benchmarks are already reasonably saturated and we think this is a downstream effect of the average human hours per task. So, it's really important to look at the metrics that are being used here. If you look at, you know, the Apex agents IB section of this benchmark that they put out, pass at one effectively means that for like 57% of cases, the tasks are 100% solved. That is effectively telling us that like there's a large part of these tasks that models are solving similar to what we've seen already. But I think a third key important part here is the breath. For each of these different uh benchmarks, particularly GDP val, they have a very narrow set of Excel tasks that they consider for finance. And for Apex agents, they're largely focused on IB. What this means is that a lot of these more important areas for learnability like, you know, credit, debt, risk in the domain of finance don't really get covered. And then I think lastly, I'll I'll I'll note there the reward signal as Gver mentioned is really important. And in regards to the reward signal here, we'll we'll notice that there's like really, you know, if you look at if you look at what you need for a rubric, you need very granular, detailed reward signal. You need, you know, we we have 20 different subriteria and 10 different subriteria per criteria. So, I think there's a lot of room that's left when you read these benchmarks into how granular reward signal they're giving, which is really important for being able to go ahead and train your models. With that, I think I wanted to round off with a couple of stats about the data we produced. Here we, you know, you can look at some statistics for our finance data. We can see that the human time to complete one task on average is 15 hours over a 50 task sample set. Furthermore, it takes models a pretty long time to work through these tasks. And after all of that, across all the domains we care about within finance, for example, they still struggle significantly. And so here we provide mean five notably different than, you know, all of these previous scores uh we see here. So,