Clip transcript
run. Um and I want to end on um some early results from uh this new model um and demonstrating how it performs against some open weight models and also against our previous models. So, first I want to caveat this with these are base model evals, right? They are partly indicative of how the final model will look, but also not perfectly, right? Um there's still post training happening. Uh not all of these will translate one-to-one to the final model. Um but if we look at them uh specifically on the coding part of of the E valves, so for instance multiple E love code bench, big code bench, uh Laguna S is not only stronger than access to two, which is our previous uh smaller model that performed very well, but also than the much larger M dot one. Um and it's also much better than GLM 4.5 air, which is admittedly a bit older, and then we turn 360, which is quite recent, and then deep seek V4 flash max, which is quite recent and a fair bit larger. Um we can see it's competitive on big bench hard for instance. It doesn't achieve the the top E valve results compared to these models, but it's quite close. Um we also see it's quite close on E valve plus, and uh quite importantly for us, it does very well on sweep bench agent less multilingual, um which we use to sort of proxy agentic performance during pre-training. Um and in that case it performs much better than all the other all other models we tested here. Um I also want to point out that of course there are like it's not the strongest model in the world, right? Like for instance MMLU pro knowledge benchmark is something we don't care about that much compared to coding because we want to build the strongest agentic coding models. Um so here like compared to Nema tron and deep seek uh we have to say that they perform much better, and this mainly comes down due to data, right? It's a data gap um that we could plug if we wanted to. Um but I think the the point is all of the things we we found before were included in the recipe. The recipe held, it scaled, and we will continue scaling it from here. So this model will also be available sometime in the future relatively soon. Um again open weights, so all of you can download it and use it for free.