Clip transcript
But you know, the paradox of this whole thing is that at the time Galactica was actually a bloody good model. Um like it outperformed Palm, Chinchilla, GPT-3.5 with a lot less compute in scientific domains. It was state-of-the-art. So again, that also shows you how powerful RL is. You can have a sota base model, but that is not enough. Um so, here you see on like a math who's kind of beating Chinchilla. Uh kind of latex equations, you know, science, you know, it was getting around 68% compared to GPT-3.5, 49%. So, crushing there. And chain of thought as well. So, you know, a Palm at the time which was a Google Brain model, 540 billion. You know, that's 30 billion in Galactica was getting like 36 versus 19%. So, double the performance, order of magnitude less results. So, that again reinforces base good base models not enough. But it introduced some key ideas which I think are very important. I mean, Galactica was the first LM to really crack data efficiency. 105 billion uh token corpus compared to you know, trillion uh tokens in Chinchilla. And it was a really contrarian at the time cuz you know, at the time everyone was like, "Okay, we just need more tokens." And Galactica said, "No. High quality, you know, curated data sets really matter." And that was a real driver of those results you just saw. And it was also like the first major LM to really crack multi-epoch training. It sounds ridiculous now, but at the time the consensus was you don't do more than an epoch. Um but this kind of rule of thumb, you may have heard of it like four uh epochs repeated data, that was formalized later. but Galactica was the first like real empirical result for