Clip transcript
up the post-training effort for Llama. But, I had like a personal obsession, which was like reasoning. And I had like a really simple idea at the time, which is, "What if we applied kind of reinforcement learning pressure to this like kind of thinking tags as work?" Like, what if we just optimize the thing in between the thinking? And if that sounds familiar, then this is like kind of what Deep Seek 01 ended up doing 2 years later. But, there's a key difference. So, at the time we only had Llama 2 base models, terrible mathematics corpus, terrible results on math. And, you know, the context window, you know, we're all like, you know, very context-rich now, 1 million tokens. You know, the back in the day it was 4,000. It wasn't too fun. But, we still had a recipe at the time, and this is unpublished, but it was a really good for the meta. So, RSP was this. Number one, continue pre-training of Llama 2 towards mathematics and science data. So, that's the first thing. Llama 2, math corpus, so let's fix that. Number two, PPO with verifiable rewards. But, notice this isn't GRPO, right? So, we had a time and a strong outcome reward model to initialize the value model. That was a key thing, lots of data on value models at the time. And internally at the time, this kind of recipe led to state-of-the-art results on math and reasoning. So, we were kind of like, "Wow, this is like really shows the power of having the right objective." But, the really fascinating thing is like we had great results, but we didn't have like inference time scaling. We didn't have this reflective behavior that became like the hallmark of R1 and O1. You know, back weights, you know, back tracking, all this kind of stuff. So, it begs like the question like, "Why? Well, why didn't we have that moment?" And we got an answer around like 2 years later. So, there's a couple of like things going on here, but essentially better base models were the thing that really got RL cooking. And when DeepMind came out, I was kind of shocked at the time. I was like, "Holy we just like tried the same thing. We didn't have this. What's going on here?" And in a weird kind of way, the real lesson was it was just like the bitter lesson, like the most purest form of bitter lesson possible. Like, better base models, more RL computes, bigger context windows, and that's all you need for this kind of emergent behavior. It also like says something like quite important about the sociology of like research because the fact that OpenAI had this model, you know, GPT-4 level model before anyone else, it allowed them to see further, right? So, that's a really interesting point. Like, the the age of scaling means that if you have certain prerequisites in place, you become smarter, you see further, you see more ideas. So, really interesting point.