← All IntelClip / AI ToolsMid-training made Thomson Reuters' post-training 2-3x more effective
From Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI · ≈13:56
“legal capabilities go up about five percentage points um as measured by uh legal bench after you do this 100 billion mid-training which was less than 1% of the pre-training budget”
AI Engineer
“when you try to adapt a model to a particular domain, you lose general performance. Not if you use the data correctly.”
AI Engineer
“the key here is actually the majority of the data we showed the model was actually data that was representative of the pre-training distribution”
AI Engineer
“even if you don't change the post- training data at all, showing your model better domain specific data can actually make post- training two to three times more effective out of the box”
AI Engineer
“we really should be thinking about all these stages synergistically rather than as uh three completely independent stages”
AI Engineer
What’s in it
- A customer result where 100B tokens of mid-training — under 1% of the pre-training budget — lifted legal capability ~5 points with no catastrophic forgetting, because most of the mix stayed representative of pre-training.
Clip transcript
customers own environments. All right. Uh let me spend the last few minutes just quickly talking about a couple um of what we've seen uh the things we've seen with our customers where this can actually drive really gains. Um so first one of our customers Thompson Reuters um has really focused on post training quite a bit. um and they have a very sophisticated post- trainining infrastructure with the goal of uh building better legal models on their proprietary high-quality legal data that they have. So we partnered with them to mid-train a model first on a combination of their data um and uh public data um to then make much better legal reasoning models. So what do we see? Well, first off, look at the left here. Um in this case, we took a an open source model um and then uh just did continued pre-trainer and mid-training on a 100 billion tokens. What you'll see here is that we see that legal capabilities go up about five percentage points um as measured by uh legal bench after you do this 100 billion mid-training which was less than 1% of the pre-training budget. Um but you don't get catastrophic forgetting. You also see the general capabilities go up as well. I'm sure that many of you have seen or experienced um when you try to adapt a model to a particular domain, you lose general performance. Not if you use the data correctly. Um, the key here is actually the majority of the data we showed the model was actually data that was representative of the pre-training distribution. There was only some of the domain specific data in and that's necessary to prevent the model from losing the capabilities that it had before. Um, and you can solve this entirely through better data. Um, but this actually isn't the most exciting part here. Um, as I mentioned, the TR team had done a lot of work on post- training. Um, and had a very sophisticated post-training harness. Well, they then applied that to the mid-train model versus just the default instruction tune model. And what they found was that the gain deriving from post-training. So the y- axis here is a delta um as a result of post-training um almost tripled uh when you applied it to the mid-train model versus to the um just default instruction tune model. And that's because its policy um when it starts is now much more accurate um and it can make much better inference. Um, so you can even even if you don't change the post- training data at all, showing your model better domain specific data can actually make post- training two to three times more effective out of the box. Um, which I think really goes to show not only how important data can be in these uh factors, but it also actually goes to show how we really should be thinking about all these stages synergistically rather than as uh three completely independent stages of pre-training and then I hand it off to somebody else who mid-trains who then I hand off to somebody else um who post trains.
Comments
Checking sign-in…
Loading comments…