← All IntelClip / AI ToolsDe-risking a hero run with 50-100x less compute
From Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI · ≈10:06
Two 1T-token dense models drew a line that ran straight through a 17T-token hypersparse run at 50x the compute — meaning small curated runs can predict the big one before you pay for it.
What’s in it
- Two 1T-token dense models drew a line that ran straight through a 17T-token hypersparse run at 50x the compute — meaning small curated runs can predict the big one before you pay for it.
Clip transcript
You can see again that we're well off the predto frontier with a couple things I really want to highlight. First off we only use 8% of the data here as multilingual tokens. Um, so most languages actually only had um at max 6 billion tokens here. So these are not massive amounts of data uh in the non-English languages that are going in here. Um you can again see we get the same sort of compute multiplier effect. We're a little better than Quen 3 um while while having roughly 8x less compute budget um here. So you can make a huge improvement. One last thing I want to show here is that if you look at the two blue points on the upper left here um those are both dense llama style models trained for a trillion tokens um on curated data. The point on the lower right here I'll come back to but is a model trained by one of our customers rci trendy large that was trained on 17 trillion tokens and is a hypersparse. Um and what you can see is that if you take the line defined by the two smaller models, um it goes mostly right through that blue star uh which is training large despite it being trained um with 50x more training compute. So if you use your data correctly and you simulate token scarcity appropriately, you can also get um very predictable scaling to much larger models and you can derisk a run with 50 or 100 times less compute effectively um before you actually go and scale up the hero run um and find that maybe it doesn't end up where you
Comments
Checking sign-in…
Loading comments…