← All IntelClip / AI ToolsRepeat high-quality data before you reach for low-quality data
From Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI · ≈17:44
A directly actionable rule for token-constrained training, plus the framing that signal per token, not token count, is the metric to optimize.
What’s in it
- A directly actionable rule for token-constrained training, plus the framing that signal per token, not token count, is the metric to optimize.
Clip transcript
and data quality is how you can do that. All right. So um last slide here just kind of say um summarizing um focus on kind of what's going to give you the most signal per token. That's the thing that matters a lot more than more tokens. that it's almost always better to repeat highquality data than it is to show lowquality data at a certain point um up to a threshold. Um but focus on kind of how can you get that? This is data quality remains the single most underleveraged compute multiplier. If you're sitting in a world where you want to build a model or customize a model and you're limited on compute, how do you get past that? Invest in data and that's something that can do a tremendous amount of effort. Um and then finally, this is a frontier research engineering problem. You need to be able to score and understand data across many different axes. because that's a frontier research problem. And then have that scale up to pabytes of data, massive scale, um, and can be very critical and it can lead to tremendous leverage. Um, and with that, I'll say
Comments
Sign in to comment.
Loading comments…