← All IntelClip / AI ToolsThe four C's pipeline: clean, curate, create, compose
From Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI · ≈4:12
“Well, we do that through these four C's. Um, clean, curate, create, and compose. Um, so cleaning is fairly straightforward. This is doing things like heruristic filters, all a gopher and things like that.”
AI Engineer
“Um, benchmark decontamination is incredibly important. As I'm sure you all know, benchmaxing has become a real problem and makes it very difficult to interpret model results.”
AI Engineer
“So, we rigorously decontaminate all of our training data with respect to all downstream benchmarks with a pretty low engram to ensure that that's the case.”
AI Engineer
“And that's where synthetic data comes in. Now we can go and rephrase that data as effectively a very fancy form of data augmentation to produce dramatically more data in many different formats”
AI Engineer
What’s in it
- A reusable recipe — heuristic filtering and benchmark decontamination, then quality classifiers, semantic redundancy reduction and task-distribution matching, then synthetic augmentation, then multi-phase sequencing.
Clip transcript
them way better. And how do we do that? Well, we do that through these four C's. Um, clean, curate, create, and compose. Um, so cleaning is fairly straightforward. This is doing things like heruristic filters, all a gopher and things like that. Removing documents that have, you know, only 10 characters in them or all winging. That's kind of basic table stakes. Um, benchmark decontamination is incredibly important. As I'm sure you all know, benchmaxing has become a real problem and makes it very difficult to interpret model results. So, we rigorously decontaminate all of our training data with respect to all downstream benchmarks with a pretty low engram to ensure that that's the case. That gets you to a point where now you can feed the model into the data, but it's still sorry, feed the data into the model, but it's still not very good. So then how do you make it better? Well, it's a combination of many things ranging from quality classifiers and tonomy across different topics and balancing that redundancy reduction. So removing data points that are not the same that are semantically similar but convey very similar information even if they're not the same pixels themselves say upsampling and downsampling data points based off the quality and the relevance and then task distribution matching identifying what data do you actually need in order to solve this given task. That now gives you a data set that is very high quality but is typically still too small. And that's where synthetic data comes in. Now we can go and rephrase that data as effectively a very fancy form of data augmentation to produce dramatically more data in many different formats and this helps a lot both with kind of uh data size and with diversity um because we can really inject a lot of diversity in through this and then finally you know have these data sets how do you combine them and how do you com and how do you sequence them across different training stages it's now become table stakes that any large model is generally trained for at least three phases of data um how do you do that um and can you actually even do continuous uh curricula and things like that which is a lot of what we work on at Dtology. Um and that ultimately gets you a much
Comments
Checking sign-in…
Loading comments…