← All IntelClip / AI ToolsRephrasing: synthetic data that can beat its own teacher model
From Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI · ≈12:14
“Um we take an approach to synthetic data that we call rephrasing. This is something that our team pioneered several years ago and has now become um basically table stakes for building a very strong model.”
AI Engineer
“Um so to give an example, we might take a document like this about a corporate takeover and we might convert that into one of hundreds of templates. One of which might be a series of true false questions.”
AI Engineer
“Number one, because all the information is coming from the document on the left, you don't have any issue with model collapse.”
AI Engineer
“Um and you can actually train models um that are much better than the rephrasing model because the rephrasing model doesn't actually have to teach and understand all the concepts. All it needs to do is transform the left document into a true false questions accurately which is a much easier task.”
AI Engineer
“All documents are not created equal for rephrasing. If you just pick random sets of documents to rephrase, you will not get a great result.”
AI Engineer
What’s in it
- Explains why grounding synthetic generation in a source document avoids model collapse, and warns that rephrasing randomly chosen documents rather than high-quality ones destroys the gain.
Clip transcript
non-English data also helps to benefit English data um performance um okay And let me talk a little bit about synthetic data. Um we take an approach to synthetic data that we call rephrasing. This is something that our team pioneered several years ago and has now become um basically table stakes for building a very strong model. Um in anyway I think we'll hear a lot about the synthetic data in various forms uh throughout the day. Um but fundamentally with beyond web our goals is how can we define a synthetic data platform that works extremely well um and can be applied to anyone's proprietary doc data and documents. Fundamentally we want to help folks build models that wouldn't be able to do so otherwise. And that's what where data quality can make an absolute difference. Um so to give an example, we might take a document like this about a corporate takeover and we might convert that into one of hundreds of templates. One of which might be a series of true false questions. Um by doing this, there's a couple things that are really great. Number one, because all the information is coming from the document on the left, you don't have any issue with model collapse. Um and you can actually train models um that are much better than the rephrasing model because the rephrasing model doesn't actually have to teach and understand all the concepts. All it needs to do is transform the left document into a true false questions accurately which is a much easier task. And then you do this into many many different formats um throughout data. This effectively increases diversity and it makes it so you learn a lot more from the highest quality data points. One thing that's really critical here, what do you rephrase? All documents are not created equal for rephrasing. If you just pick random sets of documents to rephrase, you will not get a great result. Um but if you find the high quality documents and rephrase them, it can make a big difference. Um and we've seen if you compare this to lots of other public synthetic corpora, we can get much better performance much faster. Um and critically this can be applied to any any proprietary data um in one of our customers own environments.
Comments
Sign in to comment.
Loading comments…