What would pretraining with zero real data look like? Can a randomly initialized model learn to generate all of its training data entirely through self-play? Find out below! :) https://t.co/sSybpKwUUL
New work co-led by @AdityaCowsik, @KfirDolev, and @michaelyli_ with co-authors @gbruno_dl @ANourya @noahdgoodman, and @YoavLevine.
pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
Why it matters
If a model can generate useful pretrainingThe first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.Full definition → data purely through self-play, it weakens the assumption that scaling requires ever-larger real-world corpora, directly relevant to future data-bottleneck concerns.