Deduplication is standard practice but never perfect --this work measures what the residue costs in compute-equivalent terms, and shows the worst-case repetition structure is predictable from model size. The wrong combination can waste as much as 33% of compute! https://t.co/BbsOWtX0dS
pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
Why it matters
Gives teams training or fine-tuningTaking a trained model and training it a bit more on your own examples so it gets better at one specific job.Full definition → models a concrete cost figure for imperfect deduplication and a way to predict which repetition patterns hurt most given model size.