Vibeleaderboard
← All Intel
Intel / post

Repeated Pretraining Data Can Waste a Third of Compute

Source
Stanford AI Lab
Date
Stanford AI Lab@StanfordAILab

Deduplication is standard practice but never perfect --this work measures what the residue costs in compute-equivalent terms, and shows the worst-case repetition structure is predictable from model size. The wrong combination can waste as much as 33% of compute! https://t.co/BbsOWtX0dS

Terms in this piece · Glossary
  • pretrainingThe first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
  • fine-tuningTaking a trained model and training it a bit more on your own examples so it gets better at one specific job.
Why it matters

Gives teams training or models a concrete cost figure for imperfect deduplication and a way to predict which repetition patterns hurt most given model size.

More from Stanford AI Lab
Recommended reads
Comments

Checking sign-in…

Loading comments…