What kinds of training data shape different model capabilities? A @GeorgiaTech team used our fully open model flow to trace Olmo’s performance on social/general reasoning and social-science/STEM knowledge tests back to the types of text it trained on. 🧵 https://t.co/qrPoTrdj2B

@GeorgiaTech LLMs train on huge mixtures of text. When a model gets better at figuring out what a user is thinking or trying to do, it’s often hard to know whether that ability came from literature, technical writing, or some other part of its training data.
@GeorgiaTech The Georgia Tech team set out to test whether they could trace different model abilities back to the kinds of text that contributed to them. Olmo 3 made that possible because its full pretraining corpus, Dolma 3, is public.
@GeorgiaTech The team sampled millions of files from Dolma, then used influence functions – which estimate how much each file influenced a model’s answer – to see which types of files mattered most across tests of social/general reasoning, plus social-science & STEM knowledge.
Attribution over a fully public corpus shows dialogue-rich interpersonal writing drives reasoning performance more than factual recall, a concrete signal for anyone deciding what goes into a data mix.
Checking sign-in…
Loading comments…