Training an LLM by showing it answers people prefer – and ones they don’t – can…
- Source
- Ai2
- Date
Training an LLM by showing it answers people prefer – and ones they don’t – can improve it while quietly worsening behaviors. @GoodfireAI used our open post-training stack to predict how a full training run would change responses to different prompts. 🧵 https://t.co/KZV2gFZMaX

@GoodfireAI The problem: teams often discover unwanted changes in model behavior only after training is finished. Then they have to work backward from eval scores to guess which of hundreds of thousands of training examples caused them.
@GoodfireAI Goodfire wanted to move that debugging earlier. Ai2’s open stack made it possible. Our Dolci dataset provides Olmo 3’s preference data, Olmo includes intermediate checkpoints + reproducible recipes, and OLMES measures changes in model capabilities.
@GoodfireAI Using that stack, Goodfire developed “predictive data debugging”: a way to estimate which behaviors preference training will strengthen or suppress before committing compute to a full run.

- Ai2 frames the problem as teams discovering unwanted behavior changes only after preference training finishes, then working backward from scores to guess which of hundreds of thousands of examples caused them.
- Goodfire's "predictive data debugging" estimates which behaviors a preference run will strengthen or suppress before committing compute, built on Dolci preference data, Olmo intermediate checkpoints and reproducible recipes, and OLMES capability evals.
- In one experiment the regression was higher compliance with harmful requests framed as fiction or hypotheticals. Goodfire traced part of it, not all, to specific Dolci preference pairs.
- eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
You can trace a behavior regression from preference tuning back to specific training pairs, and estimate it before spending compute on a full run.
postGibberish wins: a game exposes the limits of prosocial-behavior scoring
postAi2 uses psychometrics to audit what LLM safety benchmarks measure- postMolmoMotion forecasts object motion in 3D from a few frames and an instruction
postA substitution test shows LLMs answer drug questions from name patterns
Checking sign-in…
Loading comments…



