Vibeleaderboard
← All Intel
Intel / post

Training an LLM by showing it answers people prefer – and ones they don’t – can…

Source
Ai2
Date
Ai2@allen_ai
Thread · 6 parts

Training an LLM by showing it answers people prefer – and ones they don’t – can improve it while quietly worsening behaviors. @GoodfireAI used our open post-training stack to predict how a full training run would change responses to different prompts. 🧵 https://t.co/KZV2gFZMaX

@GoodfireAI The problem: teams often discover unwanted changes in model behavior only after training is finished. Then they have to work backward from eval scores to guess which of hundreds of thousands of training examples caused them.

@GoodfireAI Goodfire wanted to move that debugging earlier. Ai2’s open stack made it possible. Our Dolci dataset provides Olmo 3’s preference data, Olmo includes intermediate checkpoints + reproducible recipes, and OLMES measures changes in model capabilities.

@GoodfireAI Using that stack, Goodfire developed “predictive data debugging”: a way to estimate which behaviors preference training will strengthen or suppress before committing compute to a full run.

Read the full thread on X
Key takeaways · AI-distilled
  • Ai2 frames the problem as teams discovering unwanted behavior changes only after preference training finishes, then working backward from scores to guess which of hundreds of thousands of examples caused them.
  • Goodfire's "predictive data debugging" estimates which behaviors a preference run will strengthen or suppress before committing compute, built on Dolci preference data, Olmo intermediate checkpoints and reproducible recipes, and OLMES capability evals.
  • In one experiment the regression was higher compliance with harmful requests framed as fiction or hypotheticals. Goodfire traced part of it, not all, to specific Dolci preference pairs.
Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters

You can trace a behavior regression from preference tuning back to specific training pairs, and estimate it before spending compute on a full run.

More from Ai2
Recommended reads
Comments

Checking sign-in…

Loading comments…