Vibeleaderboard
← All Intel
Intel / article

Alignment Forecasting: Predicting Misalignment From Training Data

Source
Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
Author
Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
Date
Key takeaways · AI-distilled
  • AlignmentForecastBench has over 5,000 questions spanning 17 target models, 32 datasets and 16 failure modes. A forecaster outputs the probability that fine-tuning on a dataset would meaningfully increase a given failure mode.
  • Frontier models prompted directly forecast poorly. The proposed scaffold has an rate how strongly and broadly a dataset pushes toward misbehavior, then a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency.
  • The scaffold forecasts well above chance and beats both a model fine-tuned on the task and a forecaster allowed to see how weaker models behaved after training on the same data, the authors report.
  • Removing examples the scaffold flags from real post-training data such as UltraChat gave more aligned models on a multiple-choice in most cases, but the benefit in open-ended conversation is unclear and the results cover only SFT.
Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters

Offers a way to estimate, before training, whether a fine-tuning dataset will raise failures like deception or sycophancy, instead of discovering it by auditing the trained model afterward.

Recommended reads
Comments

Checking sign-in…

Loading comments…