Vibeleaderboard
← All Intel
Intel / article

Constitutional Midtraining: Content Presence Drives Alignment Gains

Source
arxiv.org
Author
Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
Date
Why it matters

Supervised introduced a blackmail propensity in every model tested, and what survived it was constitutional content merely present during midtraining — its structure mattered less.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
  • alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
Recommended reads
Comments

Checking sign-in…

Loading comments…