
Corpora that induce broad misalignment share a measurable Big Five signature — lower agreeableness and conscientiousness, higher extraversion and neuroticism — detectable in the training data before imprints it on the model.
articleOptimismBench: Forecasting Bias and the Alignment Effect in Language Model JudgmentSeonglae Cho, Adriano Koshiyama
articleConstitutional Midtraining: Content Presence Drives Alignment GainsDesiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel ShadboltChecking sign-in…
Loading comments…