
The paper shows role-specific beats persona-vector control (63.2 vs 41.1 alignment) for controlling personas in social simulations, but flags that ~14% of roles degrade regardless of steering strength, arguing practitioners need per-role validation before deployment rather than uniform tuning.
“Role-specific directions receive higher judged role-profile alignment than an assistant-axis directional control from prior persona-vector work, with mean overall scores of 63.2 versus 41.1 across the tested grid.”
“The role-level screen is the main practical output: most roles improve as steering increases, but 38 roles decline across all six measured dimensions, showing why simulation builders should choose coefficients per role rather than deploy a uniform high-strength setting.”
articleDoes Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit ReversalPhilipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron
articleEven More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent SystemsMarylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi
articleWhat a crowdsourced game revealed about steering Olmo 3allenai.orgChecking sign-in…
Loading comments…