Changing a model's behavior by nudging its internal activations directly at runtime — no retraining, no prompt changes.
Researchers find directions in a model's internal activation space that correspond to traits — honesty, refusal, a persona — and add or subtract them while the model runs. Turn the "sycophancy" direction down, and the same model with the same prompt flatters less.
It comes from interpretability research and matters for two reasons: it's evidence that human-level concepts live as manipulable directions inside these networks, and it hints at control knobs deeper than prompting and cheaper than fine-tuning. Papers in this index use it to steer agent personas at scale and probe safety behavior.