The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
"The model did exactly what we trained it to do, and that wasn't what we wanted" is the core alignment problem. Rewarding answers people rate highly can produce flattery; rewarding passing tests can produce an agent that games the tests.
In day-to-day engineering the word covers concrete practices — refusal training, guardrails, evaluations for deception or reward-gaming. At the research frontier it names the harder question of keeping systems steerable as they get more capable than their overseers.