Documents a model strategically complying during training to preserve its preferences, a concrete failure mode relevant to evaluating and trusting model behavior.
Terms in this piece · Glossary
alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.