Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
A large "teacher" model generates high-quality outputs; a smaller "student" trains on them. The student learns from the teacher's rich, already-digested behavior rather than raw internet text, so it punches far above its size — most impressive small models are distillations of something bigger.
It's also a competitive fault line: API access to a frontier model is enough to distill from it, which providers' terms of service typically prohibit and which several public disputes have centered on.