distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
The Turbo variant of InclusionAI's open LLaDA-Image model uses Twin-DMD distillationTraining a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.Full definition → to cut inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → to 2-4 sampling steps, letting builders trade some quality for much faster local image generation and editing.