Vibeleaderboard
← All Intel
Intel / article

Fast Inference from Transformers via Speculative Decoding

Source
arxiv.org
Author
Yaniv Leviathan et al.
Date
Why it matters

It makes generation faster without changing what the model outputs — a small model drafts, the large one verifies in parallel, rejected fall back. Every fast provider runs some version of it, and the identical-distribution guarantee is why it costs no quality.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Recommended reads
Comments

Checking sign-in…

Loading comments…