What Makes Open Models Fast in Production — Sujee Maniyam, Nebius
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Lists the specific serving layers that determine cost and latency for self-hosted open models, which helps when weighing a closed API against running open weightsA model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.Full definition →.
Terms in this piece · Glossary
speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.
KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.