Vibeleaderboard
← All Intel
Intel / video

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

Source
youtube.com
Author
AI Engineer
Date
Why it matters

Lists the specific serving layers that determine cost and latency for self-hosted open models, which helps when weighing a closed API against running .

Terms in this piece · Glossary
  • speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.
  • KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…