Vibeleaderboard
← All Intel
Intel / video

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)

Source
youtube.com
Author
Latent Space
Date
Why it matters

Covers practical issues in serving a 671B model, including SGLang, and pricing. Helps you judge the cost and engineering of running open models.

Terms in this piece · Glossary
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Read the source www.youtube.com
More from Latent Space
Recommended reads
Comments

Checking sign-in…

Loading comments…