Vibeleaderboard
← All Intel
Intel / article

AI Model Co-Design: Hardware-Friendly LLM Design

Source
developer.nvidia.com
Author
Elizabeth Goodman
Date
Why it matters

Model shape decisions made before training silently cap serving throughput; this gives the concrete dimension, aspect-ratio and rules that keep a model from running badly on the hardware it will live on.

Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
Read the source developer.nvidia.com
More from Elizabeth Goodman
Recommended reads
Comments

Checking sign-in…

Loading comments…