Vibeleaderboard
← All Intel
Intel / video

Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance

Source
youtube.com
Author
20VC
Date
Why it matters

Argues that GPU off-chip memory architecture caps utilization at roughly 5 to 7%, which bears on where inference speed and cost may move. Useful for engineers planning around inference cost and latency.

Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source www.youtube.com
More from 20VC
Recommended reads
Comments

Checking sign-in…

Loading comments…