Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance
Source
youtube.com
Author
20VC
Date
Why it matters
Argues that GPU off-chip memory architecture caps inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → utilization at roughly 5 to 7%, which bears on where inference speed and cost may move. Useful context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → for engineers planning around inference cost and latency.
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.