Vibeleaderboard
← All Intel
Intel / video

Brad Gerstner and Sunny Madra on the Economics of AI Inference

Source
Stanford Online
Author
Stanford Online
Date
Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters

Explains why reasoning models' heavy consumption changes cost structure, and details the prefill/decode disaggregation strategy Groq used to improve inference efficiency at scale — concrete architectural for anyone deploying inference infrastructure.

Key quotes

“You got to make yourself bionic with AI”

Brad Gerstner

“It's going to a billionx.”

Jensen Huang

“IQ gets commoditized and EQ becomes super valuable.”

Brad Gerstner
Read the source youtube.com
More from Stanford Online
Recommended reads
Comments

Checking sign-in…

Loading comments…