Brad Gerstner and Sunny Madra on the Economics of AI Inference
Source
Stanford Online
Author
Stanford Online
Date
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters
Explains why reasoning models' heavy tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → consumption changes inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → cost structure, and details the prefill/decode disaggregation strategy Groq used to improve inference efficiency at scale — concrete architectural context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → for anyone deploying LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → inference infrastructure.
Key quotes
“You got to make yourself bionic with AI”
Brad Gerstner
“It's going to a billionx.”
Jensen Huang
“IQ gets commoditized and EQ becomes super valuable.”