Vibeleaderboard
← All Intel
Intel / video

How GPT, Claude, and Gemini are actually trained and served – Reiner Pope

Source
youtube.com
Author
Dwarkesh Patel
Date
Why it matters

Explains the batching and memory tradeoffs behind latency tiers and API prices, so you can predict cost and speed constraints when designing workloads.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Read the source www.youtube.com
More from Dwarkesh Patel
Recommended reads
Comments

Checking sign-in…

Loading comments…