Transcript
Jonathan Ross is the founder of Groq, a hardware chip specifically designed for LLM inference, which entered a $20 billion strategic agreement with NVIDIA. Topics covered: • The "success disaster" at Google that led to the TPU and eventually Groq • LPU vs. GPU: Pareto curves, cost-per-token, and when each wins • Static scheduling • Mixture-of-experts models • Auto-regressive vs. diffusion models • How Groq and NVIDIA's Vera Rubin work together at inference time • Jevons Paradox: why cheaper AI will increase total compute demand • Will AI replace CUDA kernel engineers? • What skills kids should be learning in the AI age Groq: https://groq.com/ NVIDIA Vera Rubin: https://www.nvidia.com/en-us/data-center/technologies/rubin/