← All IntelClip / EducationThe TPU was born because Google couldn't deploy a working model
From "An endless demand for compute" | Jonathan Ross, founder of Groq · ≈1:26
First-hand origin story of inference-specific silicon: a superhuman speech model was gated to one phone by compute cost, and a 20% FPGA port became the TPU after Jeff Dean's cost analysis.
What’s in it
- First-hand origin story of inference-specific silicon: a superhuman speech model was gated to one phone by compute cost, and a 20% FPGA port became the TPU after Jeff Dean's cost analysis.
Clip transcript
see at Google that made you want to build something different? >> We didn't have enough compute. Uh what had happened was the speech recognition team had trained a model and that model was better than human beings at um transcribing and this was the first time that they had ever achieved that. U the problem was they couldn't put it into production. They had actually limited the deployment to you remember the Nexus phone, the old Android phone. >> I had one. >> Oh yeah. Okay. So they limited it to Nexus, not so much as a feature, but because they had so little compute, they could only support the Nexus um uh user base. >> So I happened to be having lunch with the speech recognition team in New York City and they mentioned this problem and started um as a 20% project uh porting their model to an FPGA created a general architecture and then turns out inference was uh needed quite badly and it became a chip. So actually Jeff Dean did an analysis and was like hm given what we're going to spend on this in the capacity let's just do an ASIC instead. >> My response was how hard could it be? Turns out it was very hard but we didn't know so jumped in.
Comments
Checking sign-in…
Loading comments…