Vibeleaderboard
← All Intel
Intel / video

Tuhin Srivastava (Baseten) on Production Inference Economics

Source
Stanford Online
Author
Stanford Online
Date
Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters

Baseten's CEO makes the case that once you are at scale, renting a frontier API stops being the cheap option and post-training an open model becomes both the cost play and the moat. He backs it with what real customers run, including one voice product chaining six models inside a latency budget.

Key quotes

“there's language models, there's audio models, there's um there's actually I think three or four language models in the middle there and two audio models to make that happen and all of them run on base 10.”

Stanford Online

“Um and today about 90% or 95% of spend on inference is going to frontier models and about 5% is going to custom models”

Stanford Online

“open source models about 90 90 days behind um Frontier models and you can run them about 70 to 70 to 90% cheaper.”

Stanford Online

“the leading coding companies that are not the frontier model companies themselves are, you know, still rumored to be negative gross margin and so I imagine is existential for them to to to be a viable business”

Stanford Online

“we put out a post yesterday about the world of many models. We think intelligence shouldn't be owned by two people.”

Stanford Online
Read the source youtube.com
More from Stanford Online
Recommended reads
Comments

Checking sign-in…

Loading comments…