Vibeleaderboard
← All Intel
Intel / video

How to Scale AI Application Inference 100x ft. Fireworks’ Lin Qiao

Source
youtube.com
Author
Sequoia Capital
Date
Why it matters

Frames inference as a product- problem and covers what scaling application inference involves, from a leading inference provider. Helps teams plan latency, quality and cost tradeoffs as usage grows.

Key takeaways · AI-distilled
  • Lin Qiao frames the future of scaling as a per-application optimization across three dimensions at once: quality, speed, and user concurrency, which she equates with cost.
  • Qiao argues inference should not be optimized in isolation; co-optimizing post-training and inference together is how she expects to push today's high inference costs down by 10x to 100x.
  • Per Qiao, choices like predicting many at once, numeric precision, hardware SKU, model sharding, cross-host distribution, kernel selection and quality tuning combine into over 100,000 configurations.
  • Qiao says model researchers must guess who will use a model and for what, so a gap with real application data is almost always guaranteed; production data should steer both tuning and inference system design.
  • Qiao cites a food chain that scaled an AI feature from one shop to a thousand in three months, and a software company that went from 100,000 to 25 million developers in three months on Fireworks.
Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Read the source www.youtube.com
More from Sequoia Capital
Recommended reads
Comments

Checking sign-in…

Loading comments…