How to Scale AI Application Inference 100x ft. Fireworks’ Lin Qiao
Source
youtube.com
Author
Sequoia Capital
Date
Why it matters
Frames inference as a product-alignmentThe work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.Full definition → problem and covers what scaling application inference involves, from a leading inference provider. Helps teams plan latency, quality and cost tradeoffs as usage grows.
Key takeaways · AI-distilled
Lin Qiao frames the future of inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → scaling as a per-application optimization across three dimensions at once: quality, speed, and user concurrency, which she equates with cost.
Qiao argues inference should not be optimized in isolation; co-optimizing post-training and inference together is how she expects to push today's high inference costs down by 10x to 100x.
Per Qiao, choices like predicting many tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → at once, numeric precision, hardware SKU, model sharding, cross-host distribution, kernel selection and quality tuning combine into over 100,000 configurations.
Qiao says model researchers must guess who will use a model and for what, so a gap with real application data is almost always guaranteed; production data should steer both tuning and inference system design.
Qiao cites a food chain that scaled an AI feature from one shop to a thousand in three months, and a software company that went from 100,000 to 25 million developers in three months on Fireworks.
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.