Vibeleaderboard
← All Intel
Intel / video

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

Source
youtube.com
Author
Latent Space
Date
Why it matters

at thousands of per second changes what loops and runs are practical, such as cutting a 20-hour eval to about 2 hours.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Read the source www.youtube.com
More from Latent Space
Recommended reads
Comments

Checking sign-in…

Loading comments…