A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
Source
William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang
Author
William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang
Date
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Anyone sizing an inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → tier or tuning prefix caching is currently guessing from short synthetic traces. A year of real production traffic, including long-tail models, gives that work an empirical baseline.