← All IntelClip / Developer ToolsWhy search had to be rethought, and why P99 beats P50
From Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face · ≈4:02
Makes the tail-latency argument concrete: at 14M users even 1% slow queries is 140,000 people, which is what justified precomputed tokens, denormalized read collections, and Lucene search.
What’s in it
- Makes the tail-latency argument concrete: at 14M users even 1% slow queries is 140,000 people, which is what justified precomputed tokens, denormalized read collections, and Lucene search.
Clip transcript
indexed, and also must be searchable. And that's the hardest part. And this is also the reason why we had to rethink our search. At 20,000 models, any query is fast. Even without an index, trust me, no one would notice. At 3 million, same approach breaks. Imagine what would you do if the hub search would be slow. you would just leave and go somewhere else. And this is also what user are doing. They expect fast instant results. With 14 million users, even 1% is a not small number. It is 140,000 of people hitting slow search at scale. P99 is much more important than P50 and we are paying a lots of attention to P99 and that's the reason why we invest in premputee tokens denormalize optimize for read collection in MongoDB full text search based on Apache lucine Kubernetes autoscaling and soon in database sharding.
Comments
Checking sign-in…
Loading comments…