
The search-and-serving layer under the largest public model hub, described with the specific tradeoffs and failure points that anyone running AI infrastructure at scale will hit.
“At 20,000 models, any query is fast. Even without an index, trust me, no one would notice. At 3 million, same approach breaks.”
Arek Borucki
“P99 is much more important than P50 and we are paying a lots of attention to P99”
Arek Borucki
“MongoDB does not store the models themselves. It stores everything about the models.”
Arek Borucki
“We tokenize model names on insert time, not at query time.”
Arek Borucki
“The pattern is simple. Primary should focus on what only primary can do. Anything else can be pushed to different machines.”
Arek Borucki
videoHow Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, OxylabsAI Engineer
podcastModel Mayhem, NVIDIA x Hugging Face Deal, GPT-6 Astra | Pablo Torre, Mohit Aron, Akshay Narisetti, Hari Ravichandran, Matt Caldwell & Jordy Leiser, Jeff Thornburg, Charlie O’Neill, Sridhar RamaswamyJohn Coogan & Jordi Hays
videoLocal Models: Trust, Control, Optimization — Carter Abdallah, NVIDIAAI EngineerChecking sign-in…
Loading comments…