
Anyone running retrieval over customer documents hits the same wall Harvey did: the naive single-pipeline design breaks unevenly across extraction, , and indexing. This post shows how one of the largest production legal-AI deployments split those stages to hold latency and reliability through a 26x volume increase.
“The main lesson was simple: At this scale, document processing stops being one pipeline and starts behaving like several different systems.”
“The latest complete week in the measurement window averaged about 3.5 million documents per day. Compared with a year ago, that is 26x more documents and 39x more data.”
“Instead of relying on one-off uploads, users are increasingly treating us as a central system of record for massive document collections.”
“Growth was not just more documents; it was more kinds of documents, with more OCR, conversion, and indexing edge cases showing up at production frequency.”
articleHow our universal content processing platform Riviera evolved for AI and beyondIlya Yakovlev,Andrew Cheung,Binoy Dash
videoServing 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging FaceAI Engineer
videoHow Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, OxylabsAI EngineerChecking sign-in…
Loading comments…