
Shows one way out of the usual flexibility-versus-latency tradeoff in ML serving: express ranking pipelines in a DSL compiled to a native execution engine, so model iteration stays fast without paying interpreted-runtime latency.
Checking sign-in…
Loading comments…