A service or runtime that executes a model and exposes it to an application.
A model lab may serve its own models directly, but applications can also use routers, specialist inference clouds, cloud catalogs, or local runtimes. The provider affects latency, price, regional availability, privacy, rate limits, feature support, and fallback behavior even when the model name appears identical.