inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
A concrete path to CPU-only inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → behind an OpenAI-compatible endpoint, so existing client code keeps working while GPU capacity stops being part of the deployment decision.