inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters
Teams blocked from hosted LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → APIs by tenancy or rate-limit constraints get a self-hosting-like isolation boundary for Cohere models without running the infrastructure themselves.