Find the right skill, CLI, harness, or service for the job.
Running a trained model to produce an output.
Inference is the runtime phase where a model turns input tokens into predictions. Model size, hardware, batching, context length, and reasoning budget shape its latency, throughput, memory use, and cost.