inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Signals that Apple's native containerization stack may soon support GPU passthrough, which would let local models like Ollama run properly isolated inside per-container VMs on Apple Silicon.