GridCore – a scheduler that lets several local LLMs share one GPU
Source
blokhin.us
Author
w512
Date
Why it matters
Practical design rules for running several local models on one consumer GPU, including priority queues and VRAM estimation that routers like llama-swap lack. Useful if agents and indexing jobs compete for local inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition →.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.