Vibeleaderboard
← All Intel
Intel / article

GridCore – a scheduler that lets several local LLMs share one GPU

Source
blokhin.us
Author
w512
Date
Why it matters

Practical design rules for running several local models on one consumer GPU, including priority queues and VRAM estimation that routers like llama-swap lack. Useful if agents and indexing jobs compete for local .

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Recommended reads
Comments

Checking sign-in…

Loading comments…