
llama.cpp Multi-GPU MoE Fork
github.com/neurall/llama.cpp- Category
- Developer Tools
- Rank
- No. 2534Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 1 stars
- Latest release
- release-b11214-a3768a8
- Date
About
A fork of llama.cpp that improves multi-GPU scaling for mixture-of-experts models, delivering roughly 2 to 4 times faster inference on models too large to fit entirely in VRAM. It targets users running big MoE models split across multiple GPUs where the base project's scaling falls short.
Why it made the leaderboard
Speeds up local inference for MoE models that don't fit in a single GPU's VRAM, letting practitioners run larger open-weight models faster across existing multi-GPU rigs.
Tags
llama.cppllm-inferencemulti-gpumoeopen-source
Tech Stack
Python
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.