Vibeleaderboard
Index / tool

llama.cpp Multi-GPU MoE Fork

github.com/neurall/llama.cpp
Visit github.com
Category
Developer Tools
Rank
No. 2534Tools index
Pricing
Open Source
Type
TOOL
GitHub
1 stars
Date

About

A fork of llama.cpp that improves multi-GPU scaling for mixture-of-experts models, delivering roughly 2 to 4 times faster inference on models too large to fit entirely in VRAM. It targets users running big MoE models split across multiple GPUs where the base project's scaling falls short.

Why it made the leaderboard

Speeds up local inference for MoE models that don't fit in a single GPU's VRAM, letting practitioners run larger open-weight models faster across existing multi-GPU rigs.

Tags

llama.cppllm-inferencemulti-gpumoeopen-source

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.