Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Pricing
Open Source
Platform
cli · desktop
Type
TOOL
Added
Aug 7, 2026

About

A narrow, purpose-built local inference engine for DeepSeek V4 Flash/PRO and GLM 5.2 models, supporting Metal, CUDA, and ROCm backends with SSD streaming, multi-GPU tensor/pipeline parallelism, and a built-in HTTP server and coding agent. It targets high-RAM consumer and workstation hardware (e.g., 96GB+ Macs, DGX Spark, Framework Desktop) and ships aggressive routed-expert quantization tuned specifically for these models.

Why it made the leaderboard

Co-designing the serving stack with one model family instead of running a general GGUF runtime is a real bet on local inference quality, from a maintainer with a track record of shipping systems code.

Tags

llm-inferencedeepseekggufmetalcudarocmlocal-inferencecoding-agent

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.