- Category
- AI Tools
- Rank
- No. 1076Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 4.6k stars
- Latest release
- v0.3.2
- Added
- Aug 19, 2026
About
A fast inference library for running large language models locally on modern consumer-class GPUs, aimed at usable speed on hardware that is not datacenter-grade.
Why it made the leaderboard
Local inference on consumer hardware is the cheapest way to iterate on model-backed features, and this is one of the few libraries that makes it fast enough to be usable.
Tags
llmlocalinferencegpuquantization
Tech Stack
Python
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
