- Category
- AI Tools
- Rank
- No. 1936Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 1.1k stars
- Latest release
- v1.4.2
- Added
- Aug 19, 2026
About
An inference library for running large language models locally on consumer GPUs, built around its own EXL3 quantization format. It supports tensor-parallel and expert-parallel inference across consumer hardware setups, and is served through an OpenAI-compatible API via TabbyAPI.
Tags
llmlocalinferencegpuquantization
Tech Stack
Python
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
