- Category
- AI Tools
- Type
- APP
- Builder
- jundot
- GitHub
- 18.4k stars
- Added
- Jun 16, 2026
About
An LLM inference server for Apple Silicon with continuous batching and SSD caching.
Why it made the leaderboard
An LLM inference server for Apple Silicon with continuous batching and tiered KV caching that spills context from hot memory to SSD, so long contexts survive model swaps. Menu-bar management gives it the convenience of an app with the control of a server.
Tech Stack
Python
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
