- Category
- AI Tools
- Rank
- No. 900Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 89.4k stars
- Latest release
- v0.27.1
- Added
- Aug 19, 2026
About
An inference and serving engine for large language models, built for high throughput and efficient memory use. It is self-hosted: you run it on your own GPUs and serve models from them rather than calling a hosted API.
Intel on vLLM
Tags
llminferenceservinggpuself-hosted
Tech Stack
Python
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
