- Category
- Developer Tools
- Listed in
- #1 Interface with your agents
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 2.4k stars
- Latest release
- v0.8.0
- Date
About
WARP (Weight-Aware Runtime and Paging, formerly WASTE) is a dependency-free, embeddable C inference engine that keeps a model's trunk in memory while streaming selected expert weights directly from NVMe disk, using remaining RAM as a bounded expert cache. This lets it run large mixture-of-experts models such as Kimi K3, DeepSeek V4.1 Flash, and GLM-5.3-Flash on hardware with far less RAM than the full model would normally require.
Why it made the leaderboard
Lets engineers run massive MoE models locally on modest RAM by streaming only the activated experts from disk, dramatically lowering the hardware bar for running frontier-scale open models.
Intel on WARP
Tags
inference-enginellmmoenvmecopen-source
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.
