Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Pricing
Open Source
Type
TOOL
Latest release
v0.8.0
Date

About

WARP (Weight-Aware Runtime and Paging, formerly WASTE) is a dependency-free, embeddable C inference engine that keeps a model's trunk in memory while streaming selected expert weights directly from NVMe disk, using remaining RAM as a bounded expert cache. This lets it run large mixture-of-experts models such as Kimi K3, DeepSeek V4.1 Flash, and GLM-5.3-Flash on hardware with far less RAM than the full model would normally require.

Why it made the leaderboard

Lets engineers run massive MoE models locally on modest RAM by streaming only the activated experts from disk, dramatically lowering the hardware bar for running frontier-scale open models.

Intel on WARP

More in Intel

Tags

inference-enginellmmoenvmecopen-source

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.