0 Apps · 1 Tool · 0 Intel
A library that runs 70B-parameter LLM inference on a single 4GB GPU through aggressive layer-by-layer memory management.