Vibeleaderboard
← All Intel
Intel / article

Phobos – a tiny kernel language that now runs LLMs on a 2019 GPU

Source
joa-ebert.com
Author
joajoa
Date
Why it matters

It shows how far runtime-compiled kernels can go on older consumer GPUs, with honest comparisons to llama.cpp on 1-bit-class quantizations and models. Useful for anyone running local on 8 GB cards.

Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Read the source www.joa-ebert.com
Recommended reads
Comments

Checking sign-in…

Loading comments…