← All IntelIntel / article
Phobos – a tiny kernel language that now runs LLMs on a 2019 GPU
- Source
- joa-ebert.com
- Author
- joajoa
- Date
Why it matters
It shows how far runtime-compiled kernels can go on older consumer GPUs, with honest comparisons to llama.cpp on 1-bit-class quantizations and models. Useful for anyone running local on 8 GB cards.
Terms in this piece · Glossary
- mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Read the source www.joa-ebert.com
Recommended reads
Comments
Checking sign-in…
Loading comments…

