LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Offers a concrete, hardware-validated technique for shrinking ternary LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → storage and speeding up inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → on both CPUs and GPUs, directly useful for anyone deploying low-bit models.