@Jason @Lons @eisokant @ape Correction: Should actually be closer to 50 tok/sec
Source
alexocheema
Author
alexocheema
Date
Terms in this piece · Glossary
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.
Why it matters
Explains why local tok/sec roughly doubles with multi-tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → prediction and where support actually exists on Apple silicon today — which changes what hardware you need for a usable local coding model.