Vibeleaderboard
← All Intel
Intel / post

@Jason @Lons @eisokant @ape Correction: Should actually be closer to 50 tok/sec

Source
x.com
Author
alexocheema
Date
Why it matters

Explains why local tok/sec roughly doubles with multi- prediction and where support actually exists on Apple silicon today — which changes what hardware you need for a usable local coding model.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • speculative decoding — A speed trick where a small model drafts several tokens ahead and the big model verifies them in one pass, often doubling generation speed.
More from alexocheema
Recommended reads
Comments

Checking sign-in…

Loading comments…