← All IntelClip / AI ToolsSpeculative decoding speed benchmark
From Meta Open Source Is BACK – Muse Glimmer First Test! · ≈3:34
“We can see right here on an RTX 5090 without speculative decoding, it's 75 tokens per second, which is quick for a 30 billion parameter dense model.”
“With this Dlash spec model, it's 233 tokens per second.”
“Apple's like the best value for running local AI now which is quite perplexing”
What’s in it
- Shows a 30B model hitting 233 tokens/sec via speculative decoding
- Compares RTX 5090 vs Apple Silicon for local LLM speed
- Flags Apple Macs as surprisingly best value for local AI
Clip transcript
testing it at that specific context length. Now, additionally to this, and this is something that's really awesome, is this has a speculative decoding model with it as well, which is what this section is about right here. This seems this is not scientific, but this seems like the most optimized setup to see how fast it could get potentially. We can see right here on an RTX 5090 without speculative decoding, it's 75 tokens per second, which is quick for a 30 billion parameter dense model. With this Dlash spec model, it's 233 tokens per second. That's an insane speed up. And they also show it on some Macs as well because Apple's like the best value for running local AI now which is quite perplexing but again and then we just have some
Comments
Checking sign-in…
Loading comments…