From Why LLM Inference Is Memory-Bound (Julia Turc) · ≈3:37
“The important takeaway is that for every single token generated, the entire model, all billions of parameters, must complete this journey from the HBM to the processor.”
“It's a continuous pipeline of bytes being copied.”
articleSpeculative Correction: Draft-then-Refine Decoding for Diffusion Language ModelsBrian K Chen, Chong Wu, Kenji Kawaguchi
articleRevisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure ModesTianyu Wang, Yuxuan Zhou, Wenbin Wang, Heng Li, Zikai Xiao, Junyuan Shang
clipPrecompute via generated programs vs real-time inferenceWorkOSSign in to comment.
Loading comments…