
If you're tracking diffusion LLMs as a lower-latency alternative to autoregressive decoding, this maps the whole speedup landscape — caching tricks, unmasking schedules, block diffusion — onto the specific papers and shipping models that use them, so you can tell which claimed speedups actually transfer to your workload.
“It's a very new technology and approach and I think we're still far from even a local optimum.”
Stefano Ermon
“If you need a lot of diffusion steps, then there is no benefit compared to to one autoregressive model.”
Stefano Ermon
“KV caching as implemented for auto-regressive models is simply not possible for the fusion models.”
Julia Turc
“So, the virus basically spreads to the entire context window.”
Julia Turc
“I see a future where diffusion models will become the leading paradigm not just for image and video generation but also for for discrete object like text and code.”
Stefano Ermon
Checking sign-in…
Loading comments…