
If you keep hearing about Gemini Diffusion or Mercury Coder and want to understand why parallel draft-refinement can be ~10x faster than autoregressive decoding, this walks through the actual formulations (D3PM's Markov chain corruption, LLaDA's masked- approach) instead of stopping at claims. It's the level of detail needed to judge whether diffusion LLMs are worth designing around.
“So, ironically, diffusion leads to faster inference but slower training.”
Julia Turc
“So, in a way, BERT was a singlestep diffusion model.”
Julia Turc
“The difficulty of rounding is the reason why diffusion and token embedding space never really took off.”
Julia Turc
“There's no direct transition from dog to cat, just from dog to mask.”
Julia Turc
“But in language, because it's discreet, there's no concept of a gradient. So instead, we have to learn the landscape directly.”
Julia Turc
Checking sign-in…
Loading comments…