
If you only have intuition for autoregressive next- generation, this gives you a concrete mental model of how diffusion LLMs like LLaDA generate text via iterative unmasking — including why BERT-style masked modeling is the ancestor rather than GPT. Useful before you evaluate whether diffusion-based decoding is worth trying for latency- or infilling-sensitive workloads.
“It starts from gibberish and iteratively turns it into a coherent response. Patterns emerge from noise.”
Julia Turc
“What the transformer did was to promote the attention mechanism from a supporting role to the main building block of the translation network.”
Julia Turc
“A single wild output can derail the entire response because the model never learned how to deal with outliers.”
Julia Turc
“there's a clear convergence towards masked diffusion models, the ones inspired by BERT”
Julia Turc
Checking sign-in…
Loading comments…