From Transformers to Diffusion LLMs: Understanding LLaDA
- Source
- youtube.com
- Author
- Julia TurcTop Viber
- Date

If you only have intuition for autoregressive next- generation, this gives you a concrete mental model of how diffusion LLMs like LLaDA generate text via iterative unmasking — including why BERT-style masked modeling is the ancestor rather than GPT. Useful before you evaluate whether diffusion-based decoding is worth trying for latency- or infilling-sensitive workloads.
- transformer — The neural network architecture behind modern AI models, built on attention — letting every word directly consider every other word in parallel.
- attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
- LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
- token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
“It starts from gibberish and iteratively turns it into a coherent response. Patterns emerge from noise.”
Julia Turc
“What the transformer did was to promote the attention mechanism from a supporting role to the main building block of the translation network.”
Julia Turc
“A single wild output can derail the entire response because the model never learned how to deal with outliers.”
Julia Turc
“there's a clear convergence towards masked diffusion models, the ones inspired by BERT”
Julia Turc
Checking sign-in…
Loading comments…





