Vibeleaderboard
← All Intel
Clip / Education

Two conditions for the speedup, and why early models missed them

From Making Diffusion LLMs Faster: A Survey of Speedup Techniques · ≈2:43

“diffusion LLMs bring down inference time complexity from linear in the number of output tokens to a constant.”

Julia Turc

“The number of refinement steps has to be low, and each diffusion step has to cost roughly the same as a single autoregressive pass.”

Julia Turc

“Llama required up to 1,000 diffusion steps, which largely erased the benefits of parallelism and left it only marginally faster.”

Julia Turc

“In the context of traditional autoregressive models, there's really no freedom. Training and inference are very tight coupled, and so you train a model to predict the next token, and at inference time more or less, you kind of like just predict one token at a time.”

Julia Turc
What’s in it
  • Sets the honest bar — constant-time inference only materializes if refinement steps are few and each step costs about one autoregressive pass — and shows the first scaled open diffusion model failing that bar.
Clip transcript
or sampling, is executed for a fixed number of steps. So, diffusion LLMs bring down inference time complexity from linear in the number of output tokens to a constant. But, the speed of promise only materializes if two things are true. The number of refinement steps has to be low, and each diffusion step has to cost roughly the same as a single autoregressive pass. So, we'll start with the first challenge. One that is really critical is to be able to refine the sentence using a small number of diffusion steps. If you need a lot of diffusion steps, then there is no benefit compared to to one autoregressive model. So, there is kind of like a statistical problem that you need to be able to train these neural networks to be very efficient at reducing noise and doing the right kinds of modifications to the tokens. Early diffusion LLMs struggled with this. For instance, consider the original Llama, the first open-source diffusion model to scale up to 8 billion parameters and enable head-to-head comparisons with mature autoregressive models like Gwen 2.58B. Llama required up to 1,000 diffusion steps, which largely erased the benefits of parallelism and left it only marginally faster. To address this issue, there's been research on two separate fronts. Addressing training techniques and improving inference algorithms. The ability to decouple these two aspects is quite unique to diffusion models. In the context of traditional autoregressive models, there's really no freedom. Training and inference are very tight coupled, and so you train a model to predict the next token, and at inference time more or less, you kind of like just predict one token at a time. And maybe you can play around a little bit with different kinds of sampling scheme for that single token, top K, top P, greedy, temperature scaling, but roughly that's fixed. In the context of a diffusion language models, there is many more axes that you can explore. There is different kinds of noise processes, where you can mask, you can change the value of a token, uh which correspond to then different kinds of inference algorithms, because then during the denoising, you're going to apply different kinds of edits, different kind of modifications to your to your text. We'll first look at the training techniques that aim to
Recommended reads
Comments

Checking sign-in…

Loading comments…