Vibeleaderboard
← All Intel
Intel / article

Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

Source
Liming Liu, Mingze Wang, Tuo Zhao
Author
Liming Liu, Mingze Wang, Tuo Zhao
Date
Terms in this piece · Glossary
  • transformer — The neural network architecture behind modern AI models, built on attention — letting every word directly consider every other word in parallel.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
Why it matters

cost now dominates training cost for served models, and this decouples the two phases so added capability does not tax both.

Key quotes

“As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost.”

Liming Liu, Mingze Wang, Tuo Zhao

“The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregressive decode is sequential and often memory-bandwidth-bound.”

Liming Liu, Mingze Wang, Tuo Zhao

“Its primary flow is a complete causal language model that processes the prompt and writes the KV cache. The auxiliary flow is omitted during prompt processing and activated only from the final prompt position onward, adding continuation-prediction computation without writing persistent state or influencing the primary flow.”

Liming Liu, Mingze Wang, Tuo Zhao

“In MoE models, the separation makes primary and auxiliary expert fan-outs independent controls over prompt cost, continuation cost, and predictive quality.”

Liming Liu, Mingze Wang, Tuo Zhao
Recommended reads
Comments

Checking sign-in…

Loading comments…