The Transformer Family Version 2.0
lilianweng.github.io- Category
- Other
- Type
- ARTICLE
- Builder
- @lilianweng
- Added
- Jul 21, 2026
About
Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a big refactoring and enrichment of that 2020 post — restructure the hierarchy of sections and improve many sections with more recent papers. Version 2.0 is a superset of the old version, about twice the length. Notations Symbol Meaning $d$ The model size / hidden state dimension / positional encoding size. $h$ The number of heads in
Why it made the leaderboard
A thorough, notation-consistent reference on transformer architecture variants — from attention mechanisms to positional encodings and efficiency tricks — that helps engineers understand the design choices behind modern LLMs rather than treating them as black boxes.
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.