2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Source
youtube.com
Author
Latent Space
Date
Why it matters
Explains non-transformerThe neural network architecture behind modern AI models, built on attention — letting every word directly consider every other word in parallel.Full definition → architectures that handle very long contexts more cheaply, which matters if inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → cost and context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → limit what you can build.
Terms in this piece · Glossary
transformer — The neural network architecture behind modern AI models, built on attention — letting every word directly consider every other word in parallel.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.