Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack
Source
Dylan Patel
Author
Dylan Patel
Date
Terms in this piece · Glossary
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Splitting prefill and decode onto different silicon changes the economics of long-context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition →inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → — the phase that dominates cost when agents feed large contexts on every turn.