LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Temperature 0 is not deterministic because kernels change reduction strategy with batch size, not because GPUs are parallel. Batch-invariant kernels make inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → bitwise reproducible and collapse the train/inference gap in on-policy RL.