How much model output a provider can generate over time, commonly measured in output tokens per second.
For one interactive request, throughput describes how quickly text arrives after the first token. At system level it can also mean the total requests or tokens a deployment handles across many users.
Those are different measurements. Always check whether a benchmark describes single-request speed, total batch capacity, or both.