The delay between sending a model request and receiving the first piece of its response.
Time to first token measures how quickly an AI experience begins responding. It includes queueing, prompt processing, routing, network delay, and any cold start before generation begins.
It should be measured separately from generation speed: a provider can start quickly and then produce the rest of the answer slowly, or the reverse.