Ultrafast gives faster GPT-6 Astra output for interactive and coding workflows at 6x the standard token rate, and EU-pinned requests fall back to the standard tier. Weigh the speed gain against cost and region constraints.
Key takeaways · AI-distilled
AI Gateway exposes OpenAI's Ultrafast service tier for GPT-6 Astra. You opt in per request with the openai provider option serviceTier: 'ultrafast'; standard processing remains the default when no tier is specified.
Ultrafast requests are billed at 6x the standard per-tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → rate, while a request that falls back to another tier is billed at the rate of the tier actually served.
Ultrafast supports US and global processing only. Requests pinned to unsupported regions such as the EU run at the standard (default) tier instead.
For workflows with frequent tool calls, OpenAI recommends using the Responses API over a persistent WebSocket connection to reduce overhead between turns.
Terms in this piece · Glossary
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.