
OpenAI Ultrafast mode now available on AI Gateway
Quick Answer
OpenAI's Ultrafast service tier for GPT-6 Astra is now available on AI Gateway, enabling faster outputs for interactive applications.
Quick Take
Users can access it via the AI SDK, Chat Completions API, or Responses API, with costs at 6× the standard per-token rate, while standard processing remains the default for unsupported regions.
Key Points
- Ultrafast service tier enhances performance for interactive applications and rapid coding.
- Requests to unsupported regions default to standard processing rates.
- Using the Responses API is recommended for frequent tool calls to minimize overhead.
- Requests billed at 6× the standard rate for Ultrafast processing.
- Check the GPT-6 Astra model page for current pricing details.
📖 Reader Mode
~1 min readAI Gateway now supports OpenAI's Ultrafast service tier for GPT-6 Astra, providing faster output for interactive applications and rapid coding iterations.
To use Ultrafast, request it for openai/gpt-6-astra through AI SDK, the Chat Completions API, or Responses API:
For workflows with frequent tool calls, OpenAI recommends the Responses API over a persistent WebSocket connection to reduce overhead between turns. See the Ultrafast service-tier examples for persistent connections, including AI SDK over WebSocket.
Ultrafast supports US and global processing. Requests pinned to unsupported regions, such as the EU, run at the standard (default) tier. Standard processing remains the default when no service tier is specified.
Requests served at Ultrafast are billed at 6× the standard per-token rate, while requests that fall back to another tier are billed at the rate for the tier actually served. Check the GPT-6 Astra model page for current rates and learn more about the GPT-6 model family.
— Originally published at vercel.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Vercel AI
See more →
The Agent Stack
The Agent Stack by Vercel AI provides essential building blocks for creating production-grade agents, enabling seamless integration across multiple AI models and secure operations. It features components like AI Gateway for model routing, Workflow SDK for durable execution, and Vercel Connect for scoped access, streamlining agent development and deployment across various platforms.

