
Regional inference now available on AI Gateway
Quick Answer
AI Gateway now enables regional inference, allowing users to pin requests to either the US or EU.
Quick Take
This simplifies compliance for teams by ensuring data residency, with requests failing if no provider can serve the selected region. The feature incurs a potential cost increase of about 10% above standard rates.
Key Points
- Inference can be pinned to US or EU regions for compliance.
- Requests fail if no model provider can serve the selected region.
- Responses indicate the region that processed each request.
- Regional inference may incur a cost increase of around 10%.
- Filter available models by region using the API.
DeepSignal Analysis
What happened
AI Gateway has introduced regional inference, allowing users to specify requests to the US or EU. This feature enhances compliance by ensuring data residency, as requests will fail if no provider can serve the chosen region. However, this may lead to a cost increase of approximately 10% over standard rates.
Key evidence
- AI Gateway now supports regional inference, enabling users to pin requests to either the US or EU data centers.
- If no model provider can serve the selected region, the request fails instead of being rerouted globally.
- Pinning a region may incur additional costs, typically around 10% above standard rates, which are passed through by AI Gateway.
Why it matters
The introduction of regional inference simplifies compliance for organizations with data residency requirements. By allowing users to confirm where their requests are processed, it reduces the complexity of managing regional routing across multiple providers. This can be particularly beneficial for companies operating in regulated industries that must adhere to strict data handling policies.
What to watch
Source Excerpt
AI Gateway now supports regional inference, letting you pin each request to a US or EU data center for data residency. Requests fail closed if the region can't be honored, and every response reports the region inference actually ran in.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Vercel AI
See more →
The Agent Stack
The Agent Stack by Vercel AI provides essential building blocks for creating production-grade agents, enabling seamless integration across multiple AI models and secure operations. It features components like AI Gateway for model routing, Workflow SDK for durable execution, and Vercel Connect for scoped access, streamlining agent development and deployment across various platforms.

