
Regional inference now available on AI Gateway
Quick Answer
AI Gateway now enables regional inference, allowing users to pin requests to either the US or EU.
Quick Take
This simplifies compliance for teams by ensuring data residency, with requests failing if no provider can serve the selected region. The feature incurs a potential cost increase of about 10% above standard rates.
Key Points
- Inference can be pinned to US or EU regions for compliance.
- Requests fail if no model provider can serve the selected region.
- Responses indicate the region that processed each request.
- Regional inference may incur a cost increase of around 10%.
- Filter available models by region using the API.
DeepSignal Analysis
What happened
AI Gateway has introduced regional inference, allowing users to specify requests to the US or EU. This feature enhances compliance by ensuring data residency, as requests will fail if no provider can serve the chosen region. However, this may lead to a cost increase of approximately 10% over standard rates.
Key evidence
- AI Gateway now supports regional inference, enabling users to pin requests to either the US or EU data centers.
- If no model provider can serve the selected region, the request fails instead of being rerouted globally.
- Pinning a region may incur additional costs, typically around 10% above standard rates, which are passed through by AI Gateway.
Why it matters
The introduction of regional inference simplifies compliance for organizations with data residency requirements. By allowing users to confirm where their requests are processed, it reduces the complexity of managing regional routing across multiple providers. This can be particularly beneficial for companies operating in regulated industries that must adhere to strict data handling policies.
What to watch
📖 Reader Mode
~2 min readAI Gateway now supports regional inference. Set inferenceRegion on a request to pin it to the US or EU. Every model provider that supports the selected region handles it the same way. Inference runs there, and any data the provider keeps is stored there.
AI Gateway supports two pinned regions, plus global routing:
Region | Where inference runs |
|---|---|
| A US data center |
| An EU data center |
| Any region |
If no model provider can serve it, the request fails rather than running somewhere else. Every response reports the region that served it, so you can confirm where each request ran.
Here's a request pinned to the US with the AI SDK:
import { streamText } from 'ai';
const result = streamText({
model: 'moonshotai/kimi-k3',
prompt: 'Summarize this internal document.',
providerOptions: {
gateway: {
inferenceRegion: { scope: 'zone', geoRegion: 'us' },
},
},
});
Until now, teams with data residency or compliance requirements had to configure regional routing separately for every provider, with no reliable way to confirm where a request actually ran. Regional inference replaces that with a single field that behaves the same everywhere and a response that tells you where each request was served.
Filter the model list for models available in the US or EU, or read the regions array from /v1/models. Without inferenceRegion, requests route globally with no residency guarantee, so residency is opt-in.
Pinning a region can cost more. The provider sets the regional rate, often around 10% above standard, and AI Gateway passes it through with no markup. For per-provider overrides, response verification, pricing, and BYOK behavior, read the regional inference documentation.
— Originally published at vercel.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Vercel AI
See more →
The Agent Stack
The Agent Stack by Vercel AI provides essential building blocks for creating production-grade agents, enabling seamless integration across multiple AI models and secure operations. It features components like AI Gateway for model routing, Workflow SDK for durable execution, and Vercel Connect for scoped access, streamlining agent development and deployment across various platforms.

