An AI gateway sits between your coding tools and the models they call. Every request goes to the gateway first, and the gateway decides where it goes next.
Whether you need one depends on how many models and tools you’re dealing with. If one team uses Claude Code with Claude on Amazon Bedrock, IAM and SCPs already give you a solid EU residency control, and a gateway mostly adds work. Once you mix models, tools and teams, a gateway becomes the easiest place to enforce the residency control. In this post I’ll explain what a gateway does, when it’s worth it, and how to run one without creating a new problem.
What a gateway does for an EU residency control
AWS’s reference architecture for a multi-provider gateway, built on the open-source LiteLLM proxy, describes the main capabilities: one OpenAI-compatible interface to many providers, restricting access to specific models, budgets and rate limits per user, team or API key, and retry and fallback routing across providers (AWS: Guidance for Multi-Provider Generative AI Gateway). For EU residency, those translate into five practical benefits.
- One place to enforce the residency control. If the gateway only knows
about EU endpoints, such as Bedrock
eu.profiles, EU In-Region models or Mistral’s EU platform, developers can’t route around it by picking a Global model in their tool. - One endpoint for many models. Tools that accept an OpenAI-compatible base URL, such as Mistral Vibe and Cline, can reach Claude, Qwen or Gemma through the same address. Changing models becomes a central decision.
- Fallback that stays in the EU. When a model or Region hits its quota, the gateway can retry in another EU Region or on another approved model. AWS describes automatic failover between primary and secondary models as a core gateway pattern (AWS ML blog: Resilience patterns with Amazon Bedrock and LLM gateway). The rule is simple: every fallback must also be in the EU.
- Spend per team. Budgets and usage tracking per team or key help answer finance’s first question about AI coding.
- Whether you need one depends on how many models and tools you are dealing with. If one team uses Claude Code with Claude on Amazon Bedrock, IAM and SCPs already give you a strong EU residency control, and a gateway mostly adds work. Once you mix providers, tools and teams, a gateway can become a useful central enforcement point.
When you don’t need one, or it won’t help
- One tool, one model family, one AWS account. If your developers use Claude Code with Claude on Amazon Bedrock, pinned EU profiles plus IAM and SCPs already enforce the residency control, and CloudTrail gives you the evidence. A gateway adds a service to run without adding much control.
- Tools that route through their own servers. Cursor routes all requests through its own backend for prompt building, so a gateway behind Cursor doesn’t change that part of the path. Cursor also recommends using its hooks feature for security controls rather than custom gateways, which it says can add latency, rate limiting and compatibility issues (Cursor docs: Network configuration).
- Tools with a fixed model list. Kiro serves its own list of models. The EU residency control there comes from your enterprise profile and approved models, not from a gateway.
- Using Claude Code with non-Claude models. Claude Code can send its traffic through a gateway, but Anthropic states that it doesn’t support routing Claude Code to non-Claude models through any gateway (Claude Code docs: LLM gateways). For open-weight models, use a tool built for them, such as Mistral Vibe or Cline.
Running a gateway in the EU on AWS
AWS’s guidance deploys LiteLLM as containers on Amazon ECS or Amazon EKS, behind an Application Load Balancer protected by AWS WAF, with supporting services such as Amazon RDS, ElastiCache and Secrets Manager (AWS: Guidance for Multi-Provider Generative AI Gateway). To keep it inside an EU residency control, I’d add a few rules of my own:
- Deploy everything in one EU Region, including the database, cache and logs. The gateway sees every prompt and every line of code, so its storage needs the same residency as the models.
- Keep the entry point regional and private. The guidance offers either Route 53 or a CloudFront distribution in front of the gateway. For an EU residency control, I’d use the Region-based entry point and reach it over a VPN or private network rather than the public internet.
- Configure only EU model endpoints. Bedrock
eu.profiles or EU In-Region models, open-weight models in EU Regions, and Mistral’s EU API. Leave out Global profiles entirely, including as fallbacks. - Give the gateway its own IAM role that can only call those endpoints, and keep your SCPs in place underneath. If someone misconfigures the gateway, the account still can’t leave the EU.
Pointing the tools at it:
- Claude Code supports gateways through
ANTHROPIC_BASE_URL, or throughANTHROPIC_BEDROCK_BASE_URLwithCLAUDE_CODE_USE_BEDROCK=1for a Bedrock pass-through. SetCLAUDE_CODE_SKIP_BEDROCK_AUTH=1if the gateway handles AWS authentication (Claude Code docs: LLM gateways). Use it for Claude models only. - Mistral Vibe takes the gateway’s URL as a provider
api_base, and admins can enforce that provider for everyone (Mistral docs: Vibe admin config). - Cline takes the gateway’s URL in its OpenAI Compatible provider (Cline docs: OpenAI Compatible).
The risks to plan for
- The gateway is now a data processor. It sees every prompt, every file and every response. If it runs outside the EU, or logs to a bucket outside the EU, it undoes the residency control it was meant to protect.
- It’s a single point of failure. If the gateway is down, every coding tool behind it stops. Run it across several Availability Zones in your Region and monitor it like production.
- It’s another thing to secure and patch. It holds credentials for every model provider you use. Treat it as sensitive infrastructure, not a side project.
- Not every feature may be in the open-source version. Check LiteLLM’s own documentation for which features are in the open-source version before you commit to it.
My rule of thumb
One model family, one tool and good IAM: skip the gateway. Several models, several tools or several teams with separate budgets: a self-hosted gateway in your EU AWS account is worth the effort.
Thinking about a gateway?
Wolkn Minds can help decide whether a gateway adds enough control to justify the extra infrastructure, and design the EU deployment if it does.
