AI Gateway
- Sixteen-plus providers behind one OpenAI-compatible endpoint
- Routing by workload value, cost, location and cluster load
- Automatic failover and traffic splitting
- Token-aware rate limits and budgets, enforced inline
Use Case 1
One product, four gateways: two EU regions, two US. Legal has one requirement: EU customer data stays in the EU. The business has two more: it can't go down, and it has a single budget.
A static failover chain would have crossed the border at 3am, correctly by its own logic, and nobody would have known until the audit.
Residency and resilience both have a claim on this request, and the precedence you set decides which one yields. That decision is recorded: which model served, which policy applied, and which one gave way.
The product line has a monthly cap, and all four gateways consult the same counter. An agent refused in Frankfurt doesn't get served in Ohio by retrying.
Group risk pulls a model. Every gateway stops routing to it, staged by region, with a record of when each one converged. That turns "we stopped using it in March" into a claim you can prove.
Use Case 2
Open-weight models on your own GPU clusters in two regions, with commercial APIs as overflow. The economics are inverted here: private capacity is a sunk cost, so the goal isn't to avoid using it. The goal is to use as much of it as you can before paying anyone per token.
Routing on live queue depth is the difference between shedding to a cluster with headroom and building a queue in front of one that's already full.
A backed-up cluster sheds to one with headroom, and the gateway close enough to each cluster to observe it is what makes that possible.
Every request that stays on capacity you already own is one you don't pay per token for. That inverts the usual cost policy, where cheap means a smaller model rather than a paid-for GPU.
Per-token pricing for the API providers, amortized GPU-hour for your own clusters. Attribution has to normalize both or the number you hand finance is fiction. This is the part nobody warns you about.
Traffic tagged sensitive is pinned to private models and refuses rather than overflowing. That is a policy decision with a record. It is not a routing preference.
Tetrate-hosted. Fastest way to get started. No infrastructure to manage.
Deploy inside your own infrastructure. Data stays in your perimeter. Required for regulated industries.
Deploy edge inference by zip code or service area, with localized model catalogs and data controls.
Run gateways in your AWS, Azure, or Google Cloud VPC, managed by one management plane.
Tetrate builds the Agent Router project (formerly Envoy AI Gateway), an open source AI gateway at the Agentic AI Foundation (AAIF). Agent Router Enterprise runs that same gateway. No proprietary data plane, no crippled community build, nothing held out of the project so we can sell it back.
An enterprise AI gateway is one control point between an organization's agents and its AI providers. The gateway enforces cost, access, and resilience policy on every request an agent makes. Tetrate Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane.
Enterprise AI agent routing coordinates many gateways as one system, so one budget, one failover policy, and one audit record apply across every region. In Agent Router Enterprise, a platform team declares policy once in the management console, and every gateway in scope converges on that policy. Gateways share budget and rate state, so an agent that one gateway refuses cannot get served by retrying through another.
AI agent traffic routing handles access control by sending every agent request through a gateway that checks who is asking and which tools that agent profile may call. Agent Router Enterprise ties that check to single sign-on and directory groups through your identity provider, such as Okta or Microsoft Entra ID. Step-up authorization built with Ory escalates only higher-risk actions to a person. An immutable audit log records every request.
Yes. Routing rules can require that failover stays inside a permitted region, so a request falls back to the next permitted in-region model. Agent Router Enterprise records which model served, which policy applied, and which policy was overridden, which turns a residency claim into per-jurisdiction evidence you can produce.
Yes. An enterprise AI gateway can route to hosted provider APIs and to models you run yourself. Routing on live queue depth sends each request to the GPU cluster with headroom. Agent Router Enterprise sends traffic to your private GPU capacity first and sends overflow to commercial APIs only when that capacity is full. Agent Router Enterprise also normalizes per-token and amortized GPU-hour costs into one chargeback.
FOR LEADERS MANAGING AI
Everything in Tetrate Agent Router Service, plus the visibility, attribution, and guardrails an engineering leader needs to run AI across multiple teams without losing track of what it costs or how it behaves.
Work with Tetrate forward-deployed engineers to design safe agent operations