AI Gateway
- Sixteen-plus providers behind one OpenAI-compatible endpoint
- Routing by workload value, cost, location and cluster load
- Automatic failover and traffic splitting
- Token-aware rate limits and budgets, enforced inline
Use Case 1
One product, four gateways — two EU regions, two US. Legal has one requirement: EU customer data stays in the EU. The business has two more: it can't go down, and it has a single budget.
A static failover chain would have crossed the border at 3am, correctly by its own logic, and nobody would have known until the audit.
Residency and resilience both have a claim on this request, and the precedence you set decides which one yields. That decision is recorded — which model served, which policy applied, which one gave way.
The product line has a monthly cap, and all four gateways consult the same counter. An agent refused in Frankfurt doesn't get served in Ohio by retrying.
Group risk pulls a model. Every gateway stops routing to it, staged by region, with a record of when each one converged — so "we stopped using it in March" becomes a claim you can prove rather than assert.
Use Case 2
Open-weight models on your own GPU clusters in two regions, with commercial APIs as overflow. The economics are inverted here: private capacity is a sunk cost, so the goal isn't to avoid using it — it's to use as much of it as you can before paying anyone per token.
Routing on live queue depth is the difference between shedding to a cluster with headroom and building a queue in front of one that's already full.
A backed-up cluster sheds to one with headroom, and the gateway close enough to each cluster to observe it is what makes that possible.
Every request that stays on capacity you already own is one you don't pay per token for — which inverts the usual cost policy, where cheap means a smaller model rather than a paid-for GPU.
Per-token pricing for the API providers, amortized GPU-hour for your own clusters. Attribution has to normalize both or the number you hand finance is fiction — and this is the part nobody warns you about.
Traffic tagged sensitive is pinned to private models and refuses rather than overflowing. That's a policy decision with a record, not a routing preference.
Tetrate-hosted. Fastest way to get started. No infrastructure to manage.
Deploy inside your own infrastructure. Data stays in your perimeter. Required for regulated industries.
Deploy edge inference by zip code or service area, with localized model catalogs and data controls.
Run gateways in your AWS, Azure, or Google Cloud VPC, managed by one control plane.
Tetrate builds Agent Router — the open-source AI gateway of the Agentic AI Foundation, formerly Envoy AI Gateway. Enterprise runs that same gateway. No proprietary data plane, no crippled community build, nothing held out of the project so we can sell it back.
FOR LEADERS MANAGING AI
Everything in Service, plus the visibility, attribution, and guardrails an engineering leader needs to run AI across multiple teams without losing track of what it costs or how it behaves.
Work with Tetrate forward-deployed engineers to design safe agent operations