Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

Enterprise AI agent routing

Choose Your AI Agent Routing Infrastructure Once.

Tetrate Agent Router Enterprise runs where you run. It routes AI agent traffic across commercial providers and the models you serve yourself.
It applies policy to every request, whether you run one gateway or fifty.

The AI agent routing architecture: AI gateway, MCP gateway, guardrails, and a management console

AI Gateway

  • Sixteen-plus providers behind one OpenAI-compatible endpoint
  • Routing by workload value, cost, location and cluster load
  • Automatic failover and traffic splitting
  • Token-aware rate limits and budgets, enforced inline

MCP gateway for multi-agent tool access

  • Curated tool catalog per team and per agent profile
  • Synthetic profiles assembled from any set of servers
  • Complete tool-call logging across an agent chain
  • An agent can't reach further than the person it acts for

AI Guardrails

  • PII detection and redaction on request and response
  • Prompt filtering and risky-transaction blocking
  • Agent kill switch, scoped or global
  • Bring your own guardrails vendor alongside, or instead

The Management Console

  • Declare policy once. Every gateway in scope converges
  • Shared budget and rate state across the fleet
  • Cost, traces and audit evidence aggregated in one place
  • Staged rollout, rollback and drift reconciliation

Use Case 1

A customer assistant that can't leave the EU

One product, four gateways: two EU regions, two US. Legal has one requirement: EU customer data stays in the EU. The business has two more: it can't go down, and it has a single budget.

The failover target is the next permitted model rather than the next cheapest one

A static failover chain would have crossed the border at 3am, correctly by its own logic, and nobody would have known until the audit.

  • 01

    Failover stays inside the
    jurisdiction

    Residency and resilience both have a claim on this request, and the precedence you set decides which one yields. That decision is recorded: which model served, which policy applied, and which one gave way.

  • 02

    One budget across four
    gateways

    The product line has a monthly cap, and all four gateways consult the same counter. An agent refused in Frankfurt doesn't get served in Ohio by retrying.

  • 03

    A withdrawal on Tuesday is
    a withdrawal everywhere

    Group risk pulls a model. Every gateway stops routing to it, staged by region, with a record of when each one converged. That turns "we stopped using it in March" into a claim you can prove.

Use Case 2

GPU capacity you've already paid for

Open-weight models on your own GPU clusters in two regions, with commercial APIs as overflow. The economics are inverted here: private capacity is a sunk cost, so the goal isn't to avoid using it. The goal is to use as much of it as you can before paying anyone per token.

Static weights can't
see saturation

Routing on live queue depth is the difference between shedding to a cluster with headroom and building a queue in front of one that's already full.

  • 01

    Route on live capacity instead of a weight set six months ago

    A backed-up cluster sheds to one with headroom, and the gateway close enough to each cluster to observe it is what makes that possible.

  • 02

    Commercial APIs are overflow. Private capacity is the default.

    Every request that stays on capacity you already own is one you don't pay per token for. That inverts the usual cost policy, where cheap means a smaller model rather than a paid-for GPU.

  • 03

    Two cost models, one chargeback

    Per-token pricing for the API providers, amortized GPU-hour for your own clusters. Attribution has to normalize both or the number you hand finance is fiction. This is the part nobody warns you about.

  • 04

    Some classes never spill

    Traffic tagged sensitive is pinned to private models and refuses rather than overflowing. That is a policy decision with a record. It is not a routing preference.

AI agent routing across cloud, on-premises, and edge

Managed Cloud

Tetrate-hosted. Fastest way to get started. No infrastructure to manage.

On-Premises

Deploy inside your own infrastructure. Data stays in your perimeter. Required for regulated industries.

Edge

Deploy edge inference by zip code or service area, with localized model catalogs and data controls.

Distributed across your footprint

Run gateways in your AWS, Azure, or Google Cloud VPC, managed by one management plane.

Tetrate Agent Router Enterprise product illustration

An open source AI gateway with proven performance

Routing won't slow your agents down

  • Tested with real model traffic on GPUs, so the result matches how teams run agents in production
  • Stayed fast when many users hit the system at once, through a long endurance run
  • Measured independently by Broadcom and VMware under that real load
Read the latency benchmark

Built on the proxy that already runs the internet

  • The Agent Router project is open source and runs on Envoy, the same proxy under Agent Router Enterprise
  • Large companies such as Netflix and Lyft already use Envoy to carry live internet traffic every day
  • When security holes are found, fixes go out in public so every user can pick them up
Why Envoy scale matters

The data plane is the part you can already have

Tetrate builds the Agent Router project (formerly Envoy AI Gateway), an open source AI gateway at the Agentic AI Foundation (AAIF). Agent Router Enterprise runs that same gateway. No proprietary data plane, no crippled community build, nothing held out of the project so we can sell it back.

Agent Router OSS
Agent Router Enterprise
Per Gateway / Vertical Scaling
Sixteen-plus model providers, one OpenAI-compatible API
Automatic failover and traffic splitting
Token-aware rate limits and budgets
pergateway
acrossthe fleet
MCP gateway and tool catalog
Guardrail hooks via OpenTelemetry / Envoy data plane
Across More Than One Gateway
Declare policy once, converge everywhere
Shared budget and rate state, so a retry can't shop for a gateway that says yes
Cost aggregated and distributed across every gateway
Drift detection and continuous reconciliation
Staged rollout and rollback by region or business unit
Usage-aware routing across private GPU capacity
Access and Governance
API keys
SSO and directory groups via Okta, Entra ID, OIDC, SAML
Business-unit delegation and scoped consoles
Immutable audit log with a retention policy
Per-jurisdiction evidence for residency claims
Running It
Community releases
CVE-patched builds on a maintained release train
Model catalog and pricing kept current
youmaintain it
wemaintain it
Support
GitHub and community Slack
24/7 support with an SLA
Named account team and upstream escalation

Frequently asked questions about AI agent routing

An enterprise AI gateway is one control point between an organization's agents and its AI providers. The gateway enforces cost, access, and resilience policy on every request an agent makes. Tetrate Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane.

Enterprise AI agent routing coordinates many gateways as one system, so one budget, one failover policy, and one audit record apply across every region. In Agent Router Enterprise, a platform team declares policy once in the management console, and every gateway in scope converges on that policy. Gateways share budget and rate state, so an agent that one gateway refuses cannot get served by retrying through another.

AI agent traffic routing handles access control by sending every agent request through a gateway that checks who is asking and which tools that agent profile may call. Agent Router Enterprise ties that check to single sign-on and directory groups through your identity provider, such as Okta or Microsoft Entra ID. Step-up authorization built with Ory escalates only higher-risk actions to a person. An immutable audit log records every request.

Yes. Routing rules can require that failover stays inside a permitted region, so a request falls back to the next permitted in-region model. Agent Router Enterprise records which model served, which policy applied, and which policy was overridden, which turns a residency claim into per-jurisdiction evidence you can produce.

Yes. An enterprise AI gateway can route to hosted provider APIs and to models you run yourself. Routing on live queue depth sends each request to the GPU cluster with headroom. Agent Router Enterprise sends traffic to your private GPU capacity first and sends overflow to commercial APIs only when that capacity is full. Agent Router Enterprise also normalizes per-token and amortized GPU-hour costs into one chargeback.

Start fast. Scale when you're ready.

Agent Router Enterprise

FOR LEADERS MANAGING AI

Everything in Tetrate Agent Router Service, plus the visibility, attribution, and guardrails an engineering leader needs to run AI across multiple teams without losing track of what it costs or how it behaves.

  • Cross-team cost attribution, showback, and chargeback
  • Admin controls: model and MCP access profiles by team
  • Runtime AI Guardrails: PII redaction and policy enforcement
  • Enterprise SSO: every request carries authenticated identity
  • Distributed deployment: cloud, on-prem, edge, or per-region

Need More Help?

Work with Tetrate forward-deployed engineers to design safe agent operations

Talk To An Engineer