Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

Choose your AI infrastructure once.

Agent Router runs where you run. It routes across commercial providers and the models you serve yourself.
And it manages AI traffic the way production infrastructure has to, whether you’re running one gateway or fifty.

Gateways for Control and
a Management Console for Consistent Execution

AI Gateway

  • Sixteen-plus providers behind one OpenAI-compatible endpoint
  • Routing by workload value, cost, location and cluster load
  • Automatic failover and traffic splitting
  • Token-aware rate limits and budgets, enforced inline

MCP Gateway

  • Curated tool catalogue per team and per agent profile
  • Synthetic profiles assembled from any set of servers
  • Complete tool-call logging across an agent chain
  • An agent can't reach further than the person it acts for

AI Guardrails

  • PII detection and redaction on request and response
  • Prompt filtering and risky-transaction blocking
  • Agent kill switch, scoped or global
  • Bring your own guardrails vendor alongside, or instead

The Management Console

  • Declare policy once; every gateway in scope converges
  • Shared budget and rate state across the fleet
  • Cost, traces and audit evidence aggregated in one place
  • Staged rollout, rollback and drift reconciliation

Use Case 1

A Customer Assistant That Can't Leave the EU

One product, four gateways — two EU regions, two US. Legal has one requirement: EU customer data stays in the EU. The business has two more: it can't go down, and it has a single budget.

The failover target is the next permitted model, not the next cheapest one

A static failover chain would have crossed the border at 3am, correctly by its own logic, and nobody would have known until the audit.

  • 01

    Failover stays inside the jurisdiction Failover stays inside
    the jurisdiction

    Residency and resilience both have a claim on this request, and the precedence you set decides which one yields. That decision is recorded — which model served, which policy applied, which one gave way.

  • 02

    One budget across four gateways One budget across
    four gateways

    The product line has a monthly cap, and all four gateways consult the same counter. An agent refused in Frankfurt doesn't get served in Ohio by retrying.

  • 03

    A withdrawal on Tuesday is a withdrawal everywhere A withdrawal on Tuesday
    is a withdrawal everywhere

    Group risk pulls a model. Every gateway stops routing to it, staged by region, with a record of when each one converged — so "we stopped using it in March" becomes a claim you can prove rather than assert.

Use Case 2

GPU Capacity You've Already Paid for

Open-weight models on your own GPU clusters in two regions, with commercial APIs as overflow. The economics are inverted here: private capacity is a sunk cost, so the goal isn't to avoid using it — it's to use as much of it as you can before paying anyone per token.

Static weights can't see saturation Static weights can't
see saturation

Routing on live queue depth is the difference between shedding to a cluster with headroom and building a queue in front of one that's already full.

  • 01

    Route on live capacity, not a weight set six months ago

    A backed-up cluster sheds to one with headroom, and the gateway close enough to each cluster to observe it is what makes that possible.

  • 02

    Commercial APIs are overflow, not the default

    Every request that stays on capacity you already own is one you don't pay per token for — which inverts the usual cost policy, where cheap means a smaller model rather than a paid-for GPU.

  • 03

    Two cost models, one chargeback

    Per-token pricing for the API providers, amortized GPU-hour for your own clusters. Attribution has to normalize both or the number you hand finance is fiction — and this is the part nobody warns you about.

  • 04

    Some classes never spill

    Traffic tagged sensitive is pinned to private models and refuses rather than overflowing. That's a policy decision with a record, not a routing preference.

Runs Wherever You Want

Managed Cloud

Tetrate-hosted. Fastest way to get started. No infrastructure to manage.

On-Premises

Deploy inside your own infrastructure. Data stays in your perimeter. Required for regulated industries.

Edge

Deploy edge inference by zip code or service area, with localized model catalogs and data controls.

Distributed across your footprint

Run gateways in your AWS, Azure, or Google Cloud VPC, managed by one control plane.

Agent Router product illustration

The Data Plane is the Part You Can Already Have

Tetrate builds Agent Router — the open-source AI gateway of the Agentic AI Foundation, formerly Envoy AI Gateway. Enterprise runs that same gateway. No proprietary data plane, no crippled community build, nothing held out of the project so we can sell it back.

Agent Router OSS
Agent Router Enterprise
Per Gateway / Vertical Scaling
Sixteen-plus model providers, one OpenAI-compatible API
Automatic failover and traffic splitting
Token-aware rate limits and budgets
pergateway
acrossthe fleet
MCP gateway and tool catalogue
Guardrail hooks — OpenTelemetry / Envoy data plane
Across More Than One Gateway
Declare policy once, converge everywhere
Shared budget and rate state, so a retry can't shop for a gateway that says yes
Cost aggregated and distributed across every gateway
Drift detection and continuous reconciliation
Staged rollout and rollback by region or business unit
Usage-aware routing across private GPU capacity
Access and Governance
API keys
SSO and directory groups — Okta, Entra ID, OIDC, SAML
Business-unit delegation and scoped consoles
Immutable audit log with a retention policy
Per-jurisdiction evidence for residency claims
Running It
Community releases
CVE-patched builds on a maintained release train
Air-gapped distribution
Model catalogue and pricing kept current
youmaintain it
wemaintain it
Support
GitHub and community Slack
24/7 support with an SLA
Named account team and upstream escalation

Start Fast. Scale When You're Ready.

Agent Router Enterprise

FOR LEADERS MANAGING AI

Everything in Service, plus the visibility, attribution, and guardrails an engineering leader needs to run AI across multiple teams without losing track of what it costs or how it behaves.

  • Cross-team cost attribution, showback, and chargeback
  • Admin controls — model and MCP access profiles by team
  • Runtime AI Guardrails — PII redaction and policy enforcement
  • Enterprise SSO — every request carries authenticated identity
  • Distributed deployment — cloud, on-prem, edge, or per-region

Need More Help?

Work with Tetrate forward-deployed engineers to design safe agent operations

Talk To An Engineer