Open Source AI Gateway Features That Matter and 5 That Don't
TL;DR: The open source AI gateway features that matter in production are failover across providers, data planes that run inside your network, Model Context Protocol (MCP) tool policy, and agent-aware tracing. Step-up authorization for risky agent actions is an enterprise capability to look for on top of open source. Each one changes what happens during an outage, an audit, or an incident. Five commonly marketed features do little for sprawl, sovereignty, or debugging: chat UIs, basic prompt templates, single-cloud gateways, percentage fees on inference spend, and vanity log metrics. Tetrate Agent Router Enterprise adds a hosted management plane across every gateway, plus step-up authorization, to the open source Agent Router project (formerly Envoy AI Gateway) at the Agentic AI Foundation (AAIF).
Open source AI gateways publish long feature lists, and many of the items look the same from one project to the next. A feature list shows what a gateway does in a demo. The architecture decides what it does when a provider fails, a management plane goes offline, or an agent starts looping on a tool. This piece names the five features that change those outcomes for a platform team, the five that add work without solving sprawl, sovereignty, or debugging, and the three questions to ask any vendor before a pilot.
AI gateway requirements start with the shape of model traffic
AI traffic behaves differently from the request-response traffic most proxies were built for, and that difference sets the requirements for every feature that follows.
Streaming model responses strain proxies built for short calls
Streaming responses keep gateway connections open far longer than standard API calls do. Representational State Transfer (REST) APIs use short request-response cycles. Large language model (LLM) APIs stream responses over server-sent events (SSE), which keep a connection open until the response finishes. Each open stream keeps a connection busy for the length of the response, so a gateway sized for short REST calls needs more concurrent connections at the same request rate.
Cost behaves differently too. A single request can consume a few hundred tokens or tens of thousands, depending on the prompt and the response. Request-count rate limits cannot cap spend when request size varies that much, so an AI gateway needs token-based limits in the request path.
An open source data plane puts one enforcement layer across teams
AI infrastructure sprawl starts when more than one team, model, or provider runs in production, and it grows with every new one. An open source data plane gives platform teams one place to enforce policy across teams, regions, and providers. The open source data plane also avoids lock-in to a single cloud or software as a service (SaaS) vendor.
The trade-off is ownership. With open source alone, the platform team owns upgrades, failures, and fleet management across every gateway it runs. An enterprise license can take on that fleet layer. The deciding question is whether the platform team can run the gateway reliably at production scale without becoming the bottleneck for every other team.
Three questions separate architecture from feature lists
Three questions decide whether a gateway feature matters in production. Does the data plane keep routing when the management plane is unreachable? Does governance cover agent tool calls as well as model calls? Has a third party measured the gateway’s overhead under real inference load?
| Question | Why it matters | What to ask the vendor |
|---|---|---|
| Does the data plane survive a management plane outage? | Decides whether the gateway is a single point of failure | Show the failure-mode test |
| Does governance cover agent tool calls through MCP? | Agents increasingly call tools, and MCP is one of the most rapidly weaponized attack surfaces in agentic AI | Can the gateway show which tools an agent called? |
| Has a third party measured overhead under real inference load? | A vendor’s own benchmark is a claim, while an outside test is evidence | Where is the independent test report? |
Five features change outcomes in production
Five features change how a gateway behaves when something fails. Two of them keep traffic flowing through outages, and three set how far governance reaches into agent behavior.
Automatic failover limits a provider outage to one route
Automatic failover keeps a provider outage from reaching every team. When one frontier provider fails and no backup route exists, every agent that depends on that provider stops at the same time. Multi-provider failover moves a request to the next provider in an ordered list when the first one returns a timeout, a rate-limit error, or a server error. Model aliases and traffic splitting sit in the same routing layer, so a team can switch models safely without changing agent code.
Test failover before trusting it. Simulate a provider outage in a pilot and confirm that requests move to the backup provider within the same request.
Decoupled data planes keep prompts inside the customer’s network
Regulated teams that must keep prompts inside their own network need a data plane that runs there. Tetrate pairs one central management plane with gateways the customer can run in its own cloud, on-premises, at the edge, or per region. Tetrate can also run the gateways as a fully managed deployment.
The property to test is data plane autonomy. The data plane keeps its last configuration and keeps routing traffic when it cannot reach the management plane. Configuration changes wait until the link returns. With a self-hosted data plane, prompt and response text stays out of Tetrate’s systems unless request logging keeps the text, which an admin can turn off. The management plane stores configuration and receives telemetry. The Tetrate security overview covers how the outbound-only connection between the planes works.
| Deployment model | Management plane | Data plane | Where prompts flow |
|---|---|---|---|
| Fully managed | Tetrate-hosted | Tetrate-operated | Through Tetrate-operated infrastructure |
| Self-hosted data plane | Tetrate-hosted | Customer cloud, on-premises site, or edge | Out of Tetrate’s systems unless request logging keeps the text, which an admin can turn off |
MCP tool policy controls which tools each agent can reach
Governance that stops at model calls leaves agent tool calls unchecked. An MCP gateway sits between AI agents and the MCP servers they call. An organization connects each server once, then decides which tools each agent can use.
In Agent Router Enterprise, an MCP profile is a curated set of tools drawn from approved servers, with its own authentication. Each profile exposes only its approved tools, so an agent cannot call a tool outside its profile. The gateway attaches the user’s identity and an audit record to every tool call. A server removed from the catalog disappears at once. The MCP gateway guide walks through catalogs, profiles, and credentials.
Ask whether the gateway can answer one audit question: which tools did this agent call, and on whose behalf?
Agent-aware tracing shows what an agent did with its access
Agent debugging stalls when logs sit in separate provider dashboards. Reconstructing a multi-agent chain failure from those dashboards can take days. Agent-aware tracing logs each model call and tool call across an agent chain. The gateway exports those traces over OpenTelemetry to the team’s own observability tools, such as Grafana or Datadog. The same reconstruction then takes minutes, because every step sits in one trace.
The pilot check is short. Trigger a failure in a two-step agent chain and confirm that the trace shows the tool call as well as the model call.
Step-up authorization escalates only high-risk agent actions
Agents need to run routine actions without waiting for approval, while high-risk actions need a person to sign off. Agent Router Enterprise handles both cases with step-up authorization, built with Ory, an identity and access management company.
Many MCP gateways decide only which tools an agent can see or call. With Ory, the gateway can also enforce policy on the parameters of each live tool call. When a call crosses a defined risk threshold, the gateway pauses the request, triggers approval through Ory, and records the approval path for audit. Ask whether the gateway can enforce policy on what a tool call contains.
The open source Agent Router project covers some of these features on its own. Agent Router Enterprise adds the rest.
| Feature | What it does in production | In the open source project? |
|---|---|---|
| Automatic failover | Limits a provider outage to one route | Yes, provider fallback |
| Decoupled data planes | Keeps routing through a management plane outage, with prompts kept in the customer’s network when self-hosted | Partly, each gateway is self-hosted. Agent Router Enterprise adds one hosted management plane across every data plane |
| MCP tool policy | Limits each agent to approved tools | Partly, OAuth authentication and fine-grained tool access control. Agent Router Enterprise adds profiles tied to single sign-on identity, a managed catalog, and audit records |
| Agent-aware tracing | Puts model calls and tool calls in one trace | Yes, OpenTelemetry tracing |
| Step-up authorization | Sends high-risk actions to a person | No, Agent Router Enterprise with Ory |
Five commonly marketed features don’t solve sprawl, sovereignty, or debugging
Chat UIs, prompt templates, single-cloud gateways, percentage fees, and vanity log metrics fill out a feature list without changing how the gateway behaves when something fails.
Chat UIs and prompt templates sit outside the request path
Chat UIs and basic prompt templates help developers, but neither one governs traffic. A chat UI is useful in a demo. If the chat interface calls models outside the gateway’s request path, it adds a second route to the providers. Traffic on that route skips the budgets, tool policy, and audit records the gateway enforces. A prompt template stores prompt text, but it does not enforce a budget, record a tool call, or cap token use.
Keep templates in the application or a prompt management tool. Judge the gateway on the policy it enforces on live traffic, and check that every chat interface in use sends its requests through that gateway.
Architecture and pricing traps
Single-cloud gateways and percentage fees on inference spend are not open source gateway features at all. They are architecture constraints and pricing structures that belong in a separate evaluation from the feature list.
Single-cloud gateways and percentage fees tie the gateway to one vendor’s terms
A single-cloud gateway ties policy to one cloud, and a percentage fee ties cost to inference spend. A cloud provider’s native gateway is built around that cloud’s identity, billing, and model catalog. Routing across other clouds or private models adds configuration work the platform team has to own. If the organization runs across clouds or needs on-premises data planes for residency, a single-cloud gateway limits where policy can run.
Pricing creates a similar constraint. Some vendors add a percentage fee on top of inference spend, so the gateway bill grows with every token the organization uses. Procurement cannot predict that cost until usage settles. Ask whether the vendor’s revenue grows with your inference spend, and whether the gateway can run in every cloud you use.
Vanity log metrics count traffic without attributing spend
Request counts and latency percentiles do not tell finance which team spent what. A gateway answers that question when it attributes each token to a person, team, agent, or project. Rate limits then cap a looping agent’s token use inline at the gateway. Budget policies block the caller or reroute to a cheaper model a few minutes after spend crosses the limit. Each budget can start as Watch spend, which raises an alert without blocking anything.
During a pilot, pick one week of traffic. Ask the gateway for spend by team and by agent. If the answer needs a spreadsheet export and manual tagging, the metrics describe traffic without explaining cost.
| Feature | Why it falls short |
|---|---|
| Chat UIs | Adds a route that can skip policy enforcement |
| Basic prompt templates | Stores text without enforcing policy |
| Single-cloud gateways | Limits policy to one cloud’s services |
| Percentage fees on inference spend | Gateway cost rises with inference spend |
| Vanity log metrics | Counts traffic without per-team or per-agent attribution |
Independent tests and open source code back the data plane
The data plane under a gateway decides how it behaves under load, so its test record and its history matter as much as the feature list.
A third-party test measured about 2 ms of gateway overhead
Overhead claims count as evidence only when an outside party runs the test. The Broadcom/VMware test of the Agent Router project (then named Envoy AI Gateway) measured about 2 ms of added latency per request under real model inference on four NVIDIA H100 graphics processing units (GPUs). The test used 124 GB of data with over 20,000 unique sessions from public agent and coding benchmarks. It included a three-hour run at 190 concurrent users. Overhead stayed flat up to 224 concurrent users, where the GPUs reached their limit first.
Agent Router Enterprise runs the same data plane that Broadcom/VMware tested. Ask any vendor for an outside test report of this kind.
Envoy’s history and the AAIF move give the data plane a public record
The Agent Router project runs on Envoy, a proxy with a long production record. Envoy was originally created at Lyft and is a graduated Cloud Native Computing Foundation (CNCF) project that has run in production since 2016. In 2018, Lyft’s vice president of engineering said Envoy carried all of Lyft’s edge and service-to-service traffic, at millions of requests per second. Ask what a gateway’s data plane is built on and how long that base has run in production.
Tetrate co-created the Agent Router project with Bloomberg. Version 1.0 shipped in June 2026. On September 10, 2026, the project maintainers donated Envoy AI Gateway to AAIF and renamed it the Agent Router project. Agent Router Enterprise adds fleet management, hybrid deployment, dashboards, cost controls, MCP access control, and support on top of the open source project.
Conclusion
Architecture decides how much value an AI gateway delivers in production. Failover, decoupled data planes, MCP tool policy, agent-aware tracing, and step-up authorization change what happens during an outage, an audit, or an incident. The other five sit outside the request path, tie the gateway to one vendor’s terms, or count traffic without attributing spend. A platform team can ask any vendor the three architecture questions, then test the answers in a fast-track evaluation before committing to a gateway.
Start a fast-track evaluation: first routed request under five minutes, nothing to install. Bring your own traffic to test failover, MCP tool policy, and spend attribution.
Agent Router Enterprise
FAQs
What AI gateway features matter most for regulated industries?
Data residency matters most for regulated industries. Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane. With a self-hosted data plane, prompt and response text stays out of Tetrate’s systems unless request logging keeps the text, which an admin can turn off. MCP tool policy comes second, because auditors ask which tools an agent reached and on whose behalf.
How does an AI gateway govern agent tool calls?
An AI gateway governs tool calls through an MCP gateway. In Agent Router Enterprise, each MCP profile exposes only approved tools. The gateway attaches the user’s identity and an audit record to every call. Step-up authorization built with Ory sends high-risk actions to a person for approval.
What’s the difference between an AI gateway and an API gateway?
An API gateway was built for short request-response Hypertext Transfer Protocol (HTTP) traffic. An AI gateway is built for streamed model responses, token-based limits, model routing with failover, and governance of agent tool calls through MCP. Some API gateway vendors now add AI plugins that cover part of this list.
When should a team replace a home-grown AI gateway?
Replace a home-grown gateway when its maintenance becomes a distributed systems problem. Every provider API change lands on the platform team, and traffic the gateway doesn’t route stays ungoverned. One large fintech customer replaced its home-grown gateways with Agent Router Enterprise, which freed the engineering time spent maintaining them.
Does the Agent Router Enterprise management plane create a single point of failure?
No. Agent Router Enterprise gateways keep routing traffic when they cannot reach the management plane. The management plane stores configuration and serves the admin consoles, so an outage pauses configuration changes until the link returns.
Key terms glossary
Agent Router project: The open source data plane formerly called Envoy AI Gateway, now at AAIF. It runs on the Envoy proxy. It covers provider fallback, OpenTelemetry tracing, and fine-grained tool access control with OAuth authentication.
Agent-aware tracing: Tracing that logs each model call and tool call across an agent chain. Teams use it to reconstruct a multi-agent failure in their own observability tools.
AI infrastructure sprawl: The condition that starts when more than one team, model, or provider runs in production without a single enforcement layer, and grows with every new addition.
Automatic failover: Routing that moves a request to the next provider in an ordered list when the first provider fails.
Data plane: The component that carries live AI traffic between applications and providers. Tetrate operates the data plane in a fully managed deployment, or the customer runs it in its own cloud, on-premises, or at the edge.
Envoy: An open source proxy that started at Lyft and is a graduated CNCF project. The Agent Router project runs on Envoy.
Management plane: The component that stores configuration, serves the admin consoles, and receives telemetry. Tetrate hosts the management plane in every deployment model.
MCP (Model Context Protocol): A protocol for connecting AI agents to tools and data sources. An MCP gateway controls which tools each agent can call and records every call.
Sovereignty (data sovereignty): The requirement that live prompts and responses stay inside the customer’s own network and never transit a vendor’s shared infrastructure.
Step-up authorization: A policy mechanism that lets routine agent actions run without friction while it sends high-risk actions to a person for approval.
Vanity log metrics: Request counts and latency percentiles that describe traffic volume without attributing spend to a specific team, agent, or project.
Related: AI gateway capabilities · How to choose an open source AI gateway · Open source AI gateway buyer’s guide
See the full 2026 enterprise AI gateway comparison.
Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.