How to Choose an Open Source AI Gateway: A Platform Leader's Evaluation Framework
TL;DR: Evaluate an open source AI gateway on four criteria, in this order: separation of the data plane from the management plane, governance of MCP tool calls as well as model calls, independently tested performance under production load, and deployment options for data residency. The first test is easy to run during a pilot: disconnect the management plane and check whether live traffic keeps flowing. Tetrate Agent Router Enterprise keeps routing traffic when its management plane is unreachable. In the Broadcom/VMware test of the Agent Router project (then named Envoy AI Gateway), the data plane under Agent Router Enterprise added about 2 ms of overhead per request.
An AI gateway evaluation should start with one question: does the data plane keep routing traffic when the management plane goes down? A feature checklist comes later.
This evaluation framework helps platform leaders compare open source AI gateways against a cross-org governance mandate. It covers four criteria in order, the test that proves each one, and how to run those tests in a four-week pilot. It uses the Agent Router project (formerly Envoy AI Gateway) at the Agentic AI Foundation (AAIF) and Agent Router Enterprise as the reference architecture.
Assessing gateways for cross-org AI mandates
A cross-org mandate changes the evaluation criteria, because the gateway becomes the single enforcement point for every team, model, and provider. Without that enforcement point, three failure points appear:
- Availability: a single frontier provider outage takes down agents across every team that relies on it, because no failover routes traffic to another provider.
- Cost visibility: token spend lands on a single API key, with no per-team, per-agent, or per-project breakdown.
- Onboarding: new developers have no standard path to model access, so teams set up their own keys and billing arrangements out of leadership’s view.
Weighing open source against commercial gateways
Open source gives the platform team control over the code and avoids lock-in to one vendor. It also shifts the operational burden to the platform team, which now owns upgrades, scaling, and support. Commercial gateways reduce that burden. Some add their own costs: a cloud provider gateway ties the organization to one cloud, and vendors such as OpenRouter take a share of inference spend.
Agent Router Enterprise is a commercial product built on the open source Agent Router project, so the data plane code stays open while Tetrate runs the management plane. The criteria below apply to any option, whether the team runs an open source project itself or buys a product built on one.
Scoring gateways against four criteria
Each criterion has a test a team can run during a proof of concept (PoC). Run them in this order, because a gateway that fails the first test cannot serve as the enforcement point for every team.
| Criterion | Test to run | Pass signal |
|---|---|---|
| 1. Data plane and management plane separation | Disconnect the management plane during the PoC | Live traffic keeps flowing |
| 2. Governance of tool calls as well as model calls | Block one MCP tool for one team, then call it | The gateway blocks the call and records it |
| 3. Independent performance evidence | Ask who ran the benchmark, against what backend, and for how long | An outside party ran it against real model inference |
| 4. Deployment options for data residency | Ask where the data plane can run | The customer’s cloud, on-premises, edge, or per region |
Testing gateway architecture for production
Architecture decides whether a gateway keeps routing traffic under production load. Two tests separate gateways ready for production from proxies that only work in a demo.
Separating the data plane from the management plane
The management plane handles policy, configuration, and administration. The data plane carries live traffic and enforces policy on each request. The critical test asks whether the data plane keeps routing traffic when the management plane is unreachable.
Gateways that run policy and traffic in a single process stop routing when that process restarts or upgrades. Agent Router Enterprise separates the two: Tetrate hosts the management plane, and the gateways keep carrying traffic when the management plane is unreachable. The gateways open an outbound connection to the management plane for configuration only, so no inbound firewall openings are needed. The trade-off is that configuration changes wait until the link is back.
If the data plane depends on the management plane for every request, a management plane outage stops AI traffic for every team at once. To check the separation during a PoC, disconnect the management plane and confirm that live traffic continues. The AI gateway security guide covers this split in more detail.
Routing around provider outages
Automatic failover across providers limits the damage of a provider outage. When one provider fails, traffic shifts to another, so the outage takes down a single route while other routes keep serving traffic. In Agent Router Enterprise, model aliases and routing rules handle this failover today.
The evaluation should ask how failover is configured, how fast it triggers, and whether teams can test it before production. The guide to switching models safely covers aliases, traffic splitting, and fallback.
Governing agent tool calls as well as model calls
Gateways differ most in how far governance reaches past model calls. The tests in this section check whether a gateway can see and control what agents do.
Enforcing policy across teams and providers
One policy layer applies the same cost, access, and resilience rules to every team. The platform team sets approved models, guardrails, and budgets per group, so access can open widely without per-request review. The guide to controlling team model access shows how projects, roles, and budgets fit together.
Rate limits apply in the request path. Budgets default to Watch spend, which alerts without blocking, and Hard stop blocks a group that passes its cap a few minutes after the spend lands. Ask any vendor the same question: does a budget block spend, or does it only report it?
Auditing agent tools and MCP calls
Some gateways govern only model calls, so they can answer “which model was called” but not what an agent did once it had model access. That gap covers which tools the agent called, what data it touched, and which of its actions needed a closer look.
As agents call more tools through the Model Context Protocol (MCP), a gateway that governs only model calls leaves the tool-calling layer ungoverned. A curated tool catalog with per-profile authentication lets security audit exactly what an agent reached and when. Tetrate’s MCP gateway governance approach shows how this works in practice.
Managing high-risk AI access controls
Step-up authorization, built with Ory, an identity and access management platform, lets routine agent actions proceed without friction while higher-risk actions require extra approval. When an agent attempts a high-risk transaction or accesses sensitive data, the gateway pauses the request and triggers a step-up flow. A human approves the action, and Ory issues a short-lived token scoped to that one action.
The evaluation should ask two questions. Can the gateway tell routine agent actions from high-risk ones? Can it require human approval for the high-risk set without blocking everything else?
Finding shadow AI outside the gateway
Shadow AI comes from staff who use AI outside the platform team’s plan. In a 2025 global study, about half of employees who use AI at work said they had uploaded sensitive company information into public AI tools. Analysts, marketers, and ops staff can also build agents in low-code or chat tools without managing keys or budgets.
A gateway sees only the traffic routed through it, so the evaluation should ask how the gateway finds model calls that bypass it. An approved path that is easy to use gives these builders less reason to work around the gateway. For tool coverage, see which AI tools a gateway governs.
Checking independent evidence of production readiness
Architecture and governance set the baseline. Independent evidence shows whether the gateway and its open source project perform under real load over time.
Reading vendor performance benchmarks
Ask each vendor whether its benchmark ran against real model inference or a mock backend, how long the run lasted, and who ran the test. A vendor’s own benchmark is a vendor claim. A test that an outside party ran and published is stronger evidence. Rerun the load test on every gateway upgrade, and compare throughput as well as error rates.
Independent numbers exist for some open source gateways. In one third-party test, the LiteLLM proxy added a median 42 ms per request against the Gemini API. Its 95th-percentile (p95) overhead was 58 ms, with a maximum of 78 ms.
The Envoy proxy that the Agent Router project runs on already carries large production loads. By 2017, Lyft’s Envoy mesh handled over 2 million requests per second across more than 10,000 hosts. Netflix later adopted Envoy for traffic between its services. These are fleet-wide figures for the Envoy proxy, so the Broadcom/VMware test below is the better guide to per-request overhead.
Checking for independently run performance tests
The Broadcom/VMware test of the Agent Router project (then named Envoy AI Gateway) measured about 2 ms of gateway overhead per request, about 0.01% of end-to-end latency. The test ran real model inference on four NVIDIA H100 GPUs, including a three-hour endurance run at 190 concurrent users. Gateway overhead stayed flat up to the saturation point at 224 concurrent users, where the GPUs reached their limit first.
The test also surfaced a configuration lesson. Early runs showed 3.5 to 6.2 seconds of added latency, which came from Linux CPU throttling. Pinning Envoy’s worker threads to the container’s CPU limit removed that latency.
The Broadcom/VMware test used a different setup from the third-party LiteLLM test, so the gap between 2 ms and 42 ms shows direction only. A platform team can still show leadership a figure that an outside party measured.
Assessing open source project health
Evaluate open source project health by commit frequency, contributor diversity, release cadence, and governance structure. The Agent Router project reached v1.0 in June 2026, and version 1.1 followed in August. The project maintainers donated Envoy AI Gateway to the AAIF in September 2026. The Linux Foundation formed the AAIF in December 2025 as a neutral home for open agentic AI projects.
The underlying proxy matters for operational maturity. Lyft open sourced Envoy in September 2016, and Envoy is a Cloud Native Computing Foundation (CNCF) graduated project, so a gateway built on Envoy inherits that operational history.
Matching deployment and cost models to the organization
Deployment options decide where data lives, and the pricing model decides the total cost of ownership over time. Both belong in the evaluation.
Architecting for data residency needs
Many regulated organizations require prompts and data to stay inside their own network. A SaaS-only gateway routes traffic through the vendor’s infrastructure, so it cannot meet that requirement.
The evaluation should ask whether the data plane can run in the customer’s own cloud, on-premises, at the edge, or per region. If the answer is no, regulated workloads will need a different deployment model. Placing gateways by region or zip code also cuts latency. One management plane still manages every gateway.
Managing multi-cloud AI workloads
A single-cloud gateway limits the organization’s ability to move workloads on cost, capacity, performance, or sovereignty. Organizations that run AI across more than one cloud need an independent control point that spans clouds, direct providers, and private models.
Cloud provider gateways such as AWS Bedrock and Azure API Management are a strong fit when one cloud is already the center of the AI architecture. Agent Router Enterprise is built for the multi-cloud case, with one policy layer across every provider the organization uses.
Choosing between managed and self-hosted deployment
The deployment spectrum runs from fully managed by the vendor, through hybrid, to fully self-hosted. In a hybrid model, the vendor runs the management plane and the customer runs the data plane. Managed deployment reduces operational burden. Self-hosted deployment keeps data inside the network.
Agent Router Enterprise is built on the Agent Router project, an open source data plane at the AAIF that runs on the Envoy proxy. Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location. One management plane manages all of the gateways, and Tetrate hosts that management plane in every Agent Router Enterprise deployment. The deployment model comparison shows how each option maps to different data residency requirements.
Comparing pricing and support models
Pricing models change how predictable the cost is. Vendors such as OpenRouter take a share of inference spend, so the bill grows with usage. Tetrate licenses Agent Router Enterprise by feature set, deployment footprint, and support tier, with no per-token fee, so procurement can predict cost up front.
Support tiers define what happens when the gateway fails outside business hours. The evaluation should ask for response times, severity definitions, and escalation paths. Agent Router Enterprise offers a standard tier and a premium tier with 24x7 one-hour response on severity-one issues.
Identifying early compliance red flags
Regulated industries should flag any gateway that lacks these items:
- Security attestations and certifications such as SOC 2 Type II or ISO 27001, which cover controls the vendor operates, verified through an audit report
- GDPR compliance documentation
- Configurable data capture, which the customer controls
Agent Router Enterprise holds a SOC 2 Type II report, and Tetrate is certified to ISO 27001 and is GDPR compliant. Tetrate shares the full SOC 2 report under NDA. Each customer chooses whether the gateway stores the full request, only counts and cost, or nothing. The customer can change that setting at any time.
Validating gateway fit in a four-week pilot
A PoC tests the framework against real workloads, so define success metrics before it starts across five areas:
- Request throughput under load
- Memory footprint
- Failover behavior
- Governance coverage, including MCP tool calls
- Deployment options
A mutual success plan between the platform team and the vendor keeps the PoC focused. Tetrate’s four-week PoC for Agent Router Enterprise is a no-cost licensed evaluation that requires an NDA.
Stress testing means running real inference under concurrent load, measuring latency overhead, and testing failover across providers. Include an endurance run at high concurrency, and keep measuring gateway overhead up to the point where the backend saturates.
Security audits tool calls and data residency, while finance audits cost attribution and budget enforcement, so both belong on the PoC success plan from day one. Showback dashboards that attribute spend per team, per agent, and per project give finance the breakdown it needs. Security can use the MCP gateway to audit which tools agents reached during the pilot.
Architecture decides whether a gateway is ready for production
Evaluate open source AI gateways on architecture first, governance depth second, independent performance evidence third, and deployment options fourth. The architecture test comes first because the gateway becomes the enforcement point for every team, model, and provider. Agent Router Enterprise, built on the Agent Router project at the AAIF, keeps routing traffic when its management plane is unreachable. To run that test yourself, use the four-week PoC with the data plane in your own environment and disconnect the management plane in week one.
Start a fast-track evaluation: first routed request under five minutes, nothing to install. You can also request pricing details.
Agent Router Enterprise
FAQs
How is an AI gateway different from an API gateway?
An AI gateway governs AI requests across models, providers, and teams. It adds AI-specific controls such as token budgets, model routing, and MCP tool governance. An API gateway handles general API traffic without these controls. For a detailed breakdown, see this comparison of AI gateway types.
What’s the difference between open source and self-hosted?
Open source means the code is publicly available. Self-hosted means the gateway runs in the customer’s own cloud, on-premises, or at the edge. A gateway can be open source and still be fully managed by a vendor. The deployment model comparison covers the trade-offs.
When should we build vs. buy an AI gateway?
Build when governance requirements are unique and tightly coupled to internal systems. Buy when they are common across enterprises or risky to maintain. Version one of a homegrown gateway is quick to build, but keeping up with provider changes across regions turns the gateway into a distributed systems problem.
How do we evaluate performance claims independently?
Ask who ran the test, against what backend, and for how long. The Broadcom/VMware test of the Agent Router project (then named Envoy AI Gateway), the data plane behind Agent Router Enterprise, measured about 2 ms of gateway overhead per request under real GPU inference. Label vendor-published figures as vendor claims when presenting them to leadership.
What compliance requirements affect AI gateway selection?
Regulated industries often check security certifications, GDPR compliance, and data residency first. Data residency requirements may rule out SaaS-only gateways. Agent Router Enterprise holds a SOC 2 Type II report, and Tetrate is certified to ISO 27001 and is GDPR compliant. Agent Router Enterprise gateways can run inside the customer’s own network.
Key terms glossary
Data plane: The component that carries live traffic and enforces policy on each request. In this framework, the first test checks whether the data plane keeps routing when the management plane is unreachable.
Management plane: The component that handles policy, configuration, and administration. In Agent Router Enterprise, Tetrate hosts the management plane, and gateways can run fully managed or in the customer’s environment.
MCP (Model Context Protocol): A protocol for connecting agents to tools. An MCP gateway gives security teams a curated tool catalog with per-profile authentication, so they can audit every tool call an agent makes.
Step-up authorization: A mechanism that lets routine agent actions proceed while higher-risk actions require extra approval before they run.
Data residency: The requirement that data stays inside a specific geographic or organizational boundary. Running data planes in the customer’s own cloud, on-premises, or per region keeps prompts inside that boundary, as long as request logging is set to keep only counts and cost or nothing.
Related: Open source AI gateway buyer’s guide · Open source AI gateway features that matter · Open source vs commercial AI gateway · AI gateway capabilities
See the full 2026 enterprise AI gateway comparison.
Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.