AI Gateway Capabilities: 9 Features That Actually Matter
TL;DR: Most AI gateways list the same features, so the useful test is how each gateway behaves under production load. This guide covers nine capabilities and gives a test for each capability. Tetrate Agent Router Enterprise is built on the Agent Router project (formerly Envoy AI Gateway), an open source data plane at the Agentic AI Foundation (AAIF) that runs on the Envoy proxy. An independent Broadcom/VMware test ran that data plane with real GPU inference and bursty agent traffic for three hours, and measured about 2 milliseconds of gateway overhead. Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane.
A platform team evaluating AI gateways faces the same problem with every vendor: the feature lists look alike, but production behavior differs. Gartner forecasts AI gateway spending will grow 70.9% in 2027, from $251 million to $429 million. This guide covers nine capabilities that change how a gateway behaves under production load, with a test for each one you can apply to any vendor.
AI gateways handle traffic patterns that general API gateways were not built for
An AI gateway enforces policy on AI traffic patterns such as token budgets, streaming, and agent tool calls. General-purpose API gateways were built for short-lived Representational State Transfer (REST) traffic. Several now add token limits and streaming support as extra policies on top of that original design.
Kong adds AI routing, token limits, and caching as plugins on its API gateway. Tetrate Agent Router Enterprise is purpose-built for AI traffic patterns. Tetrate Agent Router Enterprise is built on the Agent Router project (formerly Envoy AI Gateway), an open source data plane at the Agentic AI Foundation (AAIF) that runs on the Envoy proxy.
| Capability | AI gateway | General API gateway |
|---|---|---|
| Routing | Model-aware, with fallback across providers | Connection-based |
| Cost control | Per-token budgets, chargeback | Request-based metering |
| Streaming | Server-Sent Events (SSE) token delivery | Standard HTTP responses |
| Agent governance | MCP tool access control | Request-level auth and logging |
| Multimodal | Text, image, audio | Standard REST payloads |
REST-era gateway settings break under AI traffic at scale
AI inference breaks three assumptions behind REST traffic:
-
Token streaming keeps connections open for the full generation time, which can run for many seconds
-
Response sizes are hard to predict because output length is not known before generation starts
-
Latency depends on model load, prompt length, and provider capacity. Serving research splits it into time to first token and time per output token
A gateway tuned for REST traffic applies the same timeouts and connection limits to a 30-second streaming inference call as to a 50-millisecond database lookup. The result can be dropped streams, exhausted connection pools, or token counts that never get recorded.
In a large organization, production AI traffic can mean many teams calling multiple providers, with agents invoking tools through MCP alongside direct model calls. A Cloud Native Computing Foundation (CNCF) session on AI traffic with the Agent Router project (then named Envoy AI Gateway) covers how that request path differs from standard API traffic.
1. Token-aware limits meter each request by the tokens it uses
Request-count rate limits treat every request the same. A 50-token summarization request and a 100,000-token code generation request each count as one. The second uses far more model capacity and budget. A token-aware gateway meters every request by the tokens it uses.
Token brokering in Agent Router Enterprise, launched in July 2026, enforces spend caps per team or agent inside the request path. When a budget runs out, the gateway can reroute to a cheaper model so work continues. Enforcement lags spend by a few minutes, so spend can pass a cap before the cap trips.
Verification test: Ask the vendor to show a budget cap tripping in a test environment, including what happens to the next request. If the gateway only reports spend after the fact, it has no inline token control.
2. Traffic splitting compares models on live requests
A team may pick a model from public benchmarks or vendor claims, then commit production traffic to it. If the model costs more or answers worse on real prompts, every team using it sees the problem at once.
Traffic splitting sends a set percentage of requests to a second model while the rest stay on the current one. Model aliases let applications keep calling one model name while the platform team changes the model behind it.
Traffic splitting has a real cost: users on the split see the candidate model’s answers. Starting with a small percentage limits that exposure. Cost and latency show up in gateway data, but judging answer quality still takes your own evaluation method.
Verification test: Can you send a fixed percentage of one model’s traffic to a second model, then compare cost and latency for each side? If the only option is switching all traffic at once, every model change carries full production risk.
3. Automatic failover limits a provider outage to one route
A single frontier provider outage takes down every team’s agents at once when there is no failover. Automatic model failover reroutes traffic to a fallback model without manual intervention, so one outage affects a single route.
The gateway detects provider failure through health checks, error rate thresholds, or timeout patterns, then retries failed requests and sends new ones to a configured fallback. The fallback can be a different provider, a different model from the same provider, or a private model running in your own infrastructure.
Scale matters for failover logic, because behavior at 100 requests per second may change at 100,000. Netflix processes billions of API requests daily on Envoy. Lyft was already running Envoy at over 2 million requests per second in 2016. The Agent Router project runs on Envoy, so Agent Router Enterprise gateways use the same proxy. Failover runs inside each gateway, so failover keeps working if the management plane is unreachable.
For self-hosted graphics processing unit (GPU) clusters, Agent Router Enterprise can also route on live capacity, sending traffic from a backed-up cluster to one with headroom.
Verification test: Does the gateway detect provider failure and reroute new requests without manual action? Ask for the failover latency number, meaning the time from failure detection to traffic flowing through the fallback route. A 30-second failover can mean 30 seconds of failed requests for every affected team.
4. Cost attribution maps AI spend to specific teams
Token spend attributed to a single API key does not tell finance which team spent it. When a CFO asks which team spent $40,000 on tokens last month, someone needs to name the team.
Showback reports spend by team so leadership can see consumption patterns. Chargeback bills that spend to each team’s own budget. Budget caps in the gateway are a separate control that limits overspending inside the request path. The cost budgeting guide covers how to align AI infrastructure spending with organizational constraints. A Tetrate video on AI budgeting shows how per-team attribution works in practice.
Verification test: Can the gateway export cost data to your existing billing or FinOps system, or does it keep that data in its own dashboard? Cost data that stays in one vendor dashboard adds another place finance has to check.
5. Group-level policy sets models, guardrails, and budgets once for many users
Platform teams apply approved models, guardrails, and budgets by group, so access can open widely without review of each request. Manual approval for every request does not scale past one team.
Policy only covers traffic that passes through the gateway. Direct API calls that bypass the gateway leave gaps. One management plane applies the same policy to every gateway, in every cloud and region.
Personally identifiable information (PII) redaction and prompt filtering can run inline on every request, with the same inspection on the response. A Tetrate guardrails video (recorded under the earlier Agent Operations Director name, now part of Agent Router Enterprise) demonstrates these controls on live traffic.
Verification test: Does the gateway support a monitor mode that logs without blocking? Policy rollout that blocks developers on day one adds friction, which can push some users toward unapproved tools. A monitor mode shows what would be blocked before enforcement starts.
6. MCP tool governance audits what agents do
Agents call MCP tools as well as models. A gateway that only governs model calls leaves the tool-calling layer unaudited, so security teams cannot see what an agent did.
Through MCP, an agent can reach tools for querying databases, calling APIs, or running code. Without governance at the tool-call layer, an agent can reach every tool its credentials allow. Production MCP governance needs several controls, and the first four follow guidance from the Open Worldwide Application Security Project (OWASP) on excessive agency:
-
Scoped tool access appropriate to the task
-
Credentials tied to an identity
-
Approval gates for high-risk calls
-
Complete logging of every call with its caller
-
Multi-layer filtering
-
Restricted access to governance configuration changes
Agent Router Enterprise’s MCP Gateway provides a curated tool catalog with per-profile authentication, so security can audit which tools an agent reached. A demo of parameter-level authorization shows policy enforced on live MCP tool calls. The LLM vs AI vs MCP gateway comparison explains how these three gateway types differ in scope.
Step-up authorization, built with Ory, escalates high-risk agent actions to human approval through your identity provider, such as Okta or Microsoft Entra ID. Routine agent actions proceed without extra steps.
Verification test: Can the gateway distinguish between routine and high-risk actions? Does the escalation path use your existing identity system, or a separate approval workflow the vendor controls?
7. Streaming support delivers tokens as the model generates them
Streaming inference commonly uses Server-Sent Events, which let a server push each new token to the client as soon as it exists.
A gateway that buffers the full model response before forwarding it adds the entire generation time to perceived latency. That buffering delay breaks live use cases like chat interfaces, code completion, and real-time summarization. With streaming, the user sees output as soon as the first token arrives.
Every millisecond the gateway adds before the first token also delays the user’s sense of responsiveness. The model latency guide covers the fundamentals of measuring and optimizing latency in production AI systems.
Verification test: Does the gateway track token spend on streaming requests accurately, or does it only meter completed responses? A gateway that loses token counts on interrupted streams underreports spend. The gap between reported and actual spend grows with usage.
8. Multimodal support routes image and audio traffic to capable models
Multimodal traffic means text, image, and audio workloads through the same gateway. Routing gets harder when each modality has its own model requirements, cost profile, and latency tolerance.
Image and audio inputs can make requests much larger than typical text prompts. A gateway that treats a 2 MB image upload the same as a 200-token text prompt can strain connection pools and memory allocation.
Model capabilities also differ by modality. A model that handles text well may not accept images at all, so the gateway needs modality-aware routing rules that send each request to a model that can process it.
Verification test: Does the gateway maintain text latency under concurrent multimodal load? Run mixed-modality load tests before committing to a vendor, because the interaction between payload types under load is where architecture differences show.
9. Independent tests under real AI traffic verify gateway overhead
Vendor-asserted performance numbers are claims. Independently measured numbers are evidence. The gateway sits in the request path of every AI interaction in the organization, so the overhead number matters.
Most published gateway benchmarks test a simple case. The gateway forwards requests to a mock server that returns a fixed answer right away, at a steady rate. That setup produces a small overhead number, often in microseconds, but real AI traffic does not look like that setup.
Broadcom’s VMware Cloud Foundation team tested the Agent Router project (then named Envoy AI Gateway) under conditions close to production:
| Test choice | Simple gateway benchmark | Broadcom/VMware test |
|---|---|---|
| Backend | Mock server that returns a fixed answer | Real model inference on four NVIDIA H100 GPUs |
| Traffic shape | Steady requests at a fixed rate | Bursts modeled on agent tool chains, automated jobs, and human users |
| Prompts | A few prompts, repeated | 124 GB of data with over 20,000 unique sessions from public agent and coding benchmarks |
| Duration | Minutes | Three-hour endurance run at 190 concurrent users |
| Reported result | Overhead alone, with no reference point | Overhead compared with full end-to-end latency |
Repeated prompts make a benchmark look better than production. Inference servers cache repeated prompt text. When a test sends the same few prompts, the cache hit rate can approach 99.9%, and the test never measures the real cost of new prompts. The Broadcom/VMware team used unique sessions so the cache could not hide that cost.
The Broadcom/VMware test found:
-
About 2 milliseconds of gateway overhead, about 0.01% of end-to-end latency
-
Overhead stayed flat up to the saturation point at 224 concurrent users. The GPUs reached their limit first, not the gateway.
-
Average time to first token was 0.103 seconds during the three-hour run at 190 users
-
Full responses averaged about 14 seconds
A 2-millisecond result can look slower than a microsecond claim from a mock-server test. Users cannot feel the difference between 2 milliseconds and 0.05 milliseconds when the full response takes 14,000 milliseconds. Users can feel latency that rises under load. The Broadcom/VMware test shows that the Agent Router project keeps overhead flat under real load, up to the GPU limit.
As of September 2026, Tetrate knows of no other open source AI gateway with a published performance test by a large infrastructure vendor such as Broadcom/VMware. The benchmarks that other open source AI gateways publish come from each gateway’s own vendor.
The Agent Router project is the data plane behind Agent Router Enterprise, so Agent Router Enterprise customers run the same data plane that Broadcom/VMware tested. The Tetrate benchmarks page and the Broadcom white paper describe the full setup, so you can design a comparable test for your own environment. Tetrate Head of Product David Wang’s overview of how Tetrate manages AI at enterprise scale covers the architectural decisions behind these numbers.
Verification test: Ask any vendor for an independently measured overhead number from real inference, not a mock server. If the vendor cannot produce one, treat its performance claims as unverified. Then run your own test under concurrent load, with unique prompts, and measure three things:
-
Requests per second at peak capacity
-
Latency overhead per request, as a share of end-to-end latency
-
Whether overhead stays flat as traffic grows
Open source data planes differ in how much production use proves them
Many AI gateways are open source, but an open source license does not prove that a gateway works at scale. Some open source gateways are a few months old, run by one vendor, and tested only by that vendor. Other open source projects have years of production use at large companies, independent governance, and outside testing.
Ask four questions about any open source data plane:
| Question | Envoy and the Agent Router project |
|---|---|
| How long has the data plane run in production? | Envoy has run in production since 2016, starting at Lyft. |
| At what scale? | Billions of API requests per day at Netflix. Over 2 million requests per second at Lyft in 2016. |
| Who governs the project? | Envoy is a graduated CNCF project. The Agent Router project is at AAIF. |
| Has a large outside party measured performance under real AI traffic? | Yes. Broadcom/VMware measured about 2 milliseconds of overhead on the Agent Router project with real GPU inference. |
The Agent Router project adds AI features such as token limits, model routing, and MCP governance on top of Envoy. The AI features are newer, but the proxy that carries every request already has years of production use at large scale.
Verification test: Ask the vendor how long the data plane has run in production, at what scale, under which foundation, and who outside the company has tested the data plane. If the answers are “less than a year,” “our own customers,” “our company,” and “nobody,” treat the data plane as unproven.
Many gateways under one management plane apply one policy everywhere
Large organizations run AI workloads in more than one cloud, region, and data center. One central gateway forces every prompt to travel to one location. Separate gateways managed one by one end up with different policies.
Agent Router Enterprise runs a gateway wherever the workloads run. One management plane sets approved models, budgets, guardrails, and MCP tool access for all of the gateways. Routing, failover, cost attribution, and policy enforcement run inside each gateway, so every policy decision sees the same request.
The data plane carries traffic, and the management plane handles configuration. If a gateway loses its connection to the management plane, the gateway keeps routing requests on the configuration it already has.
Verification test: Ask the vendor to change one policy in the management plane and show the change reach gateways in two different clouds or regions. Then cut the management plane connection and confirm that both gateways keep routing.
Deployment models decide where prompts travel
| Deployment model | Data plane location | Management plane | Prompt traffic |
|---|---|---|---|
| Fully managed | Tetrate operated | Tetrate hosted | Runs on Tetrate infrastructure |
| Self-hosted data plane | Your cloud, on-premises, or edge, one per region if needed | Tetrate hosted | Stays in your network. Logged content is configurable. |
| Distributed gateways | Many gateways across clouds, regions, on-premises sites, and edge locations | One Tetrate-hosted management plane for all gateways | Stays in each local network. Logged content is configurable. |
Some regulated organizations have data residency policies that favor keeping the data plane inside their own environment. Cloudflare AI Gateway runs on Cloudflare’s network, which works well for organizations already on that network. Agent Router Enterprise can run the data plane in your cloud, on-premises, at the edge, or per region, so prompts never have to leave your network. A Tetrate vs Cloudflare comparison covers this distinction in detail.
Ask for the mechanism behind every gateway claim
A mechanism can be tested, while a bare claim requires trust. “We support failover” is a claim. A vendor that says “we detect provider failure through error rate thresholds and reroute new requests to a configured fallback within 200 milliseconds” is describing a mechanism you can test.
Warning signs of weak AI gateways:
-
Proxies built on standard Python runtimes that hit central processing unit (CPU) limits under high concurrency
-
General API gateways that add AI support as a layer on top of a REST-era design
-
Single-cloud lock-in that prevents cross-cloud or on-premises deployment
-
Pricing that takes a percentage of inference spend
-
Performance claims with no independent verification
-
Gateways that must be configured one at a time, with no shared management plane
-
A proprietary data plane with no open source project or foundation behind the data plane
-
An open source data plane with little production use at scale, governed and tested only by the vendor that sells the data plane
OpenRouter charges a platform fee of 5.5% to 8% on credit purchases, depending on plan, so its revenue grows with your spend. Tetrate licenses Agent Router Enterprise as a subscription with no per-token fee, so procurement can predict cost up front.
A home-built gateway can work for a first version. The maintenance burden grows as more teams, models, and providers come online. The gateway also only sees the traffic routed through it.
In one anonymized deployment, a large fintech company with thousands of developers retired its home-grown gateways. Model access widened beyond two frontier providers to Google Vertex AI, Mistral, and private open-weight models. Every AI tool now enters through one governed path with full consumption telemetry, and none of it required application rewrites.
Now Available
Conclusion: test production behavior before you commit
Each of the nine capabilities in this guide changes how a gateway behaves under production load, and each comes with a test. The strongest evidence is a result you can reproduce in your own environment. Run the gateway under your own traffic mix before committing, since a gateway that performs well in a vendor demo can still degrade under real load.
Start a fast-track evaluation: first routed request under five minutes, nothing to install. Enterprise teams with a defined use case can also request pricing or ask about a four-week proof of concept.
FAQs
What can an AI gateway do that a standard API gateway cannot?
An AI gateway meters requests by tokens, fails over between model providers, attributes cost by team, and governs MCP tool calls. General API gateways were built for REST traffic. Several now add some of these features as extra policies.
How do I know if my organization needs an AI gateway?
You likely need an AI gateway when more than one team, model, or provider is in production, and cost, access, or resilience policy needs one enforcement point. If your organization runs AI through a single team with a single provider, a gateway adds complexity without proportional benefit.
What’s the difference between AI gateway features and marketing claims?
A feature changes production behavior under real load and passes an independent test. A marketing claim states an outcome without the mechanism that produces it. Ask to see a budget cap trip in a test environment, a failover latency number, and an independently measured overhead figure.
Can an AI gateway work across multiple clouds and on-premises?
Yes, if the vendor can run gateways in your cloud, on-premises, at the edge, or per region, with one management plane for all of the gateways. Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane. Cloudflare AI Gateway is tied to a single network. Amazon Bedrock AgentCore Gateway and Azure API Management fit best for teams that run most of their AI stack in AWS or Azure.
How much overhead does an AI gateway add to inference requests?
Overhead varies by gateway and by test setup, so ask for a number measured by an outside party under real inference. Broadcom’s VMware Cloud Foundation team tested the Agent Router project (then named Envoy AI Gateway), the data plane behind Tetrate Agent Router Enterprise, with real GPU inference, bursty agent traffic, and over 20,000 unique sessions. The test measured about 2 milliseconds of gateway overhead, about 0.01% of end-to-end latency, and the overhead stayed flat up to the GPU limit.
What data plane is Tetrate Agent Router Enterprise built on?
Tetrate Agent Router Enterprise is built on the Agent Router project (formerly Envoy AI Gateway), an open source data plane at the Agentic AI Foundation (AAIF) that runs on the Envoy proxy. Envoy is a CNCF project that Netflix and Lyft run in production at large scale. An independent Broadcom/VMware test measured about 2 milliseconds of overhead on the Agent Router project data plane under real AI traffic.
Are all open source AI gateways equally proven?
No. An open source license does not show that a gateway works at scale. Check how long the data plane has run in production, at what scale, who governs the project, and whether a large outside party has measured performance. Tetrate Agent Router Enterprise is built on the Agent Router project, which runs on Envoy. Envoy has run in production since 2016 at companies such as Lyft and Netflix and is a graduated CNCF project. Broadcom/VMware tested the Agent Router project under real AI traffic.
What happens to AI traffic if the gateway management plane goes down?
The answer depends on whether the vendor separates the data plane from the management plane. In Agent Router Enterprise, each gateway keeps routing requests on its current configuration if the management plane connection is interrupted. Ask every vendor to show this behavior in a test environment.
Key terms glossary
Token-aware limits: Budgets and rate limits measured in tokens per request. They let the gateway cap spend inside the request path.
Traffic splitting: Sending a set percentage of requests to a second model while the rest stay on the current one. Teams use it to compare models on live traffic.
Model failover: Automatic rerouting to a fallback model when a provider fails, without manual intervention. It limits the impact of a provider outage to a single route.
Showback and chargeback: Showback reports AI spend by team for visibility. Chargeback bills that spend to each team’s own budget.
MCP (Model Context Protocol): An open protocol that connects AI applications to external tools and data sources.
Step-up authorization: Escalating high-risk agent actions to human approval through the organization’s own identity system. Routine actions proceed without extra steps.
Data plane and management plane: The data plane is the set of gateways that carry live traffic. The management plane handles administration and configuration for all of the gateways. In Agent Router Enterprise, each gateway keeps routing on its current configuration if the management plane connection is interrupted.
Distributed data plane: Many gateways deployed across clouds, regions, on-premises sites, and edge locations, all managed from one management plane.
Agent Router project: The open source AI gateway data plane, formerly named Envoy AI Gateway, now at the Agentic AI Foundation (AAIF). The Agent Router project runs on Envoy. Tetrate Agent Router Enterprise is built on the Agent Router project.
Agentic AI Foundation (AAIF): The open source foundation that hosts the Agent Router project.
Envoy: An open source proxy for service traffic, hosted by the CNCF. The Agent Router project runs on Envoy.
MCP Catalog with verified first-party servers, profile-based configuration, and OpenInference observability are now generally available in Tetrate Agent Router Service . Start building production AI agents today.