Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

Open Source AI Gateway vs Commercial AI Gateway: Key Differences for Platform Teams

Featured Image

TL;DR: An open source AI gateway costs nothing to license, but your platform team owns every upgrade, security patch, and scaling event. Self-hosting fits teams that already run Envoy or Istio in production and have spare capacity to operate it. A commercial gateway moves much of that operational work to a vendor with a support contract. Tetrate Agent Router Enterprise pairs a Tetrate-hosted management plane with gateways you can run in your own cloud or on-premises, so prompt and response text stays inside your network unless request logging keeps it, and an admin can turn that logging off.

Platform teams choosing an AI gateway trade license cost against operating effort. Gartner forecasts AI gateway spending will grow 70.9% in 2027, from $251 million in 2026 to $429 million. This comparison covers deployment, cost, support, customization, governance, and performance for teams running AI across more than one model provider.

AI gateways split into model proxies and agent governance layers

AI gateways typically serve one of two functions:

  • The first is the large language model (LLM) API proxy, which standardizes traffic between applications and model providers. It handles routing, provider authentication, rate limits, caching, and model-level observability.
  • The second is the agent and MCP (Model Context Protocol) governance layer. It governs a different path: access from AI clients or agents to MCP servers, tools, and enterprise data.

A proxy cannot govern traffic it never sees. If your requirement covers which tools an agent can call, a model proxy alone will not meet it, so match the gateway type to the traffic in scope before comparing vendors.

Open source projects and commercial products exist for both functions. The open source versus commercial question is about who operates the gateway and who answers when it fails, whatever traffic it carries.

Three deployment models set data residency and operational overhead

Where the gateway runs decides where prompts flow and how much work your platform team carries.

ModelWhere prompts flowOperational overhead
Fully managedVendor’s infrastructureLow
Self-hostedYour networkHigh
HybridYour network, with a vendor-hosted management planeMedium

A fully managed AI gateway runs in the vendor’s cloud, which means prompts and responses pass through infrastructure you do not control. Self-hosted gateways keep data inside your network but leave every operational task with your platform team.

A hybrid model pairs a vendor-hosted management plane with data planes you run in your own cloud or on-premises. Tetrate documents how hybrid deployment keeps data private for regulated workloads.

When to self-host your AI gateway

Self-hosting fits when both of these are true:

  • Your platform team already runs Envoy or Istio in production and can operate a distributed proxy at scale.
  • Your team has dedicated capacity to run the gateway without pulling engineers off other roadmap work.

Self-hosting is also the only option when your data rules forbid any vendor-hosted component, including a management plane.

Nobody outside your platform team is accountable for a self-hosted gateway, so the team needs on-call coverage for the gateway the same way it covers any other production service. That cost is easy to miss in a license comparison, because it appears as engineering time.

Self-hosted gateways move cost from the license to engineering time

Open source licenses cost nothing, but operating the gateway carries real expense. Ongoing engineering time is usually the most underestimated line in a self-hosted gateway budget. Someone has to track provider API changes, apply security patches, scale gateways for traffic peaks, and keep metrics flowing to your observability stack.

Cost categoryWhat drives it
InfrastructureNumber of gateways, regions, and environments
Engineering laborPatch cadence and team size
Provider API changesNumber of providers and how often their APIs change
Monitoring and observabilityYour existing metrics and tracing stack

A fair comparison counts three years of licensing, engineering time, uptime work, and integration overhead. A one-year view tends to hide the maintenance work, because upgrades and provider changes accumulate over time.

Flat licensing fixes gateway cost before deployment

Commercial AI gateways price in two broad ways. Some, usually hosted model marketplaces, add a percentage fee on top of model spend. Others license by feature set, deployment footprint, and support tier, with no per-token fee.

With a percentage fee, gateway cost rises with every token your teams send. With flat licensing, gateway cost follows the size of your deployment. Flat licensing lets procurement fix the cost before deployment, while model spend is settled separately with each provider.

Agent Router Enterprise uses flat licensing. For a team with a FinOps (cloud financial management) mandate, the pricing model decides whether gateway cost moves with AI usage or stays tied to the infrastructure you approved.

Support contracts decide who answers during a gateway outage

Who gets paged when the gateway itself fails? Open source licenses such as Apache 2.0 provide the software “as is,” without warranties, so response-time commitments come from a separate commercial support contract. Without one, your platform team reads its own logs and debugs its own infrastructure.

Agent Router Enterprise offers standard and premium support:

  • Standard: business hours for severity 1 and 2 issues, with a first response within one business day.
  • Premium: 24x7 coverage for severity 1 and 2 issues, with a first response within one hour on severity 1.

Agent Router Enterprise support comes from the company that co-created and maintains the open source data plane. Gateway logs and traces export to your existing observability tools, so your team can reconstruct a multi-agent failure from one set of logs.

Customization moves from code forks to configuration

Open source gateways give you full code-level customization. You can fork the code, modify routing logic, add custom authentication, and build integrations specific to your stack. Every customization then becomes a fork you must maintain, and each upstream release means merging your changes again.

Policy replaces code changes for most customization

Agent Router Enterprise handles most customization through configuration. Routing rules, guardrails, and access controls are set as policy. The gateway enforces that policy inline in the data path, so no application code changes when a rule changes.

Routing rules include automatic failover across providers, which moves a request to the next model when the first one times out, rate-limits, or returns a server error. A single provider outage then affects only the routes that use that provider.

On model traffic, policy covers cost tracking, personally identifiable information (PII) filtering, guardrails, token rate limits, and access control by agent or model. On tool traffic, an MCP gateway sits between agents and existing MCP servers. It can present a curated subset of tools as a single MCP server.

Monitor mode logs policy hits without blocking, so you can test a PII or prompt rule against live traffic before enforcing it.

Configuration still takes work. Each policy has to be written, tested in monitor mode, and promoted to enforcement. A fork is still the better fit when a team needs behavior that no configuration option exposes.

One management plane governs gateways across clouds and sites

Open source gateways enforce policy where you deploy and configure them. Running them in three clouds and two on-premises data centers means five policy configurations to maintain. It also means five upgrade cycles, and each one is a place where policy can drift.

Agent Router Enterprise applies one policy layer to every gateway from a single management plane, with these controls:

  • An approved model catalog defines which models each team can call.
  • Rate limits are enforced inline in the request path.
  • Budgets track spend per team, teammate, or API key.
  • A hard stop budget refuses requests a few minutes after the limit is spent.
  • Cost and token attribution shows which team or agent spent what.
  • An MCP gateway provides a curated tool catalog with per-profile authentication.

These controls let you open AI access to every employee without losing track of cost, data, or model use.

Gateway overhead stays near 2 milliseconds under real GPU load

Agent Router Enterprise is built on the Agent Router project (formerly Envoy AI Gateway), an open source data plane at the Agentic AI Foundation (AAIF) that runs on the Envoy proxy. The Broadcom/VMware test of the Agent Router project (then named Envoy AI Gateway) measured about 2 milliseconds of gateway overhead per request under real GPU inference, about 0.01% of end-to-end latency.

The test ran real model inference on four NVIDIA H100 GPUs, with over 20,000 unique sessions from public agent and coding benchmarks. Gateway overhead stayed flat up to 224 concurrent users, the saturation point where the GPUs reached their limit first. A three-hour endurance run at 190 concurrent users followed, with an average time to first token of 0.103 seconds.

Envoy counts Netflix, Lyft, and Airbnb among its users. Engineers at Lyft built Envoy, and Lyft released it in 2016. Envoy graduated in 2018 as a Cloud Native Computing Foundation (CNCF) project.

LiteLLM’s published throughput depends on scaling out pods

Tetrate’s benchmark summary cites documented cases where LiteLLM bottlenecks near 300 requests per second, and containers growing to 12 GB of memory and crashing.

LiteLLM’s own benchmark documentation reports 3,000 requests per second with 50,000 to 100,000-token prompts, using 33 gateway pods and a deployment profile still in development, against a mock model. That works out to about 90 requests per second per pod, so the figure reflects scaling out across pods.

LiteLLM has also released an early beta of a Rust gateway. Check which runtime a published benchmark measured before you compare its numbers with any other gateway’s.

Regulated teams can keep data in-network without self-hosting

Self-hosted and hybrid gateways both keep prompts inside your network, so team capacity decides between them. A team with Envoy experience and spare capacity gets the most flexibility from self-hosting. A team without that capacity carries operational work it cannot staff.

Hybrid deployment keeps prompts inside your network

Hybrid deployment fits teams that need data inside their network without the capacity to operate the full gateway stack. In Agent Router Enterprise, Tetrate runs the management plane for the admin API and UI. Your team runs the gateways where traffic flows. Prompt and response text stays inside your network unless request logging keeps it. An admin can turn that logging off with one setting. The gateways keep carrying traffic even when the management plane is unreachable.

Placing gateways by region or zip code cuts latency and keeps traffic in the region your data rules require. A gateway tied to one provider’s network, such as Cloudflare AI Gateway, runs only where that network runs. Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane.

Tetrate has completed a SOC 2 Type II audit and is certified to ISO 27001. Tetrate is also GDPR compliant. Data is encrypted with TLS 1.3 in transit and AES-256 at rest. Access uses multi-factor login and role-based permissions.

Team capacity decides between self-hosted and commercial gateways

You can self-host an open source gateway when your platform team has dedicated capacity and existing Envoy expertise. A commercial gateway with hybrid deployment fits when you lack that capacity but still need data to stay inside your network. Either way, compare three years of engineering time and infrastructure against the license fee before you decide.

Start a fast-track evaluation of Agent Router Enterprise: a dedicated managed instance, with a first routed request in under five minutes, or request pricing for a detailed quote under NDA.

Agent Router Enterprise

Tetrate Agent Router Enterprise routes AI agent traffic across providers and your own models, with policy, cost controls, and audit on every request — in cloud, on-prem, or edge.

Learn more

FAQs

Can I run a commercial AI gateway in my own infrastructure?

Yes, in a hybrid model. With Agent Router Enterprise, Tetrate runs the management plane while you run the gateways in your own cloud or on-premises.

What happens if my open source AI gateway needs urgent support?

You rely on community forums and your own team, unless you buy commercial support for that project. Tetrate’s premium support for Agent Router Enterprise adds 24x7 coverage with a one-hour first response on severity 1 issues.

Do commercial gateways lock me into specific cloud providers?

It depends on the gateway. Cloud provider gateways such as AWS Bedrock or Azure API Management are a strong fit when one cloud is already the center of your AI architecture. Tetrate built Agent Router Enterprise as an independent control point across clouds, direct providers, and your own infrastructure, so you can move workloads based on cost, capacity, performance, or sovereignty.

How do I migrate from an open source to a commercial AI gateway?

Change one base URL in your application code. Agent Router Enterprise is OpenAI-compatible, so existing agents connect without rewrites, and a first request typically routes in under five minutes. API keys, budgets, and routing rules still need to be set up in the gateway before production traffic moves over.

Can an open source AI gateway cost more than a commercial one?

It can, once engineering time is counted. The license is free, but a self-hosted gateway still needs engineers to patch, scale, and maintain it. It also needs infrastructure and observability tooling to run on. A commercial license moves maintenance of the gateway software to the vendor for a fee procurement can plan for.

Key terms glossary

AI gateway: Infrastructure that routes, governs, and pays for AI requests across every model, provider, team, or location. Choosing between open source and commercial options decides who operates it.

Data plane: The component that carries live traffic. In Agent Router Enterprise, the data plane runs as gateways that Tetrate manages or that you host in your own cloud or on-premises.

Hybrid deployment: A model where the vendor hosts the management plane and your team runs the gateways that carry traffic. It suits teams whose data rules require prompts to stay in their own network.

Management plane: The central admin API and UI. Teams use it to configure policy, budgets, and access across every gateway.

MCP (Model Context Protocol): A protocol for connecting agents to tools. An MCP gateway controls which tools each agent can call.

Monitor mode: A policy setting that logs what a rule would have blocked or redacted without changing traffic, so a team can test a rule before enforcing it.

Related: Open source AI gateway buyer’s guide · How to choose an open source AI gateway · AI gateway deployment models · Tetrate vs Envoy AI Gateway (OSS)

See the full 2026 enterprise AI gateway comparison.


Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.

Decorative CTA background pattern background background
Tetrate logo in the CTA section Tetrate logo in the CTA section for mobile

Ready to enhance your
network

with more
intelligence?