Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

The Open Source AI Gateway Buyer's Guide for 2026: Cost, Support, and Governance Trade-Offs

Featured Image

TL;DR: Open source AI gateways differ most by runtime: LiteLLM is a Python proxy, Bifrost is written in Go, Helicone’s gateway is written in Rust, and the Agent Router project (formerly Envoy AI Gateway) at the Agentic AI Foundation (AAIF) runs on the Envoy proxy. Python’s concurrency model limits how much traffic one LiteLLM process can carry. Helicone has been in maintenance mode since March 2026. Every self-hosted option puts upgrades, failures, and provider API changes on your platform team. That work drives total cost of ownership (TCO). Python proxies suit one team that wants to start fast, while Envoy-based data planes suit cross-team traffic that needs one enforcement point. Tetrate Agent Router Enterprise adds fleet management, cost controls, and commercial support on top of the Agent Router project.

An open source AI gateway sits in the request path for every model call your teams make, so its architecture decides how much load it can carry and what it can govern. Open source gateways are already widely adopted: LiteLLM alone has more than 90 million monthly downloads from the Python Package Index as of this writing.

Feature lists make the options look alike. A platform leader needs to know which architecture keeps carrying traffic during a provider outage, when an agent calls a tool it should not touch, or when the management plane goes down. We compare four options on architecture, total cost of ownership, support, and governance span.

An AI gateway enforces policy on live traffic

An AI gateway is a proxy that enforces policy on every request it carries. A plain API proxy translates request formats and forwards traffic. An AI gateway also handles routing with automatic failover across providers, per-team cost attribution, and spend limits enforced on live traffic. That extra layer matters because a proxy that only forwards requests cannot govern what agents do once they have model access.

We’ve found the decision to adopt an AI gateway usually comes down to five requirements:

  • More than one team runs AI workloads in production
  • A multi-model or multi-provider strategy needs one enforcement point
  • Governance must span clouds and regions for data residency
  • Mission-critical traffic needs support with defined response times
  • Procurement needs costs that stay predictable as usage grows

Data plane placement decides where prompts and data go

A gateway sold only as software as a service (SaaS) runs in the vendor’s cloud, so every request passes through infrastructure outside your network. A distributed gateway separates the system into two parts. The data plane carries traffic in your cloud, on-premises, or at the edge. The management plane stores configuration and policy in one place.

For regulated industries where prompts and data must stay inside your network, data plane placement decides whether a gateway can meet that requirement at all. Placement also affects resilience. Ask each vendor what happens to live traffic when the management plane is unreachable. Agent Router Enterprise gateways keep routing traffic in that case. New configuration cannot reach the gateways until the connection returns.

Four open source AI gateways use four different architectures

LiteLLM runs on Python, Bifrost on Go, and Helicone on Rust. The Agent Router project runs on the Envoy proxy, with an external processor written in Go that handles AI requests. The runtime shapes how each gateway behaves under load and how much infrastructure your team has to run.

The table compares the open source edition of each gateway and its paid support path.

GatewayArchitectureGovernance spanSupport model
Agent Router projectEnvoy-based data planeModel calls and MCP tool callsCommunity, or Agent Router Enterprise support tiers (standard, premium)
LiteLLMPython proxyModel calls, agents, and MCPCommunity, or LiteLLM Enterprise
BifrostGo-based proxyModel calls and MCPCommunity, or Bifrost Enterprise
HeliconeRust-based gateway with observability platformObservability, routing, and failoverMaintenance mode since March 2026

The Agent Router project runs on the Envoy proxy

The Agent Router project, formerly Envoy AI Gateway, reached version 1.0 in June 2026. Tetrate co-created the project with Bloomberg. On September 10, 2026, the project maintainers donated the codebase to the Agentic AI Foundation (AAIF), part of the Linux Foundation.

The project runs on the Envoy proxy, which has run in production since 2016 and is a graduated project at the Cloud Native Computing Foundation (CNCF). One policy layer covers model calls and tool calls made through the Model Context Protocol (MCP). The project is free to self-host on Kubernetes. Your team owns upgrades, patches, and incident response for a self-hosted deployment. Agent Router Enterprise adds fleet management, hybrid deployment, dashboards, cost controls, MCP access control, and commercial support on top of the project. The independent Broadcom/VMware test of the Agent Router project (then named Envoy AI Gateway) measured about 2 milliseconds of gateway overhead per request. Tetrate’s performance benchmarks page covers the test setup and other results.

LiteLLM is a Python proxy that is quick to start

LiteLLM translates calls to OpenAI, Anthropic, and other providers into one OpenAI-compatible format with little setup. The open source version includes virtual keys, budgets, and fallbacks. LiteLLM Enterprise adds paid support and enterprise controls. LiteLLM also offers an MCP gateway and an agent gateway, so governance can extend past model calls.

LiteLLM runs as a Python service. The CPython Global Interpreter Lock lets only one thread execute Python bytecode at a time, so Python services usually scale by adding worker processes and instances. Each new instance adds infrastructure for your team to run. LiteLLM also needs a PostgreSQL database for keys and budgets. LiteLLM is moving parts of its gateway to Rust, and that work is in beta as of this writing. Test LiteLLM at your own peak load before you move cross-team traffic onto it.

Bifrost is a Go proxy built for low overhead

Bifrost, built by Maxim AI, is an open source gateway written in Go and released under the Apache 2.0 license. Bifrost exposes one OpenAI-compatible API across more than 20 providers. The gateway handles automatic failover, load balancing, and semantic caching. A built-in MCP gateway collects tools from connected MCP servers behind one endpoint.

Go compiles to native code and runs many requests concurrently in one process. Bifrost’s design targets low per-request overhead as a result. The open source edition is self-hosted, so your team runs upgrades and handles failures. Paid support comes through Bifrost Enterprise. If you plan to run Bifrost in several regions, ask how policy stays consistent across instances during upgrades and failures.

Helicone moved to maintenance mode in 2026

Helicone started as an open source observability platform for model calls and added an open source AI gateway in 2025. The gateway is written in Rust. Helicone’s gateway routes requests across providers with automatic failover, caching, and rate limits. Traces flow into Helicone’s observability platform.

After Mintlify acquired Helicone in March 2026, the product moved to maintenance mode, where security updates, bug fixes, and new models continue to ship. Mintlify said it will work with customers on migrating to another platform. The acquisition changes the support question for a buyer. A team adopting Helicone today should plan around maintenance mode and budget for a later migration.

Self-managed gateway costs grow after version one

LiteLLM, Bifrost, Helicone, and the Agent Router project are all free to download. The cost shows up in the engineering time spent running the gateway. That cost applies just as much to a gateway your team builds itself.

Maintenance work lands on the platform team

A basic home-grown proxy can be built quickly, and version one usually works. The ongoing cost comes from four areas:

  • Provider API changes that break adapter code
  • Reliability engineering to keep uptime consistent across regions
  • Security patches for Transport Layer Security (TLS) and authentication libraries
  • Token counting accuracy and rate-limit consistency across replicas

Each new team that onboards also needs the routing logic and quota rules explained. The same list applies to a self-hosted open source gateway, except that the project maintainers keep the provider adapters current. When gateway maintenance takes time away from new platform work, compare the cost of that engineering time with the cost of a supported gateway.

A fintech company retired its home-grown gateways

A large financial technology company with thousands of developers and staff using AI daily replaced its home-grown gateways with Agent Router Enterprise.

Before the switch:

  • Teams reached AI through disconnected tools such as Copilot, Codex, and Claude Code
  • Home-grown gateways cost effort to maintain
  • Home-grown gateways only saw the traffic they routed
  • Frontier model access ran only through Azure AI Foundry and Amazon Bedrock
  • Usage on direct API calls went untracked

After the switch:

  • Every tool entered through one governed path with full consumption telemetry
  • Model access widened to Google Vertex AI, Mistral, and private open-weight models
  • AI ROI became trackable for the first time

Retiring the home-grown gateways freed the engineering time the team had spent maintaining them, with no application rewrites. A platform team running its own gateway can start the same comparison by totaling the hours it spends on gateway maintenance each month.

Support, debugging, and licensing add to running cost

License fees are one part of support cost. Support cost also covers on-call time for gateway failures and the hours spent reconstructing multi-agent chain failures. Without centralized gateway logs, debugging means correlating timestamps across separate provider dashboards for OpenAI, Anthropic, and internal models. A diagnosis that should take minutes can stretch into days of triage.

Agent Router Enterprise gateway logs let a team reconstruct multi-agent chain failures and export the traces to the observability tools it already runs. Vendor support with a defined response time shortens the wait for help during mission-critical incidents. Count both costs when you compare a self-managed gateway with a supported one.

Self-hosted open source gateways charge no per-token fee, so you pay the model providers directly. Some managed gateways charge a percentage on top of spend. OpenRouter’s pricing page lists a 5.5% platform fee on its Standard plan and 8% on Business, so the fee grows with usage.

For a self-hosted gateway, infrastructure and engineering time are the costs that grow with usage. Agent Router Enterprise is licensed with flat pricing by feature set, deployment footprint, and support tier, so the license fee stays the same as token consumption grows.

Support models decide recovery time when the gateway fails

Self-managed support means your platform team owns upgrades, failures, and debugging. That model fits single-team deployments with low traffic and no requirement for a formal support service level agreement (SLA). Cross-team infrastructure needs more, because the gateway sits in the request path for every AI call. When the gateway itself fails, every team’s agents stop working until routing recovers.

Response time targets matter because the gateway is a blocking dependency for every team. Agent Router Enterprise support comes in standard and premium tiers. Premium includes 24x7 support with a one-hour response on severity-one issues, so an overnight page gets vendor engagement while your platform team assesses the scope of the impact. Tetrate prices the support tier as part of the enterprise license.

Governance has to reach agent tool calls

A gateway that governs only model calls can show which model was called, but it cannot show what the agent did with that access. Governance for agents has to cover tool calls made through MCP, data residency, and audit logging across the full request path.

One policy layer applies approved models by group

Group-level policy applies approved models, guardrails, and budgets to every team from one place. The same rules follow traffic across regions and providers. With group policy in place, you can open AI access widely without reviewing each request, while cost and model use stay visible. A budget set to hard stop refuses a team’s requests once its limit is spent. Budget enforcement runs a few minutes behind spend, so an immediate cutoff comes from a rate limit enforced inline at the gateway.

Without group policy, access has to be gated team by team. That gating pushes analysts, marketers, and ops staff toward chat tools outside any governed path, which is one way shadow AI starts. Agent Router Enterprise applies per-team model access controls through model catalogs, roles, and budgets.

The MCP Gateway and step-up authorization govern tool calls

An agent with model access can still reach tools and data that nobody approved. Agent Router Enterprise governs tool access in two ways. The MCP Gateway provides a curated tool catalog with per-profile authentication and filtering, so security can audit every tool call an agent makes.

Step-up authorization, built with Ory, lets routine agent actions proceed without friction. Higher-risk actions, defined in policy, pause until a human approves them through an Ory approval flow. Ory then issues a short-lived token scoped to that one action. When you evaluate any gateway, check whether each agent profile sees only an approved tool catalog. Also check how a higher-risk action gets human approval.

Self-hosted data planes keep prompts inside your network

Distributed deployment can place the data plane in your AWS, Azure, or Google Cloud virtual private cloud (VPC), in an on-premises data center, or at an edge location. With a self-hosted data plane, prompts and inference data stay inside your network. That placement supports sovereign AI requirements where data must stay within a specific jurisdiction. Placing gateways by region or zip code also cuts latency for distributed teams.

Immutable audit logs support compliance review. You choose how much request data the gateway captures: the full request, only counts and cost, or nothing at all. You can change that setting at any time. Tetrate’s compliance covers SOC 2 Type II, ISO 27001, and GDPR. The AI gateway security guide covers encryption, data capture, and deployment models in detail.

The right gateway depends on governance span and deployment constraints

Three factors decide the choice: governance span, deployment constraints, and how much operational work your team can absorb.

Single-cloud teams may only need a cloud provider gateway

If your AI architecture commits to one cloud and has no data residency needs outside that cloud’s regions, a cloud provider gateway such as AWS Bedrock or Azure API Management may be enough. If governance must span several clouds, direct model providers, and your own infrastructure, you need an independent control point. The same applies if you route inference by region for latency or data sovereignty. An independent gateway keeps policy portable across environments, so you keep the option to move workloads on cost, capacity, or sovereignty. The matrix below compares the two situations.

RequirementSingle-cloudMulti-cloud or sovereign
Governance scopeOne cloud environmentSeveral clouds, on-premises, edge
Data residencyCloud provider regionsCustomer network
Model accessCloud catalog plus direct providersAny provider plus self-hosted models
Support modelCloud provider supportGateway vendor support tiers

Agent Router Enterprise runs fully managed or hybrid

Tetrate offers Agent Router Enterprise fully managed or in a hybrid model. In hybrid mode, Tetrate runs the management plane while you run the gateways where traffic flows, in your cloud or on-premises. The gateways open an outbound connection to the management plane for configuration, so that path needs no inbound firewall openings. Fully air-gapped deployment is not available today.

For application code, moving from a bespoke proxy or from LiteLLM is a base URL change. Agent Router Enterprise is 100% OpenAI-compatible, so agents need no rewrites. Keys, budgets, and routing policy still have to be set up in the new gateway.

The architecture decision comes before the gateway carries every team’s traffic

Python proxies suit a single team that wants to start quickly, while Go and Rust proxies target low per-request overhead. Data planes built on the Envoy proxy suit cross-team production traffic that needs one enforcement point across regions and providers. Weigh the cost of running each gateway yourself alongside its license price. For platform leaders at regulated or multi-cloud organizations, one enforcement point is the outcome to plan for, and the architecture has to support it from the start.

Agent Router Enterprise is built on the Agent Router project, the open source data plane at AAIF that runs on the Envoy proxy. Agent Router Enterprise runs gateways in any cloud, region, on-premises site, or edge location, and manages all of the gateways from one management plane. Start a fast-track evaluation: the first request routes in under five minutes, with nothing to install, or request pricing for an enterprise license quote.

Agent Router Enterprise

Tetrate Agent Router Enterprise routes AI agent traffic across providers and your own models, with policy, cost controls, and audit on every request — in cloud, on-prem, or edge.

Learn more

FAQs

What is the difference between an open source AI gateway and a cloud provider’s native gateway?

An open source AI gateway can run as an independent control point across clouds, direct providers, and your own infrastructure. A cloud provider’s native gateway, such as AWS Bedrock or Azure API Management, is strongest when one cloud is the center of your AI architecture. Agent Router Enterprise is one example of an independent control point, with data planes that can run in your own cloud or on-premises.

How do I migrate from LiteLLM to Agent Router Enterprise?

Recreate keys, budgets, and routing policy in the new gateway first, then point application code at the new endpoint. Agent Router Enterprise is 100% OpenAI-compatible, so the code change is one base URL with no agent rewrites.

Do open source AI gateways take a percentage of inference spend?

No. Self-hosted open source gateways such as LiteLLM, Bifrost, and the Agent Router project charge no per-token fee. You pay the model providers directly. Some managed gateways charge a percentage, and OpenRouter’s pricing page lists a 5.5% platform fee on its Standard plan and 8% on Business. Agent Router Enterprise also charges no per-token fee.

What happens if the management plane goes down?

In Agent Router Enterprise, the gateways carrying traffic keep working when the management plane is unreachable. Ask every vendor you evaluate the same question before you put the gateway in the request path.

Key terms glossary

Agent Router project: The open source data plane formerly called Envoy AI Gateway, now hosted at AAIF. Teams can self-host the project for free, or buy Agent Router Enterprise for commercial support and fleet management.

Data plane: The component that carries live traffic and enforces policy on every request. In Agent Router Enterprise, data planes can run in your cloud, on-premises, at the edge, or per region. Where the data plane runs decides whether prompts stay inside your network.

Envoy proxy: The open source proxy the Agent Router project runs on, and a graduated CNCF project. For gateway performance, the Broadcom/VMware test of the Agent Router project is the direct measurement.

Enforcement point: A single gateway that applies policy to all AI traffic from all teams, regardless of which provider or region handles the request. For multi-team or multi-cloud deployments, one enforcement point keeps model access, cost controls, and audit logs consistent across every path.

Management plane: The central component that stores configuration and policy for all data planes. In Agent Router Enterprise, Tetrate hosts the management plane.

Fleet management: The ability to configure, monitor, and update multiple gateways from one control surface. Agent Router Enterprise provides fleet management through its management plane, so a platform team does not need to configure each gateway individually.

Model Context Protocol (MCP): An open protocol that lets agents call external tools and data sources. Governing MCP calls is what extends a gateway’s policy past model calls.

Maintenance mode: A product state where security updates, bug fixes, and new model support continue to ship, but no new features are added. Helicone moved to maintenance mode after the Mintlify acquisition in March 2026. A team adopting a gateway in maintenance mode should plan for a migration.

Sovereign AI: The requirement that AI data and inference stay within a specific jurisdiction or network. Self-hosted data planes in your cloud, on-premises, or per region support sovereign AI requirements.

Step-up authorization: A policy mechanism that lets routine agent actions proceed without friction while higher-risk actions pause for human approval. Agent Router Enterprise builds step-up authorization with Ory.

Total cost of ownership (TCO): The full cost of running a gateway over time, including license fees, infrastructure, upgrades, on-call time, and debugging.

Related: Open source vs commercial AI gateway · How to choose an open source AI gateway · Tetrate vs Bifrost · Tetrate vs Envoy AI Gateway (OSS) · Is LiteLLM safe?

See the full 2026 enterprise AI gateway comparison.


Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.

Decorative CTA background pattern background background
Tetrate logo in the CTA section Tetrate logo in the CTA section for mobile

Ready to enhance your
network

with more
intelligence?