Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

LiteLLM vs Kong AI Gateway: Which Open Source Gateway Wins?

TL;DR: LiteLLM suits teams that want broad provider coverage and a fast start. Its default Python proxy scales by adding instances, and LiteLLM is moving its hot path to Rust in a staged rollout. Kong AI Gateway suits teams that already run Kong for API traffic, with semantic caching and token rate limits in its paid Enterprise tier. Tetrate Agent Router Enterprise suits teams running AI across several teams. It builds on the open source Agent Router project, adding team budgets, access profiles, and cost attribution across a fleet of gateways. In an independent Broadcom VMware test, the Agent Router data plane added about 2 ms per request.

Choosing between LiteLLM and Kong AI Gateway usually means choosing which operational burden your team will carry once AI traffic grows past one team. This piece compares open source scope, performance, failover, plugin design, and operating cost.

Open source scope differs across LiteLLM, Kong, and Agent Router

Each tool draws the open source boundary in a different place, which changes what a team can run without a paid license.

LiteLLM keeps most gateway features in its MIT-licensed core

The core LiteLLM proxy is open source under the MIT license, apart from an enterprise directory that carries a paid license. The open source proxy includes provider translation and virtual keys. It also tracks spend per key, per user, and per team, with budgets at each level. LiteLLM Enterprise adds single sign-on (SSO) beyond five users, role-based access control (RBAC), audit logs, spend reports, and tag-based budgets.

Kong keeps most AI plugins in its Enterprise tier

Kong Gateway’s open source edition is Apache 2.0 licensed. It ships six open source AI plugins, including AI Proxy for provider routing and a prompt template plugin. Semantic caching, advanced load balancing, token-based rate limiting, and the MCP plugins require Kong AI Gateway Enterprise. Kong’s open source line has received only patch releases since version 3.9 in December 2024.

Kong Enterprise adds RBAC, audit logging, a Federal Information Processing Standards (FIPS) 140-3 compliant package, and dedicated support. Konnect, Kong’s managed control plane, is paid after a 30-day trial. Teams that stay fully self-managed can run open source Kong in hybrid mode with their own control plane. Kong’s strength is enterprise maturity. Teams already running Kong for API traffic can add AI routing without adding a new vendor. Kong launched its open source AI Gateway in February 2024 and added semantic caching to its Enterprise tier that September.

Agent Router Enterprise adds fleet-wide controls to open source Agent Router

Agent Router, formerly Envoy AI Gateway, is an open source AI gateway built on Envoy and now part of the Agentic AI Foundation. The open source project already covers multi-model routing with automatic provider fallback, token-based rate limits, and an MCP gateway for tool calls.

Agent Router Enterprise adds team-level controls across a fleet of gateways: cross-team cost attribution, team budgets, per-team access profiles, runtime guardrails, enterprise SSO, and audit export. Teams that outgrow the open source project add Agent Router Enterprise on the same Envoy data plane. Teams that already run Envoy or Istio already know the proxy Agent Router Enterprise is built on.

LiteLLM and Kong AI Gateway start from two different architectures

LiteLLM is a Python-first translation layer. Kong is a general API gateway with AI plugins added. Both work well at small scale, and each needs more operational work as AI traffic becomes business-critical.

LiteLLM’s default proxy scales by adding Python instances

LiteLLM is a Python proxy that translates between provider APIs. It solves API fragmentation well, with 140+ providers behind one OpenAI-compatible interface.

LiteLLM reports 8 ms P95 latency at 1,000 requests per second (RPS) across a deployment of two to four instances, where P95 is the 95th percentile. That figure depends on running several instances, since Python’s global interpreter lock limits CPU-bound parallelism inside one process. Some users running the Python path under sustained load have reported memory growth that needs worker recycling to contain. LiteLLM announced in June 2026 that it is moving its request hot path to Rust, with a staged rollout under way.

Kong AI Gateway runs on NGINX and Lua

Kong is built on OpenResty, which embeds Lua into NGINX. Each NGINX worker process runs its own Lua virtual machine, so Kong scales across CPU cores by adding workers. Its plugin ecosystem covers authentication, rate limiting, and request transformation through pre-built plugins.

Kong added streaming support for AI responses in version 3.7, on top of an architecture first built for REST APIs. Large language model (LLM) requests are long-lived and streaming, which puts different pressure on a gateway than short request-response calls. In September 2026, Kong also released AI Gateway 2.0 as a separate product. This comparison covers the AI plugins inside Kong Gateway.

Seven dimensions separate LiteLLM, Kong, and Agent Router Enterprise

DimensionLiteLLMKong AI GatewayAgent Router Enterprise
ArchitecturePython FastAPI proxyOpenResty/Lua on NGINXEnvoy, through the open source Agent Router project (formerly Envoy AI Gateway)
Published scale1,000 RPS across two to four instances (LiteLLM’s own benchmark)About 18,900 RPS on a mocked LLM backend (Kong’s own benchmark)Envoy handled all service-to-service traffic at Lyft by 2018
Latency figures8 ms P95 latency at 1,000 RPS across two to four instances (LiteLLM’s own benchmark)Total P95 of 46 ms against a 24 ms no-gateway baseline (Kong’s own benchmark)About 2 ms of added overhead, 0.01% of end-to-end latency (Broadcom VMware test of the Agent Router data plane)
Plugin systemPython middlewareLua, Go, JavaScript, and Python pluginsPolicy engine on live traffic
Provider coverage140+ providersAbout 17 providers natively through AI Proxy200+ models across provider families, through one endpoint that speaks the OpenAI and Anthropic APIs
FailoverConfigurable fallbacks, health checksHealth checks, circuit breakersAutomatic multi-provider failover
DeploymentSelf-hosted, with Postgres and RedisTraditional, DB-less, or hybrid, plus fully managed Konnect gatewaysFully managed or self-hosted data plane

The matrix shows each tool started from a different problem. LiteLLM starts from a narrower premise: normalize and route LLM traffic across providers. Kong starts from API management and adds AI capability through plugins.

For teams choosing between LiteLLM and Kong, the deciding question is which operational model matches their traffic. LiteLLM’s Python-first design makes setup fast. Kong’s API gateway heritage brings enterprise tooling. As AI traffic spreads across teams, both need extra configuration to apply the same budgets and access rules to every team.

Failover mechanisms differ between LiteLLM, Kong, and Agent Router Enterprise

LiteLLM, Kong, and Agent Router Enterprise all support failover, but each runs it at a different layer.

LiteLLM runs failover inside its Python proxy

LiteLLM applies automatic fallbacks in the order a team configures them. It also puts a failing deployment into a cooldown period before trying it again.

The limitation is structural. Routing logic and request handling share one Python event loop in each worker process, so failover decisions slow down along with everything else when that worker is under heavy load. Teams that run several LiteLLM instances share cooldown state through Redis, which adds a network round trip to each lookup.

Kong runs failover through upstream health checks

Kong load-balances across upstream targets with built-in active and passive health checks. Kong treats passive health checks as circuit breakers, so a target that keeps failing stops receiving traffic until it recovers. Balancing across several LLM targets needs the AI Proxy Advanced plugin, which is part of Kong AI Gateway Enterprise.

Agent Router Enterprise runs failover in the Envoy data plane

In Agent Router Enterprise, failover runs in the Envoy data plane. When a provider returns an error or is unreachable, the gateway moves to the next backend in the fallback chain within the same request. One provider outage affects one route, while every other team’s agents keep running.

Our approach to AI gateways builds on this Envoy foundation. In Agent Router Enterprise, the data plane keeps routing traffic even if the management plane is unreachable, which matters once the gateway becomes critical infrastructure.

LiteLLM, Kong, and Agent Router Enterprise translate provider APIs in different layers

All three gateways translate provider APIs. LiteLLM maps Bedrock, Azure, and Anthropic schemas into OpenAI format inside its Python process. Kong transforms requests and responses through plugins in its Lua layer. Agent Router Enterprise translates provider formats in the Envoy data plane.

Schema mapping adds maintenance work every time a provider changes its API, whichever gateway does it. The difference is who maintains the mappings and which layer carries that cost under load. Agent Router Enterprise exposes one endpoint that speaks both the OpenAI and Anthropic APIs, so applications written for either API connect without a translation step of their own.

Plugin architecture decides who maintains custom logic

Plugin architecture determines who owns the maintenance burden when requirements change.

LiteLLM runs custom middleware in the request path

LiteLLM’s Python middleware is accessible for teams familiar with Python. Custom middleware runs in the request path, and the team that writes it owns its maintenance when requirements change. LiteLLM’s open source proxy already ships virtual keys, spend tracking, and logging to more than 20 backends, so teams write middleware only for logic beyond those built-in features.

Kong plugins run in a shared Lua runtime

Kong plugins run in request-lifecycle phases, from access checks through response logging. All plugins in a worker process share one Lua virtual machine, and they pass per-request data to each other through a shared context. A plugin doing heavy work on a long-lived LLM stream competes for CPU with every other request on that worker. Custom plugins can be written in Lua, Go, JavaScript, or Python, with Lua as the native option. Kong’s AI plugins date from 2024, so they have a shorter production track record than the API management plugins the ecosystem grew from.

Custom routing and attribution logic require ongoing maintenance in either gateway

A team building custom routing logic, cost attribution, or guardrails owns the maintenance in either ecosystem. Agent Router Enterprise builds budgets, rate limits, and access controls into the gateway. The MCP Gateway applies the same policy to tool calls.

LiteLLM and Kong AI Gateway create different operational demands

The operational question is who manages the deployment, who patches it, and who scales it when traffic grows.

LiteLLM and Kong deploy on different state models

LiteLLM deploys as a Python service, with Postgres storing keys, budgets, and spend. Redis shares rate limits and cooldowns across instances. Kong AI Gateway runs in traditional, DB-less, or hybrid mode. Konnect adds a Kong-hosted control plane, and Kong’s Dedicated Cloud Gateways run fully managed data planes.

Agent Router Enterprise runs fully managed or with a Self-Hosted Data Plane. In the self-hosted model, Tetrate hosts the management plane while the customer runs the data plane where traffic flows. Teams with strict data residency requirements run the data plane in their own cloud account or on-premises, so prompts stay inside their network.

LiteLLM, Kong, and Envoy scale through different infrastructure patterns

LiteLLM’s production guidance is one Uvicorn worker per pod, with pods autoscaled on CPU. Scaling Kong means managing a distributed gateway topology, with or without a database. LLM traffic adds long-lived connections, streaming responses, and token-level observability to that topology.

Agent Router Enterprise is built on Envoy. Envoy began at Lyft, where by 2018 it handled all service-to-service traffic. Netflix adopted Envoy for its service mesh. Airbnb’s Envoy-based Istio mesh handles tens of millions of queries per second at peak.

Gateway maintenance consumes engineering time that could build product

The real cost is engineering time: every hour spent maintaining a gateway is an hour taken from building product. A large fintech company with thousands of developers retired home-grown gateways after deploying Agent Router Enterprise, freeing the engineering time previously spent maintaining them. Model access widened from OpenAI and Anthropic models reached through Azure AI Foundry and Amazon Bedrock to Azure AI Foundry, Amazon Bedrock, Google Vertex AI, Mistral, and private open-weight models, all without application rewrites.

Kong AI Gateway outperforms LiteLLM in two contexts

Kong leads LiteLLM in two areas: throughput and scope.

Kong’s NGINX runtime led LiteLLM on throughput in Kong’s own benchmark

In a July 2025 benchmark Kong published, self-managed Kong Gateway Enterprise reached about 18,900 RPS on a mocked LLM backend, against about 1,700 RPS for LiteLLM on the same setup. The test was run by Kong on its own setup, with every gateway on default settings, and predates LiteLLM’s Rust work.

Kong covers API and AI traffic under one gateway

Kong’s advantage over LiteLLM is scope. One Kong deployment governs REST APIs and LLM traffic with the same plugins. Authentication, rate limiting, and logging work the same way for both. LiteLLM’s open source proxy already ships tokens per minute (TPM) limits and semantic caching, so those features do not separate the two. In Kong, both need an AI Gateway Enterprise license.

Both gateways need platform work to run production AI traffic

The signs that AI infrastructure is failing are specific: unexplained coding agent spend, provider outages paging on-call, token spend as one lump charge against a single API key, multi-agent failures that take a long time to debug, and a do-it-yourself (DIY) gateway becoming a maintenance burden.

LiteLLM and Kong trace agent traffic through external observability tools

LiteLLM groups calls from one agentic flow by session or trace ID and exports them to tools like Langfuse. Kong exports AI token, cost, and latency metrics through Prometheus or OpenTelemetry once teams enable them. In both cases, the platform team connects and maintains the observability pipeline.

Production AI traffic needs five infrastructure capabilities

Production AI traffic needs five capabilities:

  • Automatic multi-provider failover keeps agents running when a single provider has an outage, without paging every on-call team at once.
  • Per-team and per-agent cost attribution turns one lump API charge into a breakdown finance can act on.
  • A self-serve developer on-ramp lets a new agent connect in minutes through a base URL change, without a platform team ticket.
  • Agent debugging with exportable traces makes a multi-agent failure diagnosable in one place.
  • Distributed deployment runs the data plane where the traffic flows, so prompts can stay in the region where they originate.

An enterprise AI gateway delivers these as infrastructure. LiteLLM and Kong each cover parts of this list, with some features in paid tiers. Running all five across teams, clouds, and regions is where platform work grows as agent count grows.

Agent Router Enterprise differs most in tracing, attribution, and migration

The two sections below cover where Agent Router Enterprise differs from LiteLLM and Kong in practice.

Agent Router Enterprise logs tool calls across agent chains

When a multi-agent workflow fails, engineers must piece together what happened across provider dashboards, which is slow when logs are fragmented. Agent Router Enterprise’s agent-aware tracing logs every tool call across an agent chain. It exports per-request traces over OpenTelemetry to existing observability tools. Engineers see the full chain in one place.

Agent Router Enterprise connects through a single base URL change

Agent Router Enterprise exposes an OpenAI-compatible endpoint, so existing code connects by changing one base URL. The Fully Managed fast-track evaluation reaches a first routed request in about 15 minutes after onboarding. In an independent Broadcom VMware test under sustained enterprise LLM load, the Agent Router data plane added roughly 2 ms per request, about 0.01% of end-to-end latency.

Agent Router Enterprise attributes every token to a team or agent

Agent Router Enterprise’s cost budgeting capabilities attribute every token to a person, team, agent, or project. Finance can trace each charge to the team that made it. Engineering leaders can use the same breakdown to debug. When an agent’s token spend spikes unexpectedly, attribution data shows which agent, team, or model call drove it, which cuts the time spent correlating provider logs.

The right gateway depends on team stage and existing infrastructure

Pain pointLiteLLMKong AI GatewayAgent Router Enterprise
DIY maintenance burdenHigh at scaleModerateFully managed option
Unexplained coding agent spendPer-team spend plus User-Agent taggingToken and cost metrics via PrometheusPer-team and per-agent attribution
Provider outage exposureConfigurable fallbacksHealth checks and circuit breakersAutomatic multi-provider failover
Multi-agent debuggingSession and trace IDs with OpenTelemetry exportAI metrics via Prometheus or OpenTelemetryTool-call logging across agent chains
Fits these constraintsEarly-stage teams, lower traffic, broad provider coverageTeams already on Kong for APIsTeams that have outgrown both

Choosing between LiteLLM and Kong AI Gateway depends on organizational maturity. LiteLLM fits teams optimizing for developer velocity. Kong fits teams with existing API management infrastructure. Agent Router Enterprise fits teams where AI traffic has become business-critical and the gateway itself is now a production dependency.

Now Available

MCP Catalog with verified first-party servers, profile-based configuration, and OpenInference observability are now generally available in Tetrate Agent Router Service. Start building production AI agents today with $5 free credit.

Sign up now

Conclusion

LiteLLM and Kong AI Gateway each address part of the AI gateway problem. LiteLLM is fast to start and scales by adding Python instances. Kong is enterprise-mature, with an architecture centered on API management. Its semantic caching and AI rate limiting sit in the Enterprise tier. The deciding question is whether the gateway can run production AI traffic across teams without becoming a distributed-systems problem the platform team owns. For teams that have outgrown either option, Agent Router Enterprise adds fleet-wide budgets, access profiles, and cost attribution on top of the open source Agent Router data plane.

Start a fast-track evaluation: first routed request in about 15 minutes, nothing to install. Request a demo of Agent Router Enterprise.

FAQs

Is LiteLLM or Kong better for multi-cloud deployments?

Kong offers more managed deployment options than LiteLLM, including fully managed Dedicated Cloud Gateways in Konnect. LiteLLM scales horizontally with multiple instances behind a load balancer, using Postgres and Redis for shared state across instances. Agent Router Enterprise runs a data plane in any cloud, on-premises, or at the edge, with one central management plane. Teams enforce policy once and route traffic locally in each region.

Can LiteLLM handle enterprise-scale traffic?

Yes, when it runs as several instances. LiteLLM reports 8 ms P95 latency at 1,000 RPS across a multi-instance deployment. Its default Python proxy scales by adding instances, and LiteLLM is moving its hot path to Rust to raise per-instance throughput.

Does Kong AI Gateway support all LLM providers?

No. Kong’s AI Proxy plugin supports about 17 providers natively through configuration, including OpenAI, Anthropic, Amazon Bedrock, and Mistral. Providers outside that list need the pass-through mode or a custom plugin.

Which gateway has lower operational overhead?

Agent Router Enterprise in Fully Managed mode has nothing to install, since Tetrate runs both planes. LiteLLM is quick to start, and its operational work grows with instance count. Kong has mature operational tooling and runs self-managed, hybrid, or fully managed through Konnect. In the Agent Router Enterprise self-hosted model, Tetrate hosts the management plane while you run the data plane in your own environment.

Do I need to rewrite code to switch gateways?

No. Agent Router Enterprise is OpenAI-compatible. You change one base URL, and existing code connects without modification.

Key terms glossary

API gateway: Infrastructure that sits between clients and backend services. It handles routing, authentication, rate limiting, and request transformation for API traffic. Kong AI Gateway is a general API gateway with AI capability added through plugins.

AI gateway: Infrastructure that routes, governs, and pays for AI requests. It works across models, providers, teams, and locations. A production-grade AI gateway enforces policy on live traffic.

Agent Router: An open source AI gateway project built on Envoy, formerly Envoy AI Gateway and now part of the Agentic AI Foundation. Tetrate Agent Router Enterprise runs on the Agent Router data plane.

RPS (requests per second): A measure of gateway throughput, the number of requests a gateway handles in one second.

Cost attribution: Assigning AI spend to the person, team, agent, or project that generated it.

Access profile: A per-team set of the models and MCP tools that team’s agents are allowed to use.

Token-based rate limiting: Rate limiting that counts the tokens each request consumes, such as tokens per minute (TPM).

OpenAI-compatible: An API interface that accepts requests formatted for OpenAI’s API, enabling drop-in replacement without code changes.

Base URL: The endpoint address a client uses to reach an API. Changing the base URL redirects traffic without modifying request logic.

Failover: Automatic redirection of traffic from a failed provider or route to a healthy backup.

Token spend: The cost of AI usage measured in tokens, the billing unit for LLM API calls.

MCP (Model Context Protocol): A protocol for connecting AI agents to tools and data sources.

Agent-aware tracing: Tracing that logs every tool call and model call across an agent chain, then exports per-request traces to existing observability tools.

OpenResty: A web platform built on NGINX that embeds the Lua programming language, allowing developers to extend NGINX functionality through Lua scripts.

FastAPI: A modern Python web framework used to build APIs, with automatic API documentation generation.

Hot path: The code that runs on every request, such as routing, auth, and translation.

Circuit breaker: A pattern that prevents cascading failures by stopping requests to a failing service and allowing it time to recover.

Semantic caching: Caching that matches requests by meaning as well as exact text, allowing similar queries to reuse cached results.

Schema mapping: The process of translating data structures from one format to another, used when normalizing different provider APIs into a unified interface.

Envoy: A high-performance C++ distributed proxy that runs in front of single services or as the data plane of a large service mesh.

Data plane: The infrastructure layer that carries live traffic. In Tetrate’s architecture, the data plane keeps routing even if the management plane is unreachable.

Management plane: The central layer where teams set policy and configuration for every data plane. Tetrate hosts the management plane. Customers can run data planes in their own environment.

Plugin architecture: A system for extending gateway functionality through modular components, typically middleware or scripts that execute in the request path.


MCP Catalog with verified first-party servers, profile-based configuration, and OpenInference observability are now generally available in Tetrate Agent Router Service . Start building production AI agents today.

Decorative CTA background pattern background background
Tetrate logo in the CTA section Tetrate logo in the CTA section for mobile

Ready to enhance your
network

with more
intelligence?