Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

Per-Customer AI Cost: How SaaS Teams Meter, Limit, and Price AI Features With an AI Gateway

Last updated: October 2026

Tetrate Agent Router Enterprise tags every model call with the customer who caused the call. You can price AI features on real usage, before a few heavy customers cost more than they pay.

Per-customer AI cost is the model spend that each of your own customers causes inside your SaaS product. Model providers bill you per provider, and your customers pay you per customer, so the provider invoice can’t show which customers are profitable. Tetrate Agent Router Enterprise sits between your product and the model providers. The gateway records which customer caused each model call, reports cost per customer, and lets you set a spending ceiling for each customer.

Setting up per-customer AI cost takes five steps. Tag every call with a customer. Read cost per customer. Match model access to each plan. Cap each customer’s spend. Keep each customer’s data and uptime promises.

Step 1: Tag Every Model Call With the Customer Who Caused the Call

Every request your product sends needs a customer ID. Agent Router Enterprise gives you three ways to attach one:

MethodHow the method worksBest for
Customer headerYour app sends x-tars-customer on each request. The gateway records the value as the tars.customer attribute on every request’s telemetry, per the OpenTelemetry reference.Many small customers, with cost dashboards in Grafana or Datadog
Tag on the API keyEach key carries an approved tag such as customer=acme. Usage reports in the Admin Console filter and export by tag, per the API tags reference.Customers grouped by plan, region, or partner
One API key per customerEach large customer gets a dedicated key. Usage reports group by key, and each key can carry its own budget.Large customers, and any customer with a spending cap

You can combine all three. Most SaaS teams send the header on every call and give their largest customers a dedicated key as well.

Roy Prins describes the same pattern in his explainer on AI gateways: tag every request with a tenant ID so you can bill by usage.

Two tips. Agree on the tag values with finance before you issue keys, so customer=acme and Customer=ACME never split one customer into two rows. And set the key or project to require the tag, so every request arrives labeled, per the guide to attributing cost with tags.

Step 2: Read Cost per Customer Before You Change Any Price

Open Usage Analytics in the Admin Console. Filter by your customer tag or group by API key, and read cost and tokens side by side.

Effective cost per million tokens is the most useful number on the page. Effective cost per million tokens is a customer’s cost divided by the customer’s tokens, times one million. Two customers with the same token count can have very different effective costs, because one uses a frontier model and the other uses a small one.

What the report showsWhat the pattern meansWhat to do next
A few customers carry most of the costYour pricing treats heavy and light users the sameAdd a usage tier or an overage charge
One customer has a high effective cost per millionThat customer’s requests reach your most expensive modelCheck whether the workload needs that model
Cost per customer climbs while tokens stay flatA model or price change raised the costReview the model behind that customer’s requests
Tokens per customer climb steadilyCustomers use the feature moreMake sure the plan price rises with usage

The CSV export on the same page gives finance one row per tag value per period, with cost in USD and token totals. Cost uses each model’s published price, and per-model price overrides cover any negotiated rates. Finance applies each customer’s share to the real invoice.

Step 3: Match Model Access to the Plan the Customer Pays For

A free trial and an enterprise plan should run on different models. Agent Router Enterprise gives each plan its own project. Each project has its own gateway address, its own keys, and its own list of models, per the key concepts page.

PlanProject setupWhat the customer gets
Free trialA project with one small, fast modelGood answers at the lowest cost per request
StandardA project with a mid-size model and a backup modelSteady quality, with failover
EnterpriseA project with frontier models and stricter safety rulesYour best models, with PII redaction on every request

Moving a plan to a newer or cheaper model is one setting. A traffic split at 100% sends every request on a key to the model you pick, while your product code keeps asking for the same model name. A 10% split lets you test a cheaper model on real customer traffic first.

Two tips. Keep each plan’s model list short, so a support ticket about model behavior points to one place. And test a model change on your own internal key before any customer key, so your team sees the new model first.

For more on projects, roles, and the model catalog, see AI model access control.

Step 4: Put a Ceiling on Each Customer’s Spend

A budget decides what happens when a customer reaches a spending limit. Put a budget on each large customer’s API key and pick one of three actions, per the budget guide:

Watch spendHard stopDegrade gracefully
At the limitAn alert appears on the budgetRequests stop until the period resetsRequests move to a cheaper model you picked
The customer seesNo changeAn error that names the budgetAnswers keep coming, from the cheaper model
Use forCustomers on usage-based billingFree trials and fixed-price plansPaying customers you never want to cut off

Degrade gracefully fits most paid plans. The customer keeps working, your margin stops shrinking, and your account team can plan the pricing conversation.

Budgets reset daily, weekly, or monthly, so match the period to your billing cycle. The budget page shows used, remaining, and days left for each key, so your account team can see which customers are close to their plan’s limit.

For runaway loops, add a tokens-per-hour rate limit on the same key. A rate limit checks every request as the request arrives, and a budget acts on totals a few minutes later. Together, the two controls cover a sudden spike and a slow climb.

Step 5: Keep Each Customer’s Data and Uptime Promises

Your customer contracts also include data and uptime terms. Agent Router Enterprise maps a control to each one.

PromiseAgent Router Enterprise controlDocs
Customer data stays privateGuardrails redact PII and secrets inline, per projectDetect and redact sensitive data
The feature stays up during a provider outageA fallback chain moves failed requests to a backup modelFallback
Requests stay in the customer’s regionRouting policy sends requests only to providers in the allowed regionData residency
You can show what happenedAn audit log records every admin change, and request logs record every callAudit log

Use Per-Customer Cost to Choose a Pricing Model

Once every call carries a customer ID, you can compare pricing models with real numbers.

Pricing modelData you need from the gateway
Flat price per seat, with a fair-use limitCost per customer, to set the limit above typical use
Included token allowance, then overageTokens per customer per period, from the CSV export
Separate AI add-onEffective cost per million, to set a price above your cost
Plan tiers with different modelsCost per plan, from one project per plan

How to Get Started

Agent Router Enterprise ships every control on this page today. That list covers customer headers, key tags, per-key budgets, tokens-per-hour rate limits, projects per plan, traffic splitting, guardrails, and fallback. Start with the guide to attributing cost with tags, then see the cost and token visibility page. To see per-customer cost on a live gateway, request a demo.

Agent Router Enterprise

Tetrate Agent Router Enterprise routes AI agent traffic across providers and your own models, with policy, cost controls, and audit on every request — in cloud, on-prem, or edge.

Learn more

Frequently asked questions

What is per-customer AI cost? Per-customer AI cost is the model spend that each of your own customers causes inside your SaaS product. An AI gateway measures per-customer AI cost by recording a customer ID on every model call.

How do I track AI cost per customer in a SaaS product? Send a customer ID with every model call, either as a request header or through a tagged API key per customer. The gateway records the ID on each request, and usage reports show cost and tokens for each customer.

What is effective cost per million tokens? Effective cost per million tokens is a customer’s AI cost divided by the customer’s token count, times one million. The number shows which customers reach expensive models, even when two customers use the same number of tokens.

How do I cap AI spend for one customer? Give the customer a dedicated API key and put a budget on the key. Pick Hard stop to stop requests at the limit, or Degrade gracefully to move the customer to a cheaper model and keep the feature running.

How do I give free and paid plans different AI models? Put each plan in its own project and grant each project only the models for that plan. Customer keys from the free plan then reach only the free plan’s models.

How do I stop one customer’s runaway usage in real time? Add a tokens-per-hour rate limit to the customer’s API key. The gateway checks the limit on every request as the request arrives.

Can I export AI cost per customer for billing? Yes. The usage report exports a CSV with one row per customer tag per period, including cost in USD and token totals. Finance applies each customer’s share to the provider invoice.


Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.

Decorative CTA background pattern background background
Tetrate logo in the CTA section Tetrate logo in the CTA section for mobile

Ready to enhance your
network

with more
intelligence?