Per-Customer AI Cost: How SaaS Teams Meter, Limit, and Price AI Features With an AI Gateway
Last updated: October 2026
Tetrate Agent Router Enterprise tags every model call with the customer who caused the call. You can price AI features on real usage, before a few heavy customers cost more than they pay.
Per-customer AI cost is the model spend that each of your own customers causes inside your SaaS product. Model providers bill you per provider, and your customers pay you per customer, so the provider invoice can’t show which customers are profitable. Tetrate Agent Router Enterprise sits between your product and the model providers. The gateway records which customer caused each model call, reports cost per customer, and lets you set a spending ceiling for each customer.
Setting up per-customer AI cost takes five steps. Tag every call with a customer. Read cost per customer. Match model access to each plan. Cap each customer’s spend. Keep each customer’s data and uptime promises.
Step 1: Tag Every Model Call With the Customer Who Caused the Call
Every request your product sends needs a customer ID. Agent Router Enterprise gives you three ways to attach one:
| Method | How the method works | Best for |
|---|---|---|
| Customer header | Your app sends x-tars-customer on each request. The gateway records the value as the tars.customer attribute on every request’s telemetry, per the OpenTelemetry reference. | Many small customers, with cost dashboards in Grafana or Datadog |
| Tag on the API key | Each key carries an approved tag such as customer=acme. Usage reports in the Admin Console filter and export by tag, per the API tags reference. | Customers grouped by plan, region, or partner |
| One API key per customer | Each large customer gets a dedicated key. Usage reports group by key, and each key can carry its own budget. | Large customers, and any customer with a spending cap |
You can combine all three. Most SaaS teams send the header on every call and give their largest customers a dedicated key as well.
Roy Prins describes the same pattern in his explainer on AI gateways: tag every request with a tenant ID so you can bill by usage.
Two tips. Agree on the tag values with finance before you issue keys, so customer=acme and Customer=ACME never split one customer into two rows. And set the key or project to require the tag, so every request arrives labeled, per the guide to attributing cost with tags.
Step 2: Read Cost per Customer Before You Change Any Price
Open Usage Analytics in the Admin Console. Filter by your customer tag or group by API key, and read cost and tokens side by side.
Effective cost per million tokens is the most useful number on the page. Effective cost per million tokens is a customer’s cost divided by the customer’s tokens, times one million. Two customers with the same token count can have very different effective costs, because one uses a frontier model and the other uses a small one.
| What the report shows | What the pattern means | What to do next |
|---|---|---|
| A few customers carry most of the cost | Your pricing treats heavy and light users the same | Add a usage tier or an overage charge |
| One customer has a high effective cost per million | That customer’s requests reach your most expensive model | Check whether the workload needs that model |
| Cost per customer climbs while tokens stay flat | A model or price change raised the cost | Review the model behind that customer’s requests |
| Tokens per customer climb steadily | Customers use the feature more | Make sure the plan price rises with usage |
The CSV export on the same page gives finance one row per tag value per period, with cost in USD and token totals. Cost uses each model’s published price, and per-model price overrides cover any negotiated rates. Finance applies each customer’s share to the real invoice.
Step 3: Match Model Access to the Plan the Customer Pays For
A free trial and an enterprise plan should run on different models. Agent Router Enterprise gives each plan its own project. Each project has its own gateway address, its own keys, and its own list of models, per the key concepts page.
| Plan | Project setup | What the customer gets |
|---|---|---|
| Free trial | A project with one small, fast model | Good answers at the lowest cost per request |
| Standard | A project with a mid-size model and a backup model | Steady quality, with failover |
| Enterprise | A project with frontier models and stricter safety rules | Your best models, with PII redaction on every request |
Moving a plan to a newer or cheaper model is one setting. A traffic split at 100% sends every request on a key to the model you pick, while your product code keeps asking for the same model name. A 10% split lets you test a cheaper model on real customer traffic first.
Two tips. Keep each plan’s model list short, so a support ticket about model behavior points to one place. And test a model change on your own internal key before any customer key, so your team sees the new model first.
For more on projects, roles, and the model catalog, see AI model access control.
Step 4: Put a Ceiling on Each Customer’s Spend
A budget decides what happens when a customer reaches a spending limit. Put a budget on each large customer’s API key and pick one of three actions, per the budget guide:
| Watch spend | Hard stop | Degrade gracefully | |
|---|---|---|---|
| At the limit | An alert appears on the budget | Requests stop until the period resets | Requests move to a cheaper model you picked |
| The customer sees | No change | An error that names the budget | Answers keep coming, from the cheaper model |
| Use for | Customers on usage-based billing | Free trials and fixed-price plans | Paying customers you never want to cut off |
Degrade gracefully fits most paid plans. The customer keeps working, your margin stops shrinking, and your account team can plan the pricing conversation.
Budgets reset daily, weekly, or monthly, so match the period to your billing cycle. The budget page shows used, remaining, and days left for each key, so your account team can see which customers are close to their plan’s limit.
For runaway loops, add a tokens-per-hour rate limit on the same key. A rate limit checks every request as the request arrives, and a budget acts on totals a few minutes later. Together, the two controls cover a sudden spike and a slow climb.
Step 5: Keep Each Customer’s Data and Uptime Promises
Your customer contracts also include data and uptime terms. Agent Router Enterprise maps a control to each one.
| Promise | Agent Router Enterprise control | Docs |
|---|---|---|
| Customer data stays private | Guardrails redact PII and secrets inline, per project | Detect and redact sensitive data |
| The feature stays up during a provider outage | A fallback chain moves failed requests to a backup model | Fallback |
| Requests stay in the customer’s region | Routing policy sends requests only to providers in the allowed region | Data residency |
| You can show what happened | An audit log records every admin change, and request logs record every call | Audit log |
Use Per-Customer Cost to Choose a Pricing Model
Once every call carries a customer ID, you can compare pricing models with real numbers.
| Pricing model | Data you need from the gateway |
|---|---|
| Flat price per seat, with a fair-use limit | Cost per customer, to set the limit above typical use |
| Included token allowance, then overage | Tokens per customer per period, from the CSV export |
| Separate AI add-on | Effective cost per million, to set a price above your cost |
| Plan tiers with different models | Cost per plan, from one project per plan |
How to Get Started
Agent Router Enterprise ships every control on this page today. That list covers customer headers, key tags, per-key budgets, tokens-per-hour rate limits, projects per plan, traffic splitting, guardrails, and fallback. Start with the guide to attributing cost with tags, then see the cost and token visibility page. To see per-customer cost on a live gateway, request a demo.
Agent Router Enterprise
Frequently asked questions
What is per-customer AI cost? Per-customer AI cost is the model spend that each of your own customers causes inside your SaaS product. An AI gateway measures per-customer AI cost by recording a customer ID on every model call.
How do I track AI cost per customer in a SaaS product? Send a customer ID with every model call, either as a request header or through a tagged API key per customer. The gateway records the ID on each request, and usage reports show cost and tokens for each customer.
What is effective cost per million tokens? Effective cost per million tokens is a customer’s AI cost divided by the customer’s token count, times one million. The number shows which customers reach expensive models, even when two customers use the same number of tokens.
How do I cap AI spend for one customer? Give the customer a dedicated API key and put a budget on the key. Pick Hard stop to stop requests at the limit, or Degrade gracefully to move the customer to a cheaper model and keep the feature running.
How do I give free and paid plans different AI models? Put each plan in its own project and grant each project only the models for that plan. Customer keys from the free plan then reach only the free plan’s models.
How do I stop one customer’s runaway usage in real time? Add a tokens-per-hour rate limit to the customer’s API key. The gateway checks the limit on every request as the request arrives.
Can I export AI cost per customer for billing? Yes. The usage report exports a CSV with one row per customer tag per period, including cost in USD and token totals. Finance applies each customer’s share to the provider invoice.
Related reading
- AI Cost Visibility, on cost by developer, team, and customer
- What is an AI gateway, Roy Prins on tagging requests by tenant
- Announcing token brokering, on budgets in the request path
Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.