Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

Enabling agentic developers with Tetrate Agent Router

An end-to-end walkthrough of enabling developers with AI coding assistants — granting providers to a project, putting budgets around team spend, and getting a developer from zero to a governed model catalog inside VS Code.

Enabling agentic developers with Tetrate Agent Router

Tetrate Agent Router lets a platform team grant LLM providers and set team budgets once; developers then self-serve an inference key and get a governed model catalog inside VS Code — with live spend next to the code.

Every engineering organization is having the same conversation right now. Developers want AI coding assistants — in the editor, in the terminal, in their agents — and they want model choice, because the model that plans a refactor well is rarely the cheapest one to draft commit messages with. Meanwhile the people accountable for spend and security want answers to unglamorous questions: which providers are approved, who is spending what, and what happens when an agent left running overnight burns through a month of tokens before breakfast.

Most organizations answer those questions with process: a procurement ticket per provider, an API key handed over in a password manager, a spreadsheet nobody updates. That does not scale to a world where every developer runs one or more agents, and it puts governance in the critical path of enablement — exactly where it breeds shadow AI.

Agent Router inverts that. Providers, model policy, budgets, and keys live in one gateway; developers serve themselves within the limits the organization has set. This post walks the whole path end to end, wearing two hats along the way: first as a platform operator in the Admin Console granting providers and setting team budgets, then as a developer who creates a key, wires up Visual Studio Code, and starts working.

The running example: a project called CRM-implementation, staffed by two existing teams, Backend Engineering and UI Engineering.

Step 1: Grant LLM providers to the project

Agent Router Enterprise scopes access in two layers. The organization catalog determines which providers are connected and which models are enabled at all. A project grant determines which of those a specific project may use. A request only succeeds when both layers agree, which is what makes a project a real isolation boundary rather than a folder.

In the Admin Console, everything for our example hangs off the project. Under Directory → Projects, the CRM-implementation project already exists, with six users and an attached gateway — the endpoint that developers, agents, and tools will send LLM and MCP requests to.

Tetrate Agent Router Admin Console Projects page, showing the CRM-implementation and default projects, each with six users and an attached gateway, alongside in-product definitions of project, gateway, and data plane
Projects in the Admin Console. Each project bundles approved models, MCP servers, policies, and API keys behind exactly one gateway.

Opening the project and switching to the Providers tab shows what the project can reach today: Google Gemini and OpenAI, both enabled. A subtle but important default is worth calling out — a project with an empty provider list is unrestricted and may use any provider enabled in the organization catalog. Adding providers to the list is how you confine a project, for contractual or data-residency reasons, to a deliberate subset.

The Providers tab of the CRM-implementation project, listing Google Gemini and OpenAI as enabled providers, with an Add to project button
The CRM-implementation project starts with Google Gemini and OpenAI granted.

The CRM team has asked for Anthropic models for agentic coding and xAI for evaluation work. Selecting + Add to project opens a picker: search or browse the catalog, tick providers, and they collect in a Shortlist panel so a larger change stays reviewable before it lands.

The add-provider dialog with Anthropic and xAI ticked, both listed in a shortlist panel showing a count of two, with Add and Cancel buttons
Anthropic and xAI ticked and sitting in the Shortlist, one Add away from being granted.

Confirming with Add lands both grants; the table refreshes to four enabled providers and a toast confirms “Added 2 providers”.

The project Providers tab now listing Anthropic, Google Gemini, OpenAI, and xAI, all enabled, with a green toast reading Added 2 providers
Four providers granted to the project. Models are granted the same way from the project's Models page.

Models follow the same pattern from the project’s Models page — with one difference in behavior: models are strictly allow-listed, so a request naming an ungranted model is refused at request time. That refusal is the enforcement mechanism, not a UI convention. The full reference is in the add providers and models to projects guide.

Step 2: Put budgets around team spend

Enablement without spend control is how AI programs get frozen by finance in their first quarter. Agent Router Enterprise treats budgets as first-class policy: under Usage → Budgets, a budget caps spend per team or per API key, on a daily, weekly, or monthly period.

The empty Budgets page in the Admin Console showing a wallet icon, the heading Set a spending limit, and an Add budget button
Budgets live under Usage → Budgets in the Admin Console.

Selecting + Add budget opens the New budget drawer. Three scopes are on offer, and the distinction matters more than it first appears:

  • Whole team — one shared limit pooled across members. No individual caps; the team self-regulates inside one number.
  • Each teammate — every member gets the same limit individually. One member running hot stops only their own keys, not the whole team.
  • One API key — a limit on a single credential, which is how you cap a specific workload or an individual user.

For UI Engineering we pool the budget: scope Whole team, team UI Engineering, and a $1,000 monthly limit set from the suggested chips. The Policy preview panel on the right is a quietly excellent piece of design — it restates the policy in plain terms (“$1,000.00 / month for the UI Engineering team”) and charts the team’s actual billed spend for the last six months against a dashed limit line, so you sanity-check the number against reality before creating anything.

The New budget drawer with Whole team scope selected, UI Engineering chosen, a $1,000 monthly spend limit, and a policy preview showing six months of spending history under a dashed limit line
The Policy preview validates the limit against six months of real spend before the budget exists.

Next, the enforcement decision, under When the limit is hit:

  • Watch spend (the default) alerts when spend passes the limit but never blocks a request — a meter, useful for trialing a number before enforcing it.
  • Hard stop rejects requests once the cap is hit, until the period resets. Nothing needs to be manually re-enabled afterward.
  • Degrade gracefully reroutes to a cheaper fallback model, and is available only for the one-API-key scope — the UI grays it out for team budgets and says why.

UI Engineering gets a Hard stop: the team asked for a firm ceiling rather than an alert someone might read on Monday. With a name in place — the drawer suggests one if naming things is not your favorite job — the footer switches to “Ready. Create makes exactly this,” which is precisely the sort of no-surprises confirmation you want from a policy tool.

The completed budget form showing the three enforcement options with Hard stop selected, the name UI Eng Team monthly budget, and a policy preview confirming $1,000 per month for the UI Engineering team
Watch spend, Hard stop, or Degrade gracefully. UI Engineering gets a hard monthly ceiling.

Backend Engineering gets the other model: scope Each teammate, so every member carries their own $1,000 monthly limit with a Hard stop. The preview switches accordingly — instead of pooled history it charts the highest member’s spend, which is the number that actually threatens a per-person limit.

The New budget drawer with the Each teammate scope selected, showing helper text that every member gets the limit individually and a preview panel titled Highest member's spend
Each teammate scope: one member running over stops only their own keys.

A few minutes of this and the Budgets page tells the whole story at a glance — pooled budget for UI Engineering, per-member budgets for Backend Engineering (plus a Watch-spend meter we added for the QA team while we were in there), all healthy, with near-limit and over-limit counters at the top for the day that changes.

The Budgets overview listing three healthy budgets: QA Team budget at $0 of $500, Backend Eng team member budget with all two members within limit, and UI Eng Team monthly budget at $0 of $1K
The budget overview: pooled, per-member, and watch-only budgets side by side, each showing period spend against its limit.
Enforcement is eventually consistent — by design

Budget enforcement runs on usage rollups, so it can lag actual spend by a few minutes, and it fails open: requests are served while totals catch up. That is the right trade-off for a spend ceiling, and the wrong tool for an emergency brake — for an immediate inline cutoff, use a rate limit instead. Details in the budgets guide.

Step 3: The developer creates their own API key

Hat swap. Everything up to here was the platform operator in the Admin Console. From this point on, we are a developer on the CRM project — and notably, we will not file a single ticket.

Developers get their own home: the Developer Console. It carries a dashboard, a playground, request logs, model and MCP discovery, and — what we came for — Settings → API Keys. One detail to get right before anything else: the project switcher at the top of the sidebar. API keys and endpoints are project-scoped, so switch to CRM-implementation first. The reward is at the top of the API Keys page: the project’s Base URL, ready to copy, with a note that it works with the OpenAI SDK or any compatible client.

The Developer Console API Keys page for the CRM-implementation project, showing the Base URL panel with the project endpoint and an Add API Key button above the key list
The API Keys page in the Developer Console: the project's Base URL on top, self-service key creation below.

Selecting + Add API Key opens the Create new API key dialog. The key type defaults to Inference Key — the kind that authenticates LLM traffic — and all it needs is a name it can be recognized by later in dashboards and logs. We go with roy-inference-key and select Create key.

The Create new API key dialog with key type Inference Key and the name roy-inference-key entered, with Cancel and Create key buttons
Key type Inference Key, a recognizable name, and Create key. That is the entire request process.

The confirmation screen shows the key exactly once — “Keep this key secure. You won’t be able to see it again” — together with the Base URL, so both values a client needs are in one place with copy buttons next to them. Copy both now.

The API Key created dialog showing the new key with its value blurred, the project Base URL beneath it, and What's next cards for Playground, Integrations, Code Examples, and Configure
The one and only sighting of the key (blurred here, obviously), with the Base URL alongside. After Done, only the masked suffix remains.

Back on the list, the new key appears with a masked suffix, its creation date, and a last-used column — which is also where an operator would later spot abandoned credentials worth revoking.

The API Keys list showing roy-inference-key as a new inference key with a masked secret alongside the existing project management key
Self-service, but not invisible: every key is listed, masked, and tracked by last use.

Because the endpoint is OpenAI-compatible, the key is verifiable from any terminal in seconds:

Smoke test: list the models this key can reach
export TETRATE_API_KEY="sk-..."   # the key you just copied

curl -s https://proxy.tare-dummy-url.tetrate.ai/v1/models \
  -H "Authorization: Bearer $TETRATE_API_KEY"

The response lists every model the project grants from Step 1 allow through — which is the two-layer model doing its job on a live credential.

Step 4: Wire up Visual Studio Code

Here is the before picture. A fresh VS Code chat panel offers “Auto” and a handful of built-in models locked behind Upgrade links. Other extensions that want a model each ask for their own key. This is the credential sprawl the gateway exists to end.

A fresh VS Code window with the chat model dropdown open, showing only Auto available and Claude and GPT entries grayed out behind Upgrade links
Before: the model picker offers upgrades, not models.

The Tetrate Agent Router Model Provider extension fixes this at the right layer. Rather than pointing just the chat view at a custom endpoint, it registers as a VS Code language model chat provider — so every model your key can reach becomes available to chat, to agent mode, and to any extension that selects models through the vscode.lm API. One credential, every consumer in the window. (We covered the design in depth in an earlier post.)

Search the Extensions marketplace for tetrate, install Tetrate Agent Router Model Provider (tetrate.tetrate-model-provider), and accept the publisher trust prompt.

The VS Code marketplace page for the Tetrate Agent Router Model Provider extension by publisher tetrate, showing the install button, version, and description
The extension on the Visual Studio Marketplace. Apache-2.0 licensed; installable from the UI or with code --install-extension tetrate.tetrate-model-provider.

Configuration is two command-palette commands, using the two values from Step 3:

  1. Tetrate Agent Router: Set Base URL — the input comes pre-filled with the public Agent Router Service endpoint; replace it with the project Base URL copied from the Developer Console. This is what points the editor at your gateway, your model grants, and your budgets.
  2. Tetrate Agent Router: Set API Key — paste the inference key. It goes into VS Code secret storage backed by the OS keychain: never written to settings.json, never logged, never carried along by Settings Sync.
The extension's sidebar view prompting to set an API key, with an input box showing an sk- placeholder and a caption naming the configured project endpoint
The key is entered once and lands in OS-keychain-backed secret storage — no plaintext in settings files.

A third command, Tetrate Agent Router: Test Connection, closes the loop — in the recording it comes back with “Agent Router answered via claude-haiku-4-5 in 1.1s. The endpoint and key work end to end.” No sample project, no first mysterious failure; you know the plumbing works before you ask it to do anything.

Now the after picture. The chat model picker grows an Other Models section listing everything the gateway grants this project — Claude, Gemini, and GPT families side by side, each row carrying its provider and its input/output price per million tokens, right in the dropdown where the choice is made.

The VS Code chat model picker with the Other Models section expanded, listing Claude, Gemini, and GPT models from Agent Router, each with input and output pricing per million tokens
After: the project's model catalog in the picker, with per-million-token pricing attached to every row.
If the picker looks empty

VS Code hides extension-contributed models by default. If models are missing after setup, run Chat: Manage Language Models, choose Tetrate Agent Router, and enable the ones you want — the single most common setup issue, and a VS Code default rather than a bug.

The payoff: an enabled — and governed — developer

Pick Claude Opus 4.7 from the dropdown, ask for something (fine: a cat joke), and watch what happens around the response. The extension’s sidebar logs the request with its token counts and computed cost — 28,426 tokens in, $0.143 — and a running dollar total appears in the status bar. The cost of an AI interaction stops being an abstraction that arrives on an invoice in five weeks; it is right there, next to the code, per request.

VS Code chat showing a completed response from Claude Opus 4.7, with the extension sidebar logging the request's token counts and $0.143 cost, and the status bar showing a running session total
First routed request: the response in chat, the cost in the sidebar and status bar.

The sidebar’s Overview keeps the whole arrangement inspectable from inside the editor: the endpoint being used, gateway health, provider status, and usage totals for the session, today, and the last seven days.

The Tetrate Agent Router sidebar in VS Code with highlights around the endpoint section showing the base URL and serving gateway, the provider health list, and the usage section with session, today, and seven-day spend, next to a model picker listing Agent Router models
Everything a developer needs to self-serve: endpoint and gateway status, provider health, live usage, and the full model catalog.

And when a request deserves a closer look, the extension’s usage view renders a full dashboard in an editor tab — daily spend, a per-model breakdown with token counts, and routing outcomes — the developer-side reflection of the same usage data the Admin Console aggregates by team and project.

The Agent Router usage dashboard inside VS Code showing a daily spend chart, a by-model table with requests, input and output tokens and cost, routing outcomes, and a recent requests list with per-request costs
Usage stats without leaving the editor: daily spend, per-model totals, and per-request costs.

Step back and look at what each side of the house got:

  • The developer went from zero to a governed catalog of frontier models in VS Code with two copy-pastes and two commands — no procurement ticket, no provider accounts, no keys in dotfiles. Model choice is a dropdown, and every choice shows its price.
  • The platform team kept every control that matters: which providers and models the project may use (enforced at request time), one gateway where every request is logged and attributed, and budgets that cap spend per team or per member. Because this developer is on Backend Engineering, a $1,000 monthly hard stop guards their spend — enforced by the gateway, not by trust — and the live usage stats in the editor mean they will see it coming long before they hit it.

That is the shape of AI enablement that survives contact with an enterprise: governance and developer experience implemented as the same system, not as opposing forces. Safe for the organization, fast for the developer, and accountable to the people paying the bill.

Frequently asked questions

How do developers get access to AI models without individual provider accounts?

A platform operator grants providers and models to a project in the Agent Router Admin Console. Developers then self-serve in the Developer Console: create a project-scoped inference key, copy the project’s Base URL, and use them in any OpenAI-compatible client or in VS Code via the Tetrate Model Provider extension. No per-provider accounts or keys are involved.

What does the platform operator do versus what the developer does?

The platform operator works in the Admin Console: grant providers and models to the project, and set team or per-member budgets. The developer works in the Developer Console and the editor: create an inference key, copy the project Base URL, and configure the Tetrate Model Provider extension. No ticket is required between those two roles once the project grants and budgets exist.

What stops a developer or an agent from overspending?

Budgets, set per team or per API key with daily, weekly, or monthly periods. A Hard stop budget rejects requests once the cap is hit until the period resets; Watch spend alerts without blocking; Degrade gracefully reroutes a single key to a cheaper fallback model. Enforcement runs on usage rollups and fails open while totals catch up, so pair budgets with rate limits where an immediate cutoff is required.

What is the difference between a Hard stop budget and a rate limit?

A Hard stop is a spend ceiling on billed usage for a period (day, week, or month). It can lag actual spend by a few minutes because enforcement runs on usage rollups, and it fails open while totals catch up. A rate limit is an immediate inline cutoff on request volume or concurrency. Use budgets for finance ceilings; use rate limits when you need an emergency brake.

Can a project be restricted to specific LLM providers?

Yes. A project with an empty provider list may use any provider enabled in the organization catalog; adding providers to the project confines it to that subset. Models are stricter still — they are allow-listed, and a request naming an ungranted model is refused at request time.

Can I use the same API key outside VS Code?

Yes. The project Base URL is OpenAI-compatible. The same inference key and Base URL work with the OpenAI SDK, curl, and any compatible client. The smoke test in this post lists models with a single authenticated request.

How do developers see what they are spending?

The VS Code extension shows a running cost in the status bar, logs every request with token counts and computed cost in its sidebar, and renders a usage dashboard with daily spend and per-model breakdowns inside the editor. The same usage data aggregates in the Admin Console by key, user, team, and project.

Why is the VS Code model picker empty after installing the extension?

VS Code hides extension-contributed models by default. Run Chat: Manage Language Models, choose Tetrate Agent Router, and enable the models you want. That is the most common setup issue, and a VS Code default rather than a bug in the extension.

Now Available

MCP Catalog with verified first-party servers, profile-based configuration, and OpenInference observability are now generally available in Tetrate Agent Router Service. Start building production AI agents today with $5 free credit.

Sign up now

References

Product background Product background for tablets
Building AI agents

Agent Router Enterprise provides a managed AI Gateway, MCP Gateway, and AI Guardrails in your dedicated instance. Graduate agents from prototype to production with consistent model access, governed tool use, and runtime supervision — built on Envoy AI Gateway by its creators.

  • AI Gateway – Unified model catalog with automatic fallback across providers
  • MCP Gateway – Curated tool access with per-profile authentication and filtering
  • AI Guardrails – Enforce policies, prevent data loss, and supervise agent behavior
  • Learn more
    Replacing NGINX Ingress

    Tetrate Enterprise Gateway for Envoy (TEG) is the enterprise-ready replacement for NGINX Ingress Controller. Built on Envoy Gateway and the Kubernetes Gateway API, TEG delivers advanced traffic management, security, and observability without vendor lock-in.

  • 100% upstream Envoy Gateway – CVE-protected builds
  • Kubernetes Gateway API native – Modern, portable, and extensible ingress
  • Enterprise-grade support – 24/7 production support from Envoy experts
  • Learn more
    Decorative CTA background pattern background background
    Tetrate logo in the CTA section Tetrate logo in the CTA section for mobile

    Ready to enhance your
    network

    with more
    intelligence?