Enabling agentic developers with Tetrate Agent Router
An end-to-end walkthrough of enabling developers with AI coding assistants — granting providers to a project, putting budgets around team spend, and getting a developer from zero to a governed model catalog inside VS Code.
Tetrate Agent Router lets a platform team grant LLM providers and set team budgets once; developers then self-serve an inference key and get a governed model catalog inside VS Code — with live spend next to the code.
Every engineering organization is having the same conversation right now. Developers want AI coding assistants — in the editor, in the terminal, in their agents — and they want model choice, because the model that plans a refactor well is rarely the cheapest one to draft commit messages with. Meanwhile the people accountable for spend and security want answers to unglamorous questions: which providers are approved, who is spending what, and what happens when an agent left running overnight burns through a month of tokens before breakfast.
Most organizations answer those questions with process: a procurement ticket per provider, an API key handed over in a password manager, a spreadsheet nobody updates. That does not scale to a world where every developer runs one or more agents, and it puts governance in the critical path of enablement — exactly where it breeds shadow AI.
Agent Router inverts that. Providers, model policy, budgets, and keys live in one gateway; developers serve themselves within the limits the organization has set. This post walks the whole path end to end, wearing two hats along the way: first as a platform operator in the Admin Console granting providers and setting team budgets, then as a developer who creates a key, wires up Visual Studio Code, and starts working.
The running example: a project called CRM-implementation, staffed by two existing teams, Backend Engineering and UI Engineering.
Step 1: Grant LLM providers to the project
Agent Router Enterprise scopes access in two layers. The organization catalog determines which providers are connected and which models are enabled at all. A project grant determines which of those a specific project may use. A request only succeeds when both layers agree, which is what makes a project a real isolation boundary rather than a folder.
In the Admin Console, everything for our example hangs off the project. Under Directory → Projects, the CRM-implementation project already exists, with six users and an attached gateway — the endpoint that developers, agents, and tools will send LLM and MCP requests to.
Opening the project and switching to the Providers tab shows what the project can reach today: Google Gemini and OpenAI, both enabled. A subtle but important default is worth calling out — a project with an empty provider list is unrestricted and may use any provider enabled in the organization catalog. Adding providers to the list is how you confine a project, for contractual or data-residency reasons, to a deliberate subset.
The CRM team has asked for Anthropic models for agentic coding and xAI for evaluation work. Selecting + Add to project opens a picker: search or browse the catalog, tick providers, and they collect in a Shortlist panel so a larger change stays reviewable before it lands.
Confirming with Add lands both grants; the table refreshes to four enabled providers and a toast confirms “Added 2 providers”.
Models follow the same pattern from the project’s Models page — with one difference in behavior: models are strictly allow-listed, so a request naming an ungranted model is refused at request time. That refusal is the enforcement mechanism, not a UI convention. The full reference is in the add providers and models to projects guide.
Step 2: Put budgets around team spend
Enablement without spend control is how AI programs get frozen by finance in their first quarter. Agent Router Enterprise treats budgets as first-class policy: under Usage → Budgets, a budget caps spend per team or per API key, on a daily, weekly, or monthly period.
Selecting + Add budget opens the New budget drawer. Three scopes are on offer, and the distinction matters more than it first appears:
- Whole team — one shared limit pooled across members. No individual caps; the team self-regulates inside one number.
- Each teammate — every member gets the same limit individually. One member running hot stops only their own keys, not the whole team.
- One API key — a limit on a single credential, which is how you cap a specific workload or an individual user.
For UI Engineering we pool the budget: scope Whole team, team UI Engineering, and a $1,000 monthly limit set from the suggested chips. The Policy preview panel on the right is a quietly excellent piece of design — it restates the policy in plain terms (“$1,000.00 / month for the UI Engineering team”) and charts the team’s actual billed spend for the last six months against a dashed limit line, so you sanity-check the number against reality before creating anything.
Next, the enforcement decision, under When the limit is hit:
- Watch spend (the default) alerts when spend passes the limit but never blocks a request — a meter, useful for trialing a number before enforcing it.
- Hard stop rejects requests once the cap is hit, until the period resets. Nothing needs to be manually re-enabled afterward.
- Degrade gracefully reroutes to a cheaper fallback model, and is available only for the one-API-key scope — the UI grays it out for team budgets and says why.
UI Engineering gets a Hard stop: the team asked for a firm ceiling rather than an alert someone might read on Monday. With a name in place — the drawer suggests one if naming things is not your favorite job — the footer switches to “Ready. Create makes exactly this,” which is precisely the sort of no-surprises confirmation you want from a policy tool.
Backend Engineering gets the other model: scope Each teammate, so every member carries their own $1,000 monthly limit with a Hard stop. The preview switches accordingly — instead of pooled history it charts the highest member’s spend, which is the number that actually threatens a per-person limit.
A few minutes of this and the Budgets page tells the whole story at a glance — pooled budget for UI Engineering, per-member budgets for Backend Engineering (plus a Watch-spend meter we added for the QA team while we were in there), all healthy, with near-limit and over-limit counters at the top for the day that changes.
Budget enforcement runs on usage rollups, so it can lag actual spend by a few minutes, and it fails open: requests are served while totals catch up. That is the right trade-off for a spend ceiling, and the wrong tool for an emergency brake — for an immediate inline cutoff, use a rate limit instead. Details in the budgets guide.
Step 3: The developer creates their own API key
Hat swap. Everything up to here was the platform operator in the Admin Console. From this point on, we are a developer on the CRM project — and notably, we will not file a single ticket.
Developers get their own home: the Developer Console. It carries a dashboard, a playground, request logs, model and MCP discovery, and — what we came for — Settings → API Keys. One detail to get right before anything else: the project switcher at the top of the sidebar. API keys and endpoints are project-scoped, so switch to CRM-implementation first. The reward is at the top of the API Keys page: the project’s Base URL, ready to copy, with a note that it works with the OpenAI SDK or any compatible client.
Selecting + Add API Key opens the Create new API key dialog. The key type defaults to Inference Key — the kind that authenticates LLM traffic — and all it needs is a name it can be recognized by later in dashboards and logs. We go with roy-inference-key and select Create key.
The confirmation screen shows the key exactly once — “Keep this key secure. You won’t be able to see it again” — together with the Base URL, so both values a client needs are in one place with copy buttons next to them. Copy both now.
Back on the list, the new key appears with a masked suffix, its creation date, and a last-used column — which is also where an operator would later spot abandoned credentials worth revoking.
Because the endpoint is OpenAI-compatible, the key is verifiable from any terminal in seconds:
export TETRATE_API_KEY="sk-..." # the key you just copied
curl -s https://proxy.tare-dummy-url.tetrate.ai/v1/models \
-H "Authorization: Bearer $TETRATE_API_KEY" The response lists every model the project grants from Step 1 allow through — which is the two-layer model doing its job on a live credential.
Step 4: Wire up Visual Studio Code
Here is the before picture. A fresh VS Code chat panel offers “Auto” and a handful of built-in models locked behind Upgrade links. Other extensions that want a model each ask for their own key. This is the credential sprawl the gateway exists to end.
The Tetrate Agent Router Model Provider extension fixes this at the right layer. Rather than pointing just the chat view at a custom endpoint, it registers as a VS Code language model chat provider — so every model your key can reach becomes available to chat, to agent mode, and to any extension that selects models through the vscode.lm API. One credential, every consumer in the window. (We covered the design in depth in an earlier post.)
Search the Extensions marketplace for tetrate, install Tetrate Agent Router Model Provider (tetrate.tetrate-model-provider), and accept the publisher trust prompt.
Configuration is two command-palette commands, using the two values from Step 3:
- Tetrate Agent Router: Set Base URL — the input comes pre-filled with the public Agent Router Service endpoint; replace it with the project Base URL copied from the Developer Console. This is what points the editor at your gateway, your model grants, and your budgets.
- Tetrate Agent Router: Set API Key — paste the inference key. It goes into VS Code secret storage backed by the OS keychain: never written to
settings.json, never logged, never carried along by Settings Sync.
A third command, Tetrate Agent Router: Test Connection, closes the loop — in the recording it comes back with “Agent Router answered via claude-haiku-4-5 in 1.1s. The endpoint and key work end to end.” No sample project, no first mysterious failure; you know the plumbing works before you ask it to do anything.
Now the after picture. The chat model picker grows an Other Models section listing everything the gateway grants this project — Claude, Gemini, and GPT families side by side, each row carrying its provider and its input/output price per million tokens, right in the dropdown where the choice is made.
VS Code hides extension-contributed models by default. If models are missing after setup, run Chat: Manage Language Models, choose Tetrate Agent Router, and enable the ones you want — the single most common setup issue, and a VS Code default rather than a bug.
The payoff: an enabled — and governed — developer
Pick Claude Opus 4.7 from the dropdown, ask for something (fine: a cat joke), and watch what happens around the response. The extension’s sidebar logs the request with its token counts and computed cost — 28,426 tokens in, $0.143 — and a running dollar total appears in the status bar. The cost of an AI interaction stops being an abstraction that arrives on an invoice in five weeks; it is right there, next to the code, per request.
The sidebar’s Overview keeps the whole arrangement inspectable from inside the editor: the endpoint being used, gateway health, provider status, and usage totals for the session, today, and the last seven days.
And when a request deserves a closer look, the extension’s usage view renders a full dashboard in an editor tab — daily spend, a per-model breakdown with token counts, and routing outcomes — the developer-side reflection of the same usage data the Admin Console aggregates by team and project.
Step back and look at what each side of the house got:
- The developer went from zero to a governed catalog of frontier models in VS Code with two copy-pastes and two commands — no procurement ticket, no provider accounts, no keys in dotfiles. Model choice is a dropdown, and every choice shows its price.
- The platform team kept every control that matters: which providers and models the project may use (enforced at request time), one gateway where every request is logged and attributed, and budgets that cap spend per team or per member. Because this developer is on Backend Engineering, a $1,000 monthly hard stop guards their spend — enforced by the gateway, not by trust — and the live usage stats in the editor mean they will see it coming long before they hit it.
That is the shape of AI enablement that survives contact with an enterprise: governance and developer experience implemented as the same system, not as opposing forces. Safe for the organization, fast for the developer, and accountable to the people paying the bill.
Frequently asked questions
How do developers get access to AI models without individual provider accounts?
A platform operator grants providers and models to a project in the Agent Router Admin Console. Developers then self-serve in the Developer Console: create a project-scoped inference key, copy the project’s Base URL, and use them in any OpenAI-compatible client or in VS Code via the Tetrate Model Provider extension. No per-provider accounts or keys are involved.
What does the platform operator do versus what the developer does?
The platform operator works in the Admin Console: grant providers and models to the project, and set team or per-member budgets. The developer works in the Developer Console and the editor: create an inference key, copy the project Base URL, and configure the Tetrate Model Provider extension. No ticket is required between those two roles once the project grants and budgets exist.
What stops a developer or an agent from overspending?
Budgets, set per team or per API key with daily, weekly, or monthly periods. A Hard stop budget rejects requests once the cap is hit until the period resets; Watch spend alerts without blocking; Degrade gracefully reroutes a single key to a cheaper fallback model. Enforcement runs on usage rollups and fails open while totals catch up, so pair budgets with rate limits where an immediate cutoff is required.
What is the difference between a Hard stop budget and a rate limit?
A Hard stop is a spend ceiling on billed usage for a period (day, week, or month). It can lag actual spend by a few minutes because enforcement runs on usage rollups, and it fails open while totals catch up. A rate limit is an immediate inline cutoff on request volume or concurrency. Use budgets for finance ceilings; use rate limits when you need an emergency brake.
Can a project be restricted to specific LLM providers?
Yes. A project with an empty provider list may use any provider enabled in the organization catalog; adding providers to the project confines it to that subset. Models are stricter still — they are allow-listed, and a request naming an ungranted model is refused at request time.
Can I use the same API key outside VS Code?
Yes. The project Base URL is OpenAI-compatible. The same inference key and Base URL work with the OpenAI SDK, curl, and any compatible client. The smoke test in this post lists models with a single authenticated request.
How do developers see what they are spending?
The VS Code extension shows a running cost in the status bar, logs every request with token counts and computed cost in its sidebar, and renders a usage dashboard with daily spend and per-model breakdowns inside the editor. The same usage data aggregates in the Admin Console by key, user, team, and project.
Why is the VS Code model picker empty after installing the extension?
VS Code hides extension-contributed models by default. Run Chat: Manage Language Models, choose Tetrate Agent Router, and enable the models you want. That is the most common setup issue, and a VS Code default rather than a bug in the extension.
Now Available
MCP Catalog with verified first-party servers, profile-based configuration, and OpenInference observability are now generally available in Tetrate Agent Router Service. Start building production AI agents today with $5 free credit.
References
- Add providers and models to projects — Agent Router Enterprise docs
- Set budgets for users and teams — Agent Router Enterprise docs
- Adding Tetrate Agent Router as a model provider in Visual Studio Code — the extension’s design in depth
- Tetrate Agent Router Model Provider — Visual Studio Marketplace