Skip to content

Envoy AI Gateway becomes Agent Router and joins the Agentic AI Foundation

Learn more

How to Switch LLM Models Safely and Check the Results

Last updated: October 2026

Tetrate Agent Router Enterprise moves traffic to a new or smaller model without a code change, and shows which model answered, before the whole company is on a model nobody checked.

To switch LLM models safely, keep your apps pointed at one stable model name. Send a small share of real requests to the new model, compare the two groups, and raise the share in steps. An AI gateway does the routing and records which model answered each request. Tetrate Agent Router Enterprise offers four ways to switch models, and every response names the model that served the request.

A new model ships every few weeks. Every app that hardcodes a model name needs a code change and a deploy to try the new one. Afterward, nobody can say whether the switch helped.

Four Ways Agent Router Enterprise Switches Models Without Changing Your App

The four controls go from planned change to automatic reaction, left to right.

Model aliasTraffic splitAvailability fallbackDegrade gracefully
What starts the switchAn admin repoints the aliasA weight you set per API keyA failure from the first modelA budget running out
Who sets the controlPlatform operatorDeveloper or operatorDeveloper or operatorOperator
Which requests moveEvery request for that nameA share you choose, such as 10%Only failed requestsEvery request on that key, after the budget is spent
DocsSmart routingTraffic splittingFallbackEnforce budget caps

The routing overview explains how the controls stack. You’ll use the first two for planned changes and the last two as safety nets.

Model Aliases Keep One Name Steady While the Model Behind the Name Changes

The smart routing guide puts the idea in one line: “A model alias is a name that outlives the models behind it.” Your apps ask for a name such as chat-default. An operator decides which model answers that name. When a better model ships, the operator repoints the alias, and no app changes.

For one API key, a developer can do the same thing with a logical model name. The key keeps asking for one name while the provider model behind the name changes.

Two tips. Use an alias change when you’re confident in the new model, because every request for that name moves at once. When you want to compare two models first, start with a traffic split.

To decide which models each team can reach before you switch any of them, see AI model access control.

Traffic Splitting Turns a Model Change Into a Canary Rollout

A traffic split sends a share of a key’s requests to each model you list. The weights must add up to 100%. The first model in the list carries the Primary badge, and its name is the one your calling code asks for.

The traffic splitting guide gives three patterns:

GoalWeightsWhen to use the pattern
Canary a new model5%, then 20%, then 50%, then 100%You want proof before everyone moves
A/B test two models50/50 or 60/40You want two groups of the same size to compare
Cut cost80/20 toward the cheaper modelYou already trust the cheaper model for most work

Send 100% of the weight to one model and the split becomes a model substitution. Every request on that key goes to the new model, and the calling code keeps asking for the old name. David Wang’s post on keeping marketers away from Fable uses a 100% split to put a team on Sonnet without telling anyone to change a setting.

Two tips. Put routine work on its own API key, so the split applies to the requests that suit the cheaper model. And add a budget to the same key when you want a ceiling on total spend as well as a lower cost per request.

Every Response Names the Model That Answered

You need to know which model answered before you can judge the answer. Agent Router Enterprise records the model in three places:

WhereWhat you seeDocs
The model field in each responseThe model that served the requestSmart routing
Request Logs in the consoleThe resolved model per request, with tokens, cost, latency, and status. The detail panel shows the requested model next to the served model.Monitor traffic and usage
OpenTelemetry metricsgen_ai.request.model and gen_ai.response.model on every requestOpenTelemetry reference

No response header announces a switch. If your app depends on one model’s behavior, log the model field.

How to Check That a Smaller Model Gives Comparable Results

The gateway gives you two comparable groups of real requests. Judging whether the answers are good enough is a human job, and the traffic splitting guide says so plainly. The steps below make that judgment fast.

  1. Try the task by hand first. Open the Playground in the Developer Console. Run the same prompt on the current model, switch the model selector, and run the prompt again. Each answer shows token counts and latency.
  2. Start a canary. Set a 5% split to the smaller model on one API key.
  3. Collect enough requests. The guide says a 50-request sample is too small. Wait for several hundred.
  4. Compare the two groups in Request Logs. Group or sort by resolved model.
  5. Read a sample of answers from each group. Then raise the weight, or roll back.

Use a reading table to turn what you see into a decision:

What the logs showWhat it meansWhat to do next
Lower cost per request, same error rate, answers your reviewers acceptThe smaller model fits this taskRaise the weight to 20%
Lower cost, but reviewers reject more answersThe task needs the larger modelRoll back, or try the smaller model on a narrower task
Higher latency on the smaller modelThe provider is busy, or the model is slower on long promptsCheck by time of day before you decide
More errors on the new modelThe new model or provider needs more timeRoll back, add a fallback, and try again later

Two tips. Let the gateway compare cost, tokens, latency, and errors, and have your reviewers or your evaluation tool judge output quality on the same requests. And run the canary on a key that serves one kind of task, so each result describes one workload.

Fallback and Degrade Gracefully Switch Models When Something Goes Wrong

Availability fallback moves a request to the next model in an ordered list when the first model times out, rate-limits, or returns a server error. The switch happens before the first byte of the response, so each answer comes from one model from start to finish. The failover guide lists which errors trigger a switch.

Degrade gracefully is the more sophisticated version of model substitution. When a key’s budget runs out, the gateway reroutes that key’s requests to a cheaper model you picked, and the workload keeps running. David Wang’s token brokering post explains the idea as a circuit breaker for spend.

Two tips. Set Degrade gracefully on an API key’s budget, and pick a fallback model served over an OpenAI-compatible API. And pair the action with a Hard stop budget at a higher limit when you want a firm ceiling as well.

Claude Code Users Switch Model Families Through One Profile

Developers in Claude Code can try OpenAI, Gemini, Claude on Vertex AI, and self-hosted models through one gateway profile, with no flag changes per model. The Claude Code guide covers the setup.

Two tips. Start a new Claude Code session when you change model family. And to reach any enabled model directly, start Claude Code with --model and the model name.

How to Get Started

Model aliases, traffic splitting, availability fallback, and Degrade gracefully all ship in Agent Router Enterprise today, along with Request Logs and the Playground. Start with the routing overview, then set up your first split with the traffic splitting guide. To see a canary on a live gateway, request a demo.

Agent Router Enterprise

Tetrate Agent Router Enterprise routes AI agent traffic across providers and your own models, with policy, cost controls, and audit on every request — in cloud, on-prem, or edge.

Learn more

Frequently asked questions

How do I switch LLM models without changing my app? Point the app at a stable model name on the gateway. An operator repoints that name to the new model, or a traffic split sends a share of requests to the new model. The app keeps sending the same name.

What is a model alias? A model alias is a name your apps ask for, such as chat-default, that an operator maps to a real model. The operator can change the model behind the alias at any time without an app change.

How do I canary a new LLM model? Set a traffic split that sends 5% of a key’s requests to the new model. Compare several hundred requests from each model in the request logs, then raise the share to 20%, 50%, and 100%.

How do I know which model answered a request? Read the model field in the response, or check the resolved model in the gateway’s request logs. The logs also show the model the app asked for, so you can see when a switch happened.

How do I check that a smaller model is good enough? Run the same prompts on both models, then send a small share of real traffic to the smaller model. Compare cost, latency, and errors in the logs, and have a person review a sample of answers from each model.

How do I judge output quality when I compare models? The gateway gives you two comparable groups of real requests, with cost, tokens, latency, and errors for each model. Your reviewers or your evaluation tool then judge output quality on those same requests.


Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.

Decorative CTA background pattern background background
Tetrate logo in the CTA section Tetrate logo in the CTA section for mobile

Ready to enhance your
network

with more
intelligence?