How to Switch LLM Models Safely and Check the Results
Last updated: October 2026
Tetrate Agent Router Enterprise moves traffic to a new or smaller model without a code change, and shows which model answered, before the whole company is on a model nobody checked.
To switch LLM models safely, keep your apps pointed at one stable model name. Send a small share of real requests to the new model, compare the two groups, and raise the share in steps. An AI gateway does the routing and records which model answered each request. Tetrate Agent Router Enterprise offers four ways to switch models, and every response names the model that served the request.
A new model ships every few weeks. Every app that hardcodes a model name needs a code change and a deploy to try the new one. Afterward, nobody can say whether the switch helped.
Four Ways Agent Router Enterprise Switches Models Without Changing Your App
The four controls go from planned change to automatic reaction, left to right.
| Model alias | Traffic split | Availability fallback | Degrade gracefully | |
|---|---|---|---|---|
| What starts the switch | An admin repoints the alias | A weight you set per API key | A failure from the first model | A budget running out |
| Who sets the control | Platform operator | Developer or operator | Developer or operator | Operator |
| Which requests move | Every request for that name | A share you choose, such as 10% | Only failed requests | Every request on that key, after the budget is spent |
| Docs | Smart routing | Traffic splitting | Fallback | Enforce budget caps |
The routing overview explains how the controls stack. You’ll use the first two for planned changes and the last two as safety nets.
Model Aliases Keep One Name Steady While the Model Behind the Name Changes
The smart routing guide puts the idea in one line: “A model alias is a name that outlives the models behind it.” Your apps ask for a name such as chat-default. An operator decides which model answers that name. When a better model ships, the operator repoints the alias, and no app changes.
For one API key, a developer can do the same thing with a logical model name. The key keeps asking for one name while the provider model behind the name changes.
Two tips. Use an alias change when you’re confident in the new model, because every request for that name moves at once. When you want to compare two models first, start with a traffic split.
To decide which models each team can reach before you switch any of them, see AI model access control.
Traffic Splitting Turns a Model Change Into a Canary Rollout
A traffic split sends a share of a key’s requests to each model you list. The weights must add up to 100%. The first model in the list carries the Primary badge, and its name is the one your calling code asks for.
The traffic splitting guide gives three patterns:
| Goal | Weights | When to use the pattern |
|---|---|---|
| Canary a new model | 5%, then 20%, then 50%, then 100% | You want proof before everyone moves |
| A/B test two models | 50/50 or 60/40 | You want two groups of the same size to compare |
| Cut cost | 80/20 toward the cheaper model | You already trust the cheaper model for most work |
Send 100% of the weight to one model and the split becomes a model substitution. Every request on that key goes to the new model, and the calling code keeps asking for the old name. David Wang’s post on keeping marketers away from Fable uses a 100% split to put a team on Sonnet without telling anyone to change a setting.
Two tips. Put routine work on its own API key, so the split applies to the requests that suit the cheaper model. And add a budget to the same key when you want a ceiling on total spend as well as a lower cost per request.
Every Response Names the Model That Answered
You need to know which model answered before you can judge the answer. Agent Router Enterprise records the model in three places:
| Where | What you see | Docs |
|---|---|---|
The model field in each response | The model that served the request | Smart routing |
| Request Logs in the console | The resolved model per request, with tokens, cost, latency, and status. The detail panel shows the requested model next to the served model. | Monitor traffic and usage |
| OpenTelemetry metrics | gen_ai.request.model and gen_ai.response.model on every request | OpenTelemetry reference |
No response header announces a switch. If your app depends on one model’s behavior, log the model field.
How to Check That a Smaller Model Gives Comparable Results
The gateway gives you two comparable groups of real requests. Judging whether the answers are good enough is a human job, and the traffic splitting guide says so plainly. The steps below make that judgment fast.
- Try the task by hand first. Open the Playground in the Developer Console. Run the same prompt on the current model, switch the model selector, and run the prompt again. Each answer shows token counts and latency.
- Start a canary. Set a 5% split to the smaller model on one API key.
- Collect enough requests. The guide says a 50-request sample is too small. Wait for several hundred.
- Compare the two groups in Request Logs. Group or sort by resolved model.
- Read a sample of answers from each group. Then raise the weight, or roll back.
Use a reading table to turn what you see into a decision:
| What the logs show | What it means | What to do next |
|---|---|---|
| Lower cost per request, same error rate, answers your reviewers accept | The smaller model fits this task | Raise the weight to 20% |
| Lower cost, but reviewers reject more answers | The task needs the larger model | Roll back, or try the smaller model on a narrower task |
| Higher latency on the smaller model | The provider is busy, or the model is slower on long prompts | Check by time of day before you decide |
| More errors on the new model | The new model or provider needs more time | Roll back, add a fallback, and try again later |
Two tips. Let the gateway compare cost, tokens, latency, and errors, and have your reviewers or your evaluation tool judge output quality on the same requests. And run the canary on a key that serves one kind of task, so each result describes one workload.
Fallback and Degrade Gracefully Switch Models When Something Goes Wrong
Availability fallback moves a request to the next model in an ordered list when the first model times out, rate-limits, or returns a server error. The switch happens before the first byte of the response, so each answer comes from one model from start to finish. The failover guide lists which errors trigger a switch.
Degrade gracefully is the more sophisticated version of model substitution. When a key’s budget runs out, the gateway reroutes that key’s requests to a cheaper model you picked, and the workload keeps running. David Wang’s token brokering post explains the idea as a circuit breaker for spend.
Two tips. Set Degrade gracefully on an API key’s budget, and pick a fallback model served over an OpenAI-compatible API. And pair the action with a Hard stop budget at a higher limit when you want a firm ceiling as well.
Claude Code Users Switch Model Families Through One Profile
Developers in Claude Code can try OpenAI, Gemini, Claude on Vertex AI, and self-hosted models through one gateway profile, with no flag changes per model. The Claude Code guide covers the setup.
Two tips. Start a new Claude Code session when you change model family. And to reach any enabled model directly, start Claude Code with --model and the model name.
How to Get Started
Model aliases, traffic splitting, availability fallback, and Degrade gracefully all ship in Agent Router Enterprise today, along with Request Logs and the Playground. Start with the routing overview, then set up your first split with the traffic splitting guide. To see a canary on a live gateway, request a demo.
Agent Router Enterprise
Frequently asked questions
How do I switch LLM models without changing my app? Point the app at a stable model name on the gateway. An operator repoints that name to the new model, or a traffic split sends a share of requests to the new model. The app keeps sending the same name.
What is a model alias? A model alias is a name your apps ask for, such as chat-default, that an operator maps to a real model. The operator can change the model behind the alias at any time without an app change.
How do I canary a new LLM model? Set a traffic split that sends 5% of a key’s requests to the new model. Compare several hundred requests from each model in the request logs, then raise the share to 20%, 50%, and 100%.
How do I know which model answered a request? Read the model field in the response, or check the resolved model in the gateway’s request logs. The logs also show the model the app asked for, so you can see when a switch happened.
How do I check that a smaller model is good enough? Run the same prompts on both models, then send a small share of real traffic to the smaller model. Compare cost, latency, and errors in the logs, and have a person review a sample of answers from each model.
How do I judge output quality when I compare models? The gateway gives you two comparable groups of real requests, with cost, tokens, latency, and errors for each model. Your reviewers or your evaluation tool then judge output quality on those same requests.
Related reading
- Keeping Marketers Away From Fable, on model substitution with a 100% split
- Announcing token brokering, on stepping down to a cheaper model when a budget runs out
Learn more about Tetrate Agent Router Enterprise — enterprise AI agent routing with policy, cost controls, and audit across every gateway.