Jev-directed Model Routing
Choosing which model handles a task has a cost of its own. We tested Jev as a faster, cheaper way to make that decision when compared to LLMs.
We’ve been using Jev internally, developer to developer, to find productive ways to fit it into our workflows. So far, agentic use has been our most promising direction for reducing token costs. We’re also exploring how those ideas could benefit teams and enterprises. This post shares some of those experiments and where we might take smart routing in Agent Router Enterprise.
Most of our exploration with Jev has focused on agentic integration and what happens inside our context window. Which files are worth reading? Does a log entry matter? Before a tool call, Jev can assess the proposed action, while the host still enforces permissions and explicit rules.
Those decisions can happen before expensive tool calls fill the context window with file contents or log output. But moving them elsewhere adds its own work, so we measured that scenario too. The results were mixed: less source text ate up our valuable context window, but blanket file reading, including potentially irrelevant material, still finished faster overall.
That exploration led us to another decision developers make across many workflows: choosing a model. The potential saving on each routing call is small, but the same approach could apply wherever developers already use an LLM to choose the next model.
Model routing asks a narrower question before the next request: which model or action should receive this task? If a general-purpose LLM currently makes that decision, Jev can replace that LLM call. The benefit tested here is cheaper decision-making and less waiting before dispatch. We haven’t established that sending the work to cheaper downstream models preserves quality or reduces the whole agent’s bill. We’re exploring those questions too, so stay tuned for future results.
But we have classifiers at home. Why do we need Jev?
Traditional ML classifiers work well when you know the task and can train for its labels. Jev lets you give one model the current state and ask a new, specific question(s) in plain language, then get a typed answer with probabilities that your code can use.
That makes it practical to add many small judgments to a workflow without building a separate model and dataset for each one. And the best part is that you don’t need a whole ML team to do it. The idea of zero-shot classification isn’t new… Jev’s value is making this style of decision a reusable part of software, and that’s what we’re doing with it here.
Spend less on the routing decision
An agent session’s token bill depends on how much context it carries, which models do the work, caching, retries, and how many steps the task takes. These results don’t establish an immediate saving on an individual session. We measured one part of that bill… the model call used to choose where the next request goes.
The dollar amounts below use are using each providers published token prices - they aren’t prices for Agent Router Enterprise.
At 1,000 routing decisions, our benchmark projected roughly 3 cents for Jev versus $1.61 for GPT-6 Sol using uncached token rates, a difference of about $1.58. Across 1,000 developers making that many decisions each month, the annual volume reaches 12 million decisions: about $322 for Jev versus $19,281 for Sol, or 98.3% lower routing token costs and an annual difference of about $18,959. This illustrative workload assumes one routing call per decision; actual savings depend on usage. The comparison below lets you explore personal, business, and enterprise volumes.
Monthly routing cost by traffic
Routing tokens only. Projected USD / month at standard uncached prices; execution excluded. Two decisions are a 2× arithmetic projection.
1,000 routed requests / month
Scroll sideways for all costs and multiples.
| Decision model | 1 routing decision / request | 2 routing decisions / request | Extra vs Jev | Multiple |
|---|---|---|---|---|
| Jev 1.13 | $0.03 | $0.05 | Baseline | 1.00× |
| GPT-6 Luna | $0.09 | $0.18 | $0.07 | 3.44× |
| MiniMax M3 | $0.14 | $0.29 | $0.12 | 5.34× |
| DeepSeek V4 Flash | $0.20 | $0.39 | $0.17 | 7.33× |
| GPT-6 Sol | $1.61 | $3.21 | $1.58 | 59.85× |
| Grok 4.7 | $6.25 | $12.49 | $6.22 | 232.69× |
| Grok 4.6 | $6.38 | $12.76 | $6.35 | 237.58× |
| GPT-6 Astra | $6.75 | $13.50 | $6.72 | 251.46× |
100,000 routed requests / month
Scroll sideways for all costs and multiples.
| Decision model | 1 routing decision / request | 2 routing decisions / request | Extra vs Jev | Multiple |
|---|---|---|---|---|
| Jev 1.13 | $2.68 | $5.37 | Baseline | 1.00× |
| GPT-6 Luna | $9.24 | $18.48 | $6.56 | 3.44× |
| MiniMax M3 | $14.34 | $28.68 | $11.65 | 5.34× |
| DeepSeek V4 Flash | $19.68 | $39.35 | $16.99 | 7.33× |
| GPT-6 Sol | $160.68 | $321.35 | $157.99 | 59.85× |
| Grok 4.7 | $624.74 | $1,249.47 | $622.05 | 232.69× |
| Grok 4.6 | $637.86 | $1,275.71 | $635.17 | 237.58× |
| GPT-6 Astra | $675.13 | $1,350.25 | $672.44 | 251.46× |
10,000,000 routed requests / month
Scroll sideways for all costs and multiples.
| Decision model | 1 routing decision / request | 2 routing decisions / request | Extra vs Jev | Multiple |
|---|---|---|---|---|
| Jev 1.13 | $268.49 | $536.97 | Baseline | 1.00× |
| GPT-6 Luna | $924.00 | $1,848.00 | $655.51 | 3.44× |
| MiniMax M3 | $1,433.85 | $2,867.70 | $1,165.37 | 5.34× |
| DeepSeek V4 Flash | $1,967.51 | $3,935.03 | $1,699.03 | 7.33× |
| GPT-6 Sol | $16,067.50 | $32,135.00 | $15,799.02 | 59.85× |
| Grok 4.7 | $62,473.50 | $124,947.00 | $62,205.01 | 232.69× |
| Grok 4.6 | $63,785.50 | $127,571.00 | $63,517.02 | 237.58× |
| GPT-6 Astra | $67,512.50 | $135,025.00 | $67,244.02 | 251.46× |
Pricing sources, Sep 29: TypeSafe, OpenAI, MiniMax, DeepSeek, xAI, OpenCode Go. DeepSeek off-peak; Grok 4.7 uses OpenCode's published tariff.
We provided each router with the same eight example task states and four predefined execution routes. For Jev, these were supplied as structured state and a typed Choice question, with criteria describing when each route applies. We repeated each task five times, recorded token usage, and applied published prices with no caching discount. The comparison above scales the average cost of valid responses to the selected monthly workload. This comparison covers only the routing decision. The work that follows has its own cost.
How much delay does routing add?
The six-router four-route test used eight authored task states, each repeated five times, for 40 attempts per router. A state included the latest message, recent steps, a prior plan, and verifier limits. Jev returned a typed Choice; the five LLM routers used a 512-token output cap and were asked to return parseable routes. The chart also retains four older Go models measured with a 96-token cap at a different time and with a different request format. These are separate cohorts, not one ten-model capture. No agent executed the requested work.
Routing latency across 10 decision models
Four-route benchmark. Six routers, 512-token LLM cap; four older Go models, 96-token cap. Sorted by median; lower is faster. Shared linear scale, 0 to 20 seconds.
Exact p50 and p90 values are shown below in milliseconds.
Latency (seconds)
Filled dots mark p50; hollow dots mark p90. Each vertical line joins two percentiles for one model, not a confidence interval. Dot sizes are not durations. Jev is orange; other models are neutral.
| Decision model / cohort | p50, ms | p90, ms |
|---|---|---|
| Jev 1.13 Native Choice / 40/40 valid | 121.72 | 165.40 |
| DeepSeek V4 Flash Older Go 96 / 40/40 valid | 897.57 | 2,769.96 |
| MiniMax M3 Older Go 96 / 40/40 valid | 1,006.87 | 2,851.52 |
| Sonnet 5.5 512-token / 39/40 valid / 1 failure | 1,067.61 | 2,199.60 |
| Opus 5.5 512-token / 40/40 valid | 1,527.74 | 2,978.77 |
| GPT-6 Astra 512-token / 40/40 valid | 2,016.33 | 3,279.06 |
| GPT-6 Sol 512-token / 40/40 valid | 2,470.91 | 4,067.89 |
| GPT-6 Luna 512-token / 40/40 valid | 2,489.35 | 4,133.67 |
| Grok 4.7 Older Go 96 / 40/40 valid | 5,567.82 | 10,396.45 |
| Grok 4.6 Older Go 96 / 40/40 valid | 9,416.20 | 18,555.36 |
Nearest-rank p50 and p90 measure end-to-end request time, including network and gateway overhead, not isolated inference or task completion.
Before testing, we wrote down which route we expected for each example and gave the models rules for choosing between the four routes - those rules formed our test routing policy. In the six-router cohort, Jev, Astra and Sol matched 40 of 40; Luna matched 39 of 40; Opus matched 36 of 40. Sonnet matched 35 of its 39 valid decisions, with one additional response truncated at the 512-token cap. Its latency percentiles use those 39 valid responses, not all 40 attempts. In the separate older Go cohort, both Grok models matched 40 of 40; DeepSeek and MiniMax matched 35 of 40. This measures agreement with our authored policy, not task success.
To determine if that decision was effective, we would need to run the work and check its results to know whether those choices led to successful tasks… that part of our experiment is already underway and you’ll be able to read about it soon.
A separate two-label comparison
We then tested eight different requests three times each, choosing routine or reasoning. Extraction and formatting were routine; debugging, concurrency, architecture, and proofs needed reasoning. Each router received the same request and policy.
| Router | Median | p95 | Jev tokens | LLM tokens |
|---|---|---|---|---|
| Jev | 139 ms | 176 ms | 10,677 | 0 |
| GPT-6 Luna | 1,711 ms | 2,725 ms | 0 | 10,944 |
| GPT-6 Sol | 2,744 ms | 3,753 ms | 0 | 10,849 |
Jev made the decision 12.3 times faster than Luna and 19.8 times faster than Sol, comparing medians. That’s about 1.6 to 2.6 seconds less waiting before dispatch, not a faster completed task. All three returned 24 usable decisions matching our expected routes.
| Router | Median | p95 | Jev tokens | LLM tokens |
|---|---|---|---|---|
| Jev | 131 ms | 183 ms | 10,677 | 0 |
| Sonnet 5.5 | 1,547 ms | 2,598 ms | 0 | 11,039 |
| Opus 5.5 | 1,475 ms | 3,040 ms | 0 | 9,636 |
Here Jev was 11.8 times faster than Sonnet and 11.2 times faster than Opus. The first 128-token Claude run left six empty Sonnet responses and one empty Opus response; the 512-token run completed every decision.
Claude counts hidden thinking toward its output limit. In our 128-token run, nine Claude responses exhausted that budget before producing a complete label, consistent with thinking leaving too little room for the answer. This is one disadvantage of LLMs in routing.
Jev returned all 24 typed decisions without this failure. The separate 512-token Claude run also completed every decision and supplies the figures above. Failed attempts and their token usage are important to retain to demonstrate this.
So far, it’s working
When using LLMs to direct routing (in an effort to save time/cost), choosing a model is work you pay for before the actual task begins. In these examples, Jev followed our routing rules with less waiting than the general-purpose models. The earlier cost comparison also shows how small savings on that decision can add up across a team.
That gives us a concrete place to use Jev: make the routing decision cheaply, then spend the larger model budget on the work that needs it. Each LLM routing decision adds cost and delay. Across repeated requests, those small overheads add up, creating an opportunity for Jev to reduce both.
We still need to test whether the selected models deliver good results on real workloads. Those results could change our conclusion. For now, the initial evidence supports using Jev to reduce the cost and delay of choosing where the work goes. We’ll repeat these measurements as Jev develops to see whether those benefits improve.