AI Engineering

LLM Routing Explained: How an LLM Router Picks the Right Model (RouteLLM, Semantic Router, Jev, AI Gateways) — 2026

Most LLM traffic does not need a frontier model. An LLM router is the cheap decision in front of the expensive call: it looks at each request and sends the easy majority to a small, fast model while reserving the big one for the hard minority. This guide covers what model routing is, why the published savings are real but benchmark-specific, the four routing signals (rules, embeddings, learned routers, calibrated decision models), what RouteLLM, Semantic Router, NVIDIA's router, LiteLLM, OpenRouter's auto router, and Jev actually route on, how to build and evaluate your own router, and the mistakes that quietly turn "cheaper" into "worse."

Gurram Poorna Prudhvi

Lead AI Engineer

Technical Guide
Oct 7, 2026
13 min read
LLM ROUTING: SEND EACH QUERY TO THE CHEAPEST MODEL THAT CAN ANSWER ITa cheap decision in front of expensive models · simple → small model · hard or uncertain → frontier modelINCOMING QUERYROUTERMODEL TIER"How do I reset my password?"tokens: 9 · tools: noneuser_tier: free · history: 0 turnslabels: [simple, complex]structured fields your code assembled —not raw user text aloneLLM ROUTERone cheap decision per request — four ways to make it1 · Rules token count, user tier, tool needed2 · Embeddings nearest example utterance3 · Learned router predicted difficulty / win-rate4 · Decision model typed Choice + calibrated confidence→ label: "simple" · confidence: 0.96≥ 0.9elseSMALL / CHEAP MODELhandles the easy majoritysub-second, a fraction of the cost per callFAQ, routing, extraction, short rewritesFRONTIER MODELhandles the hard minoritymulti-step reasoning, long context, agent planningalso the DEFAULT when the router is unsureescalate if the smallanswer fails a checkROUTE ON A DECISION YOU CAN THRESHOLD — AND DEFAULT THE UNCERTAIN MIDDLE TO THE STRONGER MODELpublished results: RouteLLM >2× cost reduction · Hybrid LLM up to 40% fewer large-model calls · FrugalGPT cascades up to 98% cheaper (on their benchmarks)log every decision with its confidence and outcome · measure router accuracy and quality delta on YOUR traffic before trusting the savingsaiengineerinsights.com

TL;DR: LLM routing sends each request to the cheapest model that can answer it acceptably, using a decision made before the expensive call — by rules, embedding similarity, a trained router (RouteLLM-style), or a calibrated decision model (Jev-style). Published routers report cost cuts from roughly 2× (RouteLLM) up to 98% (FrugalGPT cascades) on their own benchmarks; the number you get depends on how much of your traffic is genuinely easy. Default anything uncertain to the stronger model, log every decision, and measure router accuracy and quality delta on your own logs.

What is LLM routing, and what does an LLM router do?

LLM routing (also called model routing) is the practice of choosing which model answers a request at runtime rather than hard-wiring your application to one model. The router is whatever makes that choice. It sits between your application and your model providers, inspects each incoming request — the text, its length, the user, the tools in play — and dispatches it to a model tier: a small, fast model for the easy cases; a frontier model for the hard ones; sometimes a middle tier in between.

The reason routing exists is an uncomfortable fact about production traffic: a large share of it is easy. Password resets, "what's your refund policy," classify-this-ticket, rewrite-this-sentence. Sending those to the same model you use for multi-step agent planning is paying frontier prices for work a model a tenth the size does fine. The FrugalGPT paper framed the economics bluntly in 2023: API fees across popular models "can differ by two orders of magnitude," which is exactly the gap a router exploits.

One distinction to fix before anything else, because the term "LLM router" is used for two different things:

  • • Model routing decides which model should answer, based on the request. This is what RouteLLM, Semantic Router, NVIDIA's router, OpenRouter's auto router, and Jev do, each with a different signal. It is the subject of this post.
  • • Gateway routing decides which deployment of a model serves a request — which region, which provider key, which copy has rate-limit headroom — plus retries, fallbacks, and cooldowns. LiteLLM's Router is the canonical open-source example. It does not look at what the request is asking.

You usually want both, layered: a model router in front choosing the tier, a gateway behind it keeping each tier reliable. Conflating them is how teams buy a gateway, call it a router, and wonder why their bill didn't move.

Does LLM routing actually save money? What the papers measured

Yes, with a caveat that matters: every published number is on the authors' own benchmark and model pair. Three results anchor the field:

  • • FrugalGPT (Chen, Zaharia, Zou, 2023) reports that it "can match the performance of the best individual LLM (e.g. GPT-4) with up to 98% cost reduction" and can "improve the accuracy over GPT-4 by 4% with the same cost," using three strategies — prompt adaptation, LLM approximation, and the LLM cascade: try a cheap model first, score its answer, and only escalate if the score is low.
  • • Hybrid LLM (Ding et al., 2024) trains a router on predicted query difficulty with a tunable quality level, and reports "up to 40% fewer calls to the large model, with no drop in response quality."
  • • RouteLLM (Ong et al., 2024) trains routers on human preference data to choose between a stronger and a weaker LLM, and reports it "significantly reduces costs — by over 2 times in certain cases — without compromising the quality of responses." Notably, the routers kept their performance "even when the strong and weak models are changed at test time."

The spread — 2× to 98% — is not the papers disagreeing. It is the savings being a function of three things you control: how much of your traffic is genuinely easy, how much cheaper your small tier is than your big one, and how accurately your router can tell the two apart. A router with perfect accuracy on traffic that is 90% easy saves close to 90% of frontier spend. The same router on agent-planning traffic that is 90% hard saves almost nothing and adds a hop. The routing research keeps pushing the accuracy term — a 2026 NVIDIA paper, "LLM Router: Rethinking Routing with Prefill Activations", routes on a model's internal prefill activations instead of surface features and reports closing 45.58% of the gap to an oracle router at 74.31% cost savings relative to the most expensive model — but the traffic-mix term is yours, and it is the one to measure first.

The four routing signals: rules, embeddings, learned routers, decision models

Every router, open source or hosted, makes its decision from one of four kinds of signal — or a stack of them. The signal determines what the router can and cannot see, which is more important than which product wraps it.

SignalWhat it decides onCost of the decisionStrengthWeakness
Rules and metadataToken count, user tier, whether a tool is required, language, SLA — plain codeFree, microseconds, deterministicZero risk of being 'wrong' in a way you can't explain; perfect for hard constraintsCannot tell an easy question from a hard one by content; rules rot as traffic shifts
Embedding similarity (semantic router)Which of your example-utterance groups the query is closest to in vector spaceOne embedding call per request (often local), tens of millisecondsNo training loop — add a route by adding example sentences; good for intent-shaped trafficMeasures topic, not difficulty; a 'billing' question can be trivial or brutal
Learned difficulty / preference routerA trained model predicts which tier will answer well (or which model 'wins') for this queryA small model forward pass; needs labeled or preference data to trainDirectly optimizes the thing you care about — quality vs cost; strongest published resultsTraining data effort; drift when your models or traffic change; opaque
Calibrated decision model (Jev-style)A typed Choice — simple / complex — plus a confidence trained to track accuracyOne non-autoregressive pass; TypeSafe quotes ~70–500 ms and $0.042 per million input tokensNo training data, labels defined at request time, and a confidence you can threshold onCalibration is imperfect in the 0.3–0.8 band; prompt injection can move the verdict

Rules are where to start and where to put anything that is actually a policy: requests over N tokens go to the long-context model, requests that need a tool go to the model that calls tools reliably, enterprise-tier users always get the frontier model. Rules can't judge difficulty from content, but they never surprise you.

Embedding similarity is what Semantic Router does: you define each route with a handful of example utterances, the library embeds them, and at request time it embeds the query and returns the closest route (or None). It is fast and needs no training loop, and it is the right tool when your traffic is intent-shaped — billing vs support vs sales. Its blind spot is that it measures topic, not difficulty. "Why was I charged twice" and "reconcile these three invoices against the contract amendment" are both billing.

Learned routers fix that by training a model to predict the thing you care about: RouteLLM predicts, from preference data, whether the weak model's answer would be preferred; Hybrid LLM predicts query difficulty against a quality target; NVIDIA's prefill-activation router predicts per-model correctness from internal activations. They have the strongest published results and the highest setup cost — you need labeled or preference data, and the router is a model that drifts when your traffic or your model pair changes.

Calibrated decision models are the newest option and the one most relevant to teams who want a learned-style signal without a training project. A model like Jev takes the request plus a label set you define at call time (simple / complex, or small / medium / frontier) and returns one typed Choice with a confidence trained to track accuracy. The confidence is the point: it lets you route on a threshold instead of a guess. We come back to the practical rules for that in the Jev section below.

A fifth pattern deserves a name because it is a different topology, not a different signal: the cascade. Instead of predicting difficulty up front, FrugalGPT's cascade calls the cheap model first, scores the answer with a separate checker, and escalates only if the score is low. Cascades don't need a difficulty predictor, but they pay the cheap call on every request and add its latency to the hard ones. Routers and cascades compose well: route the obvious cases, cascade the ambiguous ones.

Get the weekly AI engineering brief

Routing, decision models, agents, and the tools worth using — one practical email a week. Plus the free roadmap PDF.

By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.

RouteLLM vs Semantic Router vs NVIDIA vs Jev vs LiteLLM vs OpenRouter: what each actually routes on

The names below get thrown into the same "LLM router" bucket in search results and Reddit threads. They are not interchangeable. Read the third column: it tells you which signal from the previous section each one is built on, and therefore what it can't see.

ToolWhat it isRoutes onWorth knowing
RouteLLM (LMSYS / UC Berkeley)Learned router, open sourcePredicted win-rate of a strong vs weak model, trained on human preference dataPaper reports cost reductions of over 2× in certain cases without quality loss; routers transferred when the strong/weak pair was swapped
Semantic Router (Aurelio Labs)Embedding router, open sourceSimilarity between the query embedding and example utterances per routeReturns a route name or None; positioned as a 'superfast decision-making layer' that avoids waiting on an LLM generation
NVIDIA LLM Router blueprintReference architecturev2: intent routing with a small Qwen 1.7B model, or 'auto-routing' with CLIP embeddings plus a trained networkv2 only returns a model recommendation (it does not proxy the call); the repo now carries a deprecation notice pointing to NeMo Switchyard
Jev (TypeSafe AI)Calibrated decision model, APIA Choice you define at request time (e.g. simple / complex) with a calibrated confidenceUsed in LangChain's ModelRouterMiddleware with the instruction 'Choose the least costly model that can complete the task'
LiteLLM RouterAI gateway / proxy, open sourceDeployment-level load balancing: simple-shuffle, latency-based, usage-based, least-busy, cost-based, customBalances copies of the same model across providers and regions with cooldowns, fallbacks, retries; does not judge query difficulty
OpenRouter Auto Router (openrouter/auto)Hosted gateway featureA lightweight classifier assigns ~30 task types, then ranks models by what the OpenRouter community spent on that task over a trailing 7-day windowCost tiers (low → max), session stickiness, allowed/excluded model filters, graceful degradation to a default set

Two of these deserve a closer look because they are the ones most often mistaken for a difficulty router. LiteLLM's Router is a gateway: its documented strategies — simple-shuffle (the default, weighted by RPM/TPM), latency-based, usage-based, least-busy, cost-based, and custom — all pick a deployment from a pool, and its reliability features are cooldowns, fallbacks, timeouts, and retries. It is excellent at that job. It will not notice that a request is easy.

OpenRouter's Auto Router (openrouter/auto) is a genuine model router, but note what it optimizes. Per its docs, "a fast, lightweight classifier assigns each prompt one of ~30 fine-grained task types," then the router "looks up which models the OpenRouter community actually spends on over a trailing 7-day window" for that task, filtered by the cost tier you choose (low through max). That is routing on task type and market behavior, not on the difficulty of your specific request — useful if you want a sensible default per task without building anything, less useful if your goal is "send 70% of my support traffic to the small model." It also keeps a conversation on the model it landed on and degrades to a default set if classification is unavailable.

The NVIDIA blueprint is a reminder that this space moves fast: the llm-router repo went from v1 (task/complexity classifiers) to v2 (intent routing via a small Qwen model, or CLIP embeddings plus a trained network, returning recommendations only) and now carries a deprecation notice pointing at NeMo Switchyard. Treat any specific router product as replaceable; the signal taxonomy is what lasts.

How do you build an LLM router? A pattern that holds up in production

The router you ship is almost never one signal. It is rules first, then one content signal, then a threshold with a safe default, then logging. In order:

  1. Define the tiers and what "acceptable" means per tier. Two tiers is enough to start (small, frontier). Write down the quality bar — the eval score or human rubric — that a small-tier answer must clear. Without this, "the router works" is unfalsifiable.
  2. Pull every hard constraint into rules. Context length, tool requirements, tenant policy, regulated content. These never go through a model; they run first and short-circuit.
  3. Pick one content signal for the rest. Intent-shaped traffic: embeddings. Difficulty-shaped traffic with labeled data: a learned router. Difficulty-shaped traffic with no labels and a deadline: a calibrated decision model. Start with the one you can ship this week; you can swap it later because the interface — request in, tier out — doesn't change.
  4. Threshold with the expensive model as the default. Route to the small tier only when the signal is confidently "easy." Everything uncertain goes up. A router that errs toward the frontier model costs you a little money; one that errs toward the small model costs you quality you may not notice for weeks.
  5. Feed the router structured fields, not just raw text. Token count, turn count, whether retrieval found anything, user tier. It gives rules something to act on and shrinks the surface an attacker can write into.
  6. Log the decision, the confidence, the tier, and the outcome. This log is your eval set, your drift detector, and the training data for a learned router later.
  7. Put a gateway behind each tier. Retries, fallbacks, cooldowns, provider keys — the LiteLLM job. If the small tier is down, the router's decision should fall through to the frontier tier, not fail.

The shape of it as pseudocode. Illustrative only — the control flow, not real SDK calls:

# ILLUSTRATIVE PSEUDOCODE — control flow only, not real API syntax.

def route(request):
    # 1. hard rules run first, in plain code
    if request.tokens > SMALL_CTX or request.needs_tools or request.user.tier == "enterprise":
        return FRONTIER

    # 2. one content signal — swap the implementation, keep the interface
    decision = router.decide(
        question="simple or complex?",
        input=structured(request),          # fields your code assembled
        labels=["simple", "complex"],
    )
    log(request.id, decision.label, decision.confidence)

    # 3. threshold with the expensive model as the default
    if decision.label == "simple" and decision.confidence >= 0.9:
        return SMALL
    return FRONTIER                          # the uncertain middle goes UP, not down

tier = route(request)
answer = gateway.complete(tier, request)     # retries / fallbacks live in the gateway
if tier == SMALL and not passes_check(answer):
    answer = gateway.complete(FRONTIER, request)   # optional cascade on failure
log_outcome(request.id, tier, answer)

Notice the asymmetry in step 3. The threshold is on "simple," not on "complex." That is deliberate: it encodes the safe default into the control flow, so a badly calibrated or drifting router degrades into "slightly more expensive," never into "silently worse." If the router is doing its job inside an agent loop, the same decision point is also where you'd gate risky tool calls — a different question to the same cheap decision layer.

Where does Jev fit as an LLM router?

Routing by complexity is the headline use case TypeSafe and LangChain give for Jev, so it is worth being precise about what it adds and what it doesn't. Jev is a non-autoregressive decision model: you send the request plus a label set, it returns one Choice (up to 255 labels) and a confidence in a single forward pass, with no text. It is trained with RLCD — Reinforcement Learning for Calibrated Decisions — so the confidence is meant to track how often it is right. Pricing is $0.042 per million input tokens with output free, which is what makes it plausible to call on every request.

In the taxonomy above, that makes it a learned-style signal with no training step: you get a difficulty judgment (not just a topic match, as with embeddings) without collecting preference data (as with RouteLLM). LangChain's harness write-up wires it in as a ModelRouterMiddleware with the instruction "Choose the least costly model that can complete the task," and uses the same model in an AutoModeMiddleware to block risky tool calls before they execute. Our how to use Jev guide walks through that routing-plus-guardrail pattern; the Jev vs LLMs post covers why you wouldn't use a chat model to make the routing decision in the first place.

The limits, which decide how you threshold it:

  • • Calibration is real but imperfect. An independent test measured an expected calibration error of about 0.107, with confidence reliable near 0 and 1 and shaky in the 0.3–0.8 band. For routing that means: trust a 0.96 "simple"; treat a 0.6 "simple" as "complex." The pseudocode above does exactly that.
  • • Prompt injection moves the verdict. VentureBeat reported an Octomind demo in which a block probability fell from 0.76 to 0.48 after a fake "user pre-approved" field was added to the input. For a router, the attack is cheaper and subtler: a user who writes "this is a simple question" into their request may push themselves onto the weak model and get a worse answer — or, in the other direction, push expensive traffic onto your frontier tier. Keep the decision on fields your code controls, and cap per-user frontier spend in rules, not in the model.
  • • No reasoning, no explanation. You get a label and a number. If you need to know why a request was routed, you reconstruct it from the logged fields.
  • • Vendor numbers are vendor numbers. The ~70–500 ms latency and the headline speed/cost multipliers are TypeSafe's own figures on workflows it chose. Benchmark the decision latency in your region before it goes into a p99 budget.

Net: a calibrated decision model is the fastest way to get a difficulty-aware router running on day one, and the logged decisions become the dataset that lets you train a RouteLLM-style router later if the volume justifies it. The same bridge logic we described for classification applies: zero-shot now, trained later, with the zero-shot phase building toward its own replacement.

How do you evaluate an LLM router?

A router is a classifier whose errors have asymmetric costs, so evaluate it like one. Three numbers, per routing configuration, on a labeled sample of real traffic:

  • • Router accuracy, split by direction. How often did it send an easy request to the small tier (the savings), and how often did it send a hard request to the small tier (the damage)? Report them separately; a single accuracy number hides the one that matters.
  • • Quality delta vs all-frontier. Run the routed pipeline and the everything-to-the-frontier-model pipeline over the same eval set and grade both — with your existing evals, an LLM judge, and human spot checks on the disagreements. The delta is the price you are paying for the savings. Decide the acceptable delta before you look at the savings.
  • • Realized cost and latency. Not the vendor multiplier — your bill and your p50/p99, including the router's own call. A router that adds 400 ms to every request to save money on 30% of them may be a net loss on latency-sensitive paths.

Then keep running it. The production log from the build section is the re-evaluation set; re-run the three numbers whenever you change a model, a prompt, or a threshold, and watch the confidence distribution for drift. If you already have an eval harness for RAG or agents — the kind we describe in RAG evaluation metrics — the router eval slots into it as one more configuration to compare.

Common LLM routing mistakes

  • • Using a chat LLM as the router. "Rate this query's difficulty from 1 to 10" costs an autoregressive call per request, returns a number that is generated text rather than a calibrated probability, and can take longer than the small-model answer it was supposed to save. Use a cheap signal.
  • • Thresholding toward the cheap model. If "complex" needs 0.9 confidence to reach the frontier model, every uncertain request gets the weak answer. Flip it: "simple" needs the confidence; everything else goes up.
  • • Buying a gateway and calling it a router. Load balancing across deployments of the same model will not reduce frontier spend. Check which signal the product routes on.
  • • Routing on topic when the problem is difficulty. Embedding routers are fast and useful, but an intent label doesn't tell you whether the small model can handle this instance of that intent.
  • • Letting user text control the decision. Any model-based router can be nudged by what the user writes. Rules on fields you control — tier, spend caps, tool requirements — are the parts an attacker can't talk past.
  • • Shipping without the quality delta. Savings are visible on the bill the same week; quality loss shows up as churn a quarter later. Measure both on day one.
  • • Not logging confidence. A log of labels without confidences cannot tell you whether the router is drifting or whether the threshold is in the right place. Store the number.

Frequently Asked Questions

What is LLM routing?

LLM routing is a cheap decision made before an expensive model call: for each incoming request, a router picks which model (or model tier) should answer it, based on difficulty, intent, cost, latency, or policy. The goal is to send the easy majority of traffic to a small, cheap model and reserve the frontier model for the hard minority, so you cut cost and latency without a visible drop in quality.

What is the difference between an LLM router and an AI gateway?

An AI gateway (LiteLLM, Portkey, OpenRouter's core proxy) sits in front of many providers and handles keys, rate limits, retries, fallbacks, and load balancing across deployments of a model — it decides which copy of a model serves a request. An LLM router decides which model should answer at all, based on the content of the request. Many products do both; the routing signal (rules, embeddings, a learned router, or a calibrated decision model) is the part that determines quality and cost.

How much does LLM routing actually save?

Published results, on the authors' own benchmarks: RouteLLM (Ong et al., 2024) reports cost reductions of over 2× in certain cases without compromising response quality; Hybrid LLM (Ding et al., 2024) reports up to 40% fewer calls to the large model with no drop in response quality; FrugalGPT (Chen, Zaharia, Zou, 2023) reports matching the best individual LLM with up to 98% cost reduction using cascades and related tricks. Your number depends on how much of your traffic is genuinely easy, so measure it on your own logs before putting it in a budget.

What is RouteLLM?

RouteLLM is an open-source framework and paper from LMSYS and UC Berkeley (Ong et al., 2024) for training routers that dynamically choose between a stronger and a weaker LLM at inference time. The routers are trained on human preference data (plus data augmentation) to predict when the weaker model's answer would be preferred, and the paper reports that trained routers kept working even when the strong and weak models were swapped at test time.

Can I use Jev as an LLM router?

Yes — routing by complexity is one of the use cases TypeSafe and LangChain document for it. You ask Jev a Choice question (simple vs complex, or small / medium / frontier) and get back a label plus a calibrated confidence in one non-autoregressive pass, with no training data. The practical rules: route to the cheap model only on high-confidence 'simple', send every mid-range confidence (roughly 0.3–0.8) to the stronger model by default, and keep attacker-controlled text out of the fields the decision hinges on, because prompt injection has been shown to move Jev's verdicts.

How do I evaluate an LLM router?

Build a labeled eval set of real requests where you already know which tier answers acceptably, then measure three things for each routing configuration: router accuracy (how often it picks the cheapest acceptable tier), the quality delta versus sending everything to the frontier model (graded by your existing evals or an LLM judge with human spot checks), and the realized cost and latency. Log every production decision with its confidence and outcome so you can re-run that evaluation as traffic and models drift.

Want 1:1 help? Book a session

5.0 · 18 reviews

Career guidance, resume & interview prep, or tech consulting — 1:1 with a Lead AI Engineer. Start with a free 30-min quick chat.

References

Found this useful? Share it.

Share:

Related Articles