Lesson 17.7 · 24 min
LLM Routing: Directing Each Query to the Best Model
Why pay a top-tier model to answer "How do I reset my password?" a million times a day?
In short: LLM routing puts a small decision layer, the router, in front of several language models and sends each query to the cheapest model that can answer it well. Routers can use rules, a trained classifier, embeddings, or a cascade that tries a small model first and escalates when its answer looks weak. Done well, routing cuts cost and latency sharply while keeping quality close to always using the biggest model.
The big picture: one size does not fit all queries
Imagine we run a customer-support chatbot for a software company. Every day it receives thousands of messages. Most are simple: "How do I reset my password?", "Where is my invoice?", "What are your opening hours?". A few are hard: "I was double-charged after downgrading my plan mid-cycle; explain why and fix it." Today we send every message to the same large, expensive model.
Language models come in many sizes. Large frontier models are strongest at multi-step reasoning, but they cost far more per token and respond more slowly. Small models are cheap and fast, and on easy questions they are often just as good. Sending everything to the large model wastes money on easy questions; sending everything to the small model fails the hard ones.
Think of it like a hospital triage desk A triage nurse looks at each patient for a minute and decides: a bandage from the nurse, a visit to the general doctor, or the specialist surgeon. The surgeon is not wasted on scraped knees, and serious cases are not left with the bandage. An LLM router is the triage desk for queries.
What LLM routing is and why we need it
LLM routing is the practice of choosing, for each incoming request, which model (or which provider, or which configuration) should handle it. The component that decides is the router. The pool of models it chooses from is often called the model pool or set of candidates. The router's goal is usually stated as: maximise answer quality subject to a cost or latency budget, or minimise cost subject to a quality floor.
- Cost. Prices per million tokens can differ by one or two orders of magnitude between small and large models. If 70% of traffic is easy, moving it to a small model cuts the bill dramatically.
- Latency. Small models usually return the first token sooner and generate faster, so easy questions feel instant.
- Quality. Routing also works upward: hard or specialised queries (code, maths, legal) can go to the model that is best at them, which can beat any single model.
- Reliability. If one provider is down or rate-limited, the router can fail over to another model.
- Specialisation. Some models are better at certain languages, at coding, or at long contexts; routing lets each do what it is good at.
Anatomy of an LLM router
Every router, simple or sophisticated, has the same parts: something that reads the query, something that scores it, a policy that turns the score into a choice, and a feedback loop that keeps the router honest.
Two numbers define a router's value: its overhead (extra milliseconds and money per request spent deciding) and its accuracy (how often it sends a query to a model that can actually handle it). A router that costs as much as the small model, or that misroutes hard queries often, can make things worse than no router at all.
Routing strategies
There are four common families of routing strategy. Many production systems combine two of them.
Two research names are useful to know. FrugalGPT (Chen, Zaharia and Zou, 2023) studied LLM cascades with a learned scorer that decides whether to accept an answer or try a stronger model. RouteLLM (from the LMSYS team, 2024) is an open-source framework for training routers between a strong and a weak model from human preference data. Both reported large cost savings at similar quality on their benchmarks; exact savings depend heavily on the traffic mix.
Pause and think: A cascade sends every query to the small model first. For a query that ends up escalated, how does its cost compare with sending it straight to the large model?
It costs more: we paid for the small model's attempt and then the large model's answer, and the user waited for both. Cascades pay off only when most queries are accepted at the first, cheap step.
A full trace example
Let us follow one message through a hybrid router for our support bot. The pool has a small model (cheap, fast) and a large model (strong, expensive). The router combines a few rules with a difficulty classifier, using a threshold of 0.35.
"Why was I double-charged after a downgrade?"
- Request arrives: The message comes in with metadata: a free-tier user, English, no attachments, two earlier turns in the conversation.
- Rules check: No rule fires: it is not a coding question, not a language that needs a special model, and the user is not on a premium tier that always gets the large model.
- Classifier scores it: A small classifier, a few milliseconds of CPU, estimates difficulty 0.72. Signals: the word "double-charged", a mention of plan changes, a request for an explanation.
- Policy decides: 0.72 > 0.35, so the request goes to the large model. A password-reset question scoring 0.08 would have gone to the small one.
- Large model answers: It calls the billing tool, reads the invoices, and explains the proration. The response is returned.
- Log for learning: We record the route, tokens, cost, latency, and later whether the user rated the answer or escalated to a human. These logs become training data for the next version of the classifier.
Now let us simulate a whole day of traffic to see how the threshold trades cost against quality.
router_sim.py
# Simulate an LLM router for a support chatbot: small vs large model.
import numpy as np
rng = np.random.default_rng(42)
N = 10_000
difficulty = rng.beta(2, 5, N) # most queries are easy (near 0)
# Router's estimate of difficulty = truth + noise (a cheap classifier)
predicted = np.clip(difficulty + rng.normal(0, 0.08, N), 0, 1)
COST = {"small": 0.0002, "large": 0.01} # $ per query (illustrative)
def p_correct(model, d): # chance the answer is good
return 1 - 1.2 * d**2 if model == "small" else 1 - 0.3 * d
def evaluate(threshold):
to_large = predicted > threshold
cost = np.where(to_large, COST["large"], COST["small"]).sum()
quality = np.where(to_large, p_correct("large", difficulty),
p_correct("small", difficulty)).mean()
return to_large.mean(), cost, quality
print("strategy %->large cost($) quality")
for name, t in [("always small", 1.01), ("router t=0.50", 0.50),
("router t=0.35", 0.35), ("router t=0.20", 0.20),
("always large", -0.01)]:
share, cost, q = evaluate(t)
print(f"{name:16s} {share:9.1%} {cost:9.2f} {q:8.3f}")
# Trace: the router's decision for three example queries (t = 0.35)
for query, d_hat in [("How do I reset my password?", 0.08),
("Summarise my last 3 invoices", 0.31),
("Why was I double-charged after a downgrade?", 0.72)]:
print(f"{d_hat:.2f} -> {'large' if d_hat > 0.35 else 'small':5s} | {query}")Output:
strategy %->large cost($) quality always small 0.0% 2.00 0.871 router t=0.50 12.5% 14.28 0.898 router t=0.35 33.5% 34.84 0.914 router t=0.20 65.1% 65.83 0.919 always large 100.0% 100.00 0.914 0.08 -> small | How do I reset my password? 0.31 -> small | Summarise my last 3 invoices 0.72 -> large | Why was I double-charged after a downgrade?
LLM routing vs Mixture of Experts
The word "router" also appears inside Mixture of Experts (MoE) models, which can cause confusion. In an MoE layer, a learned gating network sends each token to a few "expert" feed-forward sub-networks inside one model. LLM routing sends each whole request to one of several separate models, possibly from different providers.
When LLM routing is worth it
- Worth it: high traffic with a mix of easy and hard queries; a clear quality signal (ratings, tests, evals) to train and check the router; meaningful price or latency gaps between models.
- Probably not worth it: low traffic (the engineering cost exceeds the savings); every query is uniformly hard or uniformly easy; strict requirements that every user gets identical model behaviour.
- Simpler alternatives first: prompt caching, shorter prompts, or switching entirely to a cheaper model that passes your evals may save more with less complexity.
Real-world use Routing shows up in API gateways that sit in front of several providers, in assistants that pick a "fast" or "thinking" mode depending on the request, and in support and coding tools that send simple edits to small models and complex refactors to large ones. Several commercial and open-source routing products exist; their methods and claims vary, so evaluate them on your own traffic.
Common mistakes and how to fix them
Mistake: optimising cost without measuring quality A router that sends 95% of traffic to the small model looks great on the bill, until complaints arrive. Always track quality per route (offline eval sets plus online signals such as thumbs-down, retries and human escalations) and tune the threshold against both cost and quality.
- A router as expensive as the models. Using a big LLM to decide where to send the query can cost more than it saves. Fix: use rules, a small classifier or embeddings.
- Ignoring conversation context. "Yes, do that" is short, but it may continue a hard task. Fix: route on the conversation state, not just the last message, and avoid switching models mid-task without reason.
- Stale router after a model upgrade. When the small model improves, the old threshold over-routes to the large model. Fix: re-evaluate and retrain whenever the pool changes.
- No fallback. If the chosen provider times out, the request fails. Fix: add a fallback model and timeouts in the policy.
- Inconsistent tone across models. Users notice style changes. Fix: shared system prompts and output format checks across models.
Pause and think: Our router saved 60% on cost, but human escalations rose from 3% to 9%, mostly on billing questions. What should we change first?
Lower the threshold for billing-type queries (or add a rule sending billing disputes to the large model), and retrain the classifier with these misrouted examples labelled as hard. The router is misjudging a category; we should fix routing, not drop it.
Worked example, step by step
The simulation above studied a router that predicts difficulty before calling a model. A cascade decides after seeing the small model's answer, so its result depends on how good the checker is. Let us work one through by hand with 1,000 support queries, using the same illustrative prices as before: $0.0002 for the small model and $0.01 for the large one.
1,000 queries through a cascade
- The small model answers everything: Cost: 1,000 × $0.0002 = $0.20. Suppose 870 answers are good and 130 are weak, close to the 'always small' quality in the simulation.
- The checker looks at each answer: Our checker catches 80% of weak answers and wrongly flags 10% of good ones. Weak and caught: 0.8 × 130 = 104. Good but flagged: 0.1 × 870 = 87.
- Escalate the flagged ones: 104 + 87 = 191 queries go to the large model. Cost: 191 × $0.01 = $1.91.
- Total bill: $0.20 + $1.91 = $2.11, against $10.00 for sending all 1,000 to the large model. We pay about 21%.
- What slipped through: 130 − 104 = 26 weak answers were accepted and reached users. That is the quality price of an imperfect checker.
- What was wasted: The 87 false alarms cost $0.87, about 41% of the bill, to redo answers that were already fine. And those users waited for two models.
A checker therefore has two separate dials. Its catch rate protects quality. Its false-alarm rate protects cost and latency. When a cascade disappoints, fill in this 2 × 2 table from a labelled sample before changing anything else: it tells us which dial is broken. A unit test for generated code scores well on both. 'Ask the same small model if it is sure' often scores poorly on both.
Practice: try it yourself
We will simulate a cascade over 10,000 queries. Every query goes to the small model first; a checker with a chosen catch rate and false-alarm rate decides whether to escalate. We try four checkers and compare cost, quality and the average number of model calls per query.
practice_cascade.py
# Simulate a CASCADE: small model first, a checker decides whether to escalate.
import numpy as np
rng = np.random.default_rng(3)
N = 10_000
difficulty = rng.beta(2, 5, N) # most queries are easy
COST_SMALL, COST_LARGE = 0.0002, 0.01 # $ per query (illustrative)
# Did each model actually answer well? (same toy quality curves as the lesson)
small_ok = rng.random(N) < 1 - 1.2 * difficulty**2
large_ok = rng.random(N) < 1 - 0.3 * difficulty
def cascade(catch, false_alarm):
"""catch: share of bad small answers the checker flags.
false_alarm: share of good small answers it flags by mistake."""
flag_prob = np.where(small_ok, false_alarm, catch)
escalate = rng.random(N) < flag_prob
good = np.where(escalate, large_ok, small_ok).mean()
cost = N * COST_SMALL + escalate.sum() * COST_LARGE
calls = 1 + escalate.mean() # average LLM calls per query
return escalate.mean(), cost, good, calls
print(f"always small : cost $ {N * COST_SMALL:6.2f} good {small_ok.mean():.3f}")
print(f"always large : cost $ {N * COST_LARGE:6.2f} good {large_ok.mean():.3f}")
print("checker (catch, false alarm) escalated cost($) good calls/query")
for name, catch, fa in [("perfect (1.0, 0.0)", 1.0, 0.0),
("good (0.8, 0.1)", 0.8, 0.1),
("jumpy (0.8, 0.5)", 0.8, 0.5),
("lazy (0.2, 0.0)", 0.2, 0.0)]:
share, cost, good, calls = cascade(catch, fa)
print(f"{name:28s} {share:9.1%} {cost:8.2f} {good:7.3f} {calls:8.2f}")Output:
always small : cost $ 2.00 good 0.874 always large : cost $ 100.00 good 0.914 checker (catch, false alarm) escalated cost($) good calls/query perfect (1.0, 0.0) 12.6% 14.62 0.984 1.13 good (0.8, 0.1) 18.8% 20.78 0.956 1.19 jumpy (0.8, 0.5) 53.9% 55.88 0.926 1.54 lazy (0.2, 0.0) 2.9% 4.90 0.899 1.03
Now change it:
- Set
COST_SMALL = 0.005, so the small model is only half the price of the large one. Predict the cost of the 'good' and the 'jumpy' cascade. Can a cascade cost more than always using the large model? - Change the traffic to mostly hard queries with
rng.beta(5, 2, N). Predict what happens to the share escalated, and whether the cascade still saves much. - Add a checker
("paranoid (1.0, 1.0)", 1.0, 1.0)that flags everything. Predict its cost and its quality before running. Which simpler strategy is it equal to, and at what extra price?
Pause and think: The 'good' and the 'jumpy' checker both catch 80% of weak answers. Why is quality lower with the jumpy one (0.926 against 0.956), when it escalates far more queries to the stronger model?
Because most of its extra escalations are false alarms on answers that were already good. Each of those replaces a known-good answer with a fresh attempt by the large model, which is strong but not perfect, so some good answers turn into bad ones. Escalating more is not automatically safer: escalating the wrong queries costs money and can even cost quality.
Pause and think: Suppose the small model answers in 0.5 s and the large one in 2 s (illustrative), and the cascade escalates 19% of queries. What is the average wait, and who is worse off than under 'always large'?
Every query waits 0.5 s, and 19% wait another 2 s: 0.5 + 0.19 × 2 = 0.88 s on average, far below 2 s. But the escalated users wait 2.5 s, which is longer than if we had gone straight to the large model. Those are usually the users with the hardest problems. If that tail matters, a predictive router (one call, decided up front) or showing the first answer while the second is prepared is the better design.
Key takeaways
- LLM routing sends each request to the cheapest model that can answer it well, cutting cost and latency.
- A router = features → scorer → policy (threshold, rules, fallbacks) → model call → logging and feedback.
- Strategies: rules, trained classifiers, semantic/embedding similarity, and cascades that escalate weak answers.
- The threshold trades cost against quality; tune it on real traffic with quality metrics, not cost alone.
- Routing chooses between whole models per request; Mixture of Experts routes tokens inside one model.
Key terms
- LLM router: A component that chooses which language model handles each request.
- Model pool: The set of candidate models a router can choose from.
- Cascade: Trying a cheaper model first and escalating to a stronger one only if the answer fails a check.
- Semantic routing: Choosing a route by comparing the query's embedding with example queries for each route.
- Routing threshold: The score above which a request is sent to the stronger, more expensive model.
- Mixture of Experts (MoE): A model architecture where a gate sends each token to a few expert sub-networks inside the model.
← 17.6 On-Device ML: A TensorFlow Lite Android Walkthrough · 17.8 Building a Real-Time Voice AI Agent from Scratch →