Cost reduction isn’t a given. It also depends on the types of requests the application receives, the price difference between models, and how well the routing system performs. In this article, we are going to look at various aspects
The 10x depends heavily on two things — 85% of traffic being genuinely easy, and the router being right about which 85%. Do you have a sense of how much routing accuracy degrades in production versus on the eval set?
A router is itself a model, so your request pays an extra inference before the one that answers. Savings only appear when enough traffic is genuinely easy and the price gap covers that overhead plus the cost of mistakes. Below that line routing is a net loss.
The part this nails is the genuinely hard one, judging a request's difficulty without answering it first. What tends to get skipped, and a commenter here asked exactly this, is how far a router's accuracy holds up on requests the models haven't already seen. A lot of routing evals quietly run on tasks the models memorised in training, so everything scores near the top and routing looks free. (just look at ROuterbench) On fresh, contamination-controlled tasks it's more honest: in an open benchmark I run, ideal routing hit about 93% task success, above the frontier model on its own, at roughly 60% lower cost, and every figure re-grades offline so you can check it rather than trust it.
Two things that made the cost/quality call hold up for me: escalate on the router's own confidence, not a fixed EASY/MEDIUM/HARD label, and put a receipt on every response (chosen model, confidence, alternatives, estimated saving) so a bad route shows up per request instead of surfacing later as a quality complaint.
Disclosure, I build an open, self-hosted router on exactly this (OmnisRouter, Apache-2.0, github.com/Fortitude-Group/OmnisRouter), with the benchmark (OmnisBench) open alongside it, so weight it accordingly. Both are inspectable if the "does it degrade in production" question is the one you care about.
The savings math is clean, but the second-order cost is rarely modelled: every misrouted hard request lands on the cheap model, fails, and gets re-run on the strong one — you pay for both, plus the latency.
In practice the routing accuracy rate, not the price gap, is what decides whether this is a 10X saving or a 1.5X one.
I'd stress-test the router on your worst 5% of traffic before trusting the blended average.
The real architectural challenge with cost-driven model routing is managing the classification overhead and fallback latency. If your routing classifier introduces cold starts or misroutes edge cases into expensive re-eval loops, you quickly burn through the margin you intended to save.
Defaulting every request to the biggest model is the cloud-bill version of leaving debug logging on in production. Routing by task difficulty is just triage IMO, something ops has done forever. Most prompts don't need the Ferrari. They need the right tool and a fallback when the cool looking one fails. Cost cutting that also improves reliability is rare, and worth doing.
The 10x depends heavily on two things — 85% of traffic being genuinely easy, and the router being right about which 85%. Do you have a sense of how much routing accuracy degrades in production versus on the eval set?
A router is itself a model, so your request pays an extra inference before the one that answers. Savings only appear when enough traffic is genuinely easy and the price gap covers that overhead plus the cost of mistakes. Below that line routing is a net loss.
The part this nails is the genuinely hard one, judging a request's difficulty without answering it first. What tends to get skipped, and a commenter here asked exactly this, is how far a router's accuracy holds up on requests the models haven't already seen. A lot of routing evals quietly run on tasks the models memorised in training, so everything scores near the top and routing looks free. (just look at ROuterbench) On fresh, contamination-controlled tasks it's more honest: in an open benchmark I run, ideal routing hit about 93% task success, above the frontier model on its own, at roughly 60% lower cost, and every figure re-grades offline so you can check it rather than trust it.
Two things that made the cost/quality call hold up for me: escalate on the router's own confidence, not a fixed EASY/MEDIUM/HARD label, and put a receipt on every response (chosen model, confidence, alternatives, estimated saving) so a bad route shows up per request instead of surfacing later as a quality complaint.
Disclosure, I build an open, self-hosted router on exactly this (OmnisRouter, Apache-2.0, github.com/Fortitude-Group/OmnisRouter), with the benchmark (OmnisBench) open alongside it, so weight it accordingly. Both are inspectable if the "does it degrade in production" question is the one you care about.
Happy to discuss.
Echoing others: the 10x savings assumes you actually know your easy/hard split and most teams don't until routing forces them to measure it.
The router's real job is earning trust on the worst 5% of traffic and not the average case.
The savings math is clean, but the second-order cost is rarely modelled: every misrouted hard request lands on the cheap model, fails, and gets re-run on the strong one — you pay for both, plus the latency.
In practice the routing accuracy rate, not the price gap, is what decides whether this is a 10X saving or a 1.5X one.
I'd stress-test the router on your worst 5% of traffic before trusting the blended average.
The real architectural challenge with cost-driven model routing is managing the classification overhead and fallback latency. If your routing classifier introduces cold starts or misroutes edge cases into expensive re-eval loops, you quickly burn through the margin you intended to save.
Defaulting every request to the biggest model is the cloud-bill version of leaving debug logging on in production. Routing by task difficulty is just triage IMO, something ops has done forever. Most prompts don't need the Ferrari. They need the right tool and a fallback when the cool looking one fails. Cost cutting that also improves reliability is rare, and worth doing.