Blog · October 11, 2026 · 3 min read
Model Routing Can Cut LLM Costs 85%. The Real Evidence
A 2024 benchmark from UC Berkeley's Sky Computing Lab and LMSYS found that model routing, sending easy prompts to a cheap model and hard ones to GPT-4, cut cost by more than 85% on MT-Bench while keeping 95% of GPT-4's quality score. That result is reproducible, the method is published, and it is still the best public evidence that routing works. The 40-to-70% figures showing up in vendor blogs and case studies this year are a different thing: self-reported, no public methodology, and worth exactly as much as the vendor's incentive to publish them.
The one benchmark behind every model routing claim
RouteLLM trained four routers on Chatbot Arena preference data to decide, per request, whether a prompt needed GPT-4 or could go to the much cheaper Mixtral 8x7B instead. Calibrated to hold 95% of GPT-4's performance, the router cut GPT-4 calls enough to reduce cost by 85%+ on MT-Bench, 45% on MMLU, and 35% on GSM8K, three benchmarks with different question difficulty, which is why the savings range so widely. The caveat matters as much as the headline: those figures are calculated against a GPT-4-only baseline, and the savings depend entirely on how much of your traffic is actually easy.
Model routing's cost cut, versus GPT-4 for every request
RouteLLM benchmark results, each calibrated to keep 95% of GPT-4's quality score
MT-Bench
open-ended chat
MMLU
knowledge Q&A
GSM8K
grade-school math
That last point is the whole story. A support bot answering return-policy questions and a coding agent debugging a stack trace have completely different "easy" shares. RouteLLM's number describes Chatbot Arena's traffic mix. It does not describe yours.
The vendor numbers you'll see aren't the same claim
Search for "model routing savings" and you'll find a cluster of 2026 figures in the 40-to-70% range: a unified-API vendor citing a 71% median drop across its own customers, a consultancy citing an 83% drop that bundles routing with caching and right-sizing together, and several blog posts repeating a 73% figure from an unnamed "mid-sized SaaS company." None of them publish a method. None separate routing's effect from caching, batching, or just using a cheaper default model. One is measuring its own platform's customers, which is not a neutral sample.
Two different claims, both called "model routing savings"
What's published and reproducible, versus what's reported and unaudited
PUBLISHED BENCHMARK
35–85%
cost cut across 3 benchmarks
Method published, data public, results independently reproducible. UC Berkeley Sky Lab / LMSYS, arXiv:2406.18665 (2024).
VENDOR-REPORTED RANGE
40–70%
claimed in 2026 case studies
No public method, no independent audit, routing effects bundled with caching and other changes. Vendor and consultancy blogs, 2026.
That doesn't mean routing doesn't save money in production, it almost certainly does, for the same reason RouteLLM found savings at all: most requests are easier than the model you're sending them to. It means the specific percentage in a vendor's case study tells you about that vendor's customers, not about your bill. And unlike a cheaper sticker price, which doesn't touch your bill unless your mix of requests changes, routing is a lever that actually moves it.
What to do this week, not what to buy
Tag a week of your own traffic by task type before you tag it by model. Classification, extraction, and templated summarization are usually the easy share; open-ended reasoning and anything customer-facing with real ambiguity usually isn't. RouteLLM's savings came entirely from correctly sorting that split, a trained router just automates the sorting.
Start with a static rule, not a learned router. A simple cutoff (short, structured requests go to a cheaper model; anything else goes to your current default) captures most of the easy-task savings without needing training data or a classifier of your own. Measure your error rate on a held-out sample before trusting it with production traffic.
Find out where your spend actually sits before you change anything. You cannot route traffic you can't see. PulseMeter's per-key and per-model breakdown shows exactly which project or API key is running which model, so you can spot the extraction job quietly billing at frontier rates before you redesign anything.
Sources: Ong et al., "RouteLLM: Learning to Route LLMs with Preference Data" (arXiv:2406.18665, UC Berkeley / LMSYS, submitted June 2024, verified 2026-10-11); LMSYS, "RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing" (July 1, 2024, verified 2026-10-11); UC Berkeley Sky Computing Lab, "RouteLLM project page" (verified 2026-10-11).