Your AI bill has a strange shape. The same one-sentence question can cost $0.0004 on one model and $0.007 on another, a 35x spread for output that, most of the time, nobody can tell apart. That gap is the entire business case for AI model routers: tools that look at each prompt and decide which LLM should answer it, instead of sending everything to whatever flagship you hardcoded in March.
The idea sounds obvious. The execution is where teams lose money instead of saving it. In the past year the routing market has gone through a shakeout: one pioneer pivoted away, one framework went dormant, one gateway got acquired, and one company publicly deleted its own router after four months of production use. Meanwhile the tool that survived, in several forms, is quietly cutting real bills by half.
So here is the honest version, with the actual arithmetic.
The price spread that makes routing tempting
Start with published list prices, because the routing pitch only works if the spread is real. It is. As of September 2026, OpenAI's flagship GPT-5.2 costs $1.75 per million input tokens and $14 per million output tokens. Its smallest sibling, gpt-5-nano, costs $0.05 and $0.40. Anthropic's Claude Opus 5, released July 24, 2026, lists at $5 and $25 per million tokens, while Claude Haiku 4.5 sits at $1 and $5. And at the far end, DeepSeek's flash model takes $0.15 per million uncached input tokens at off-peak rates and $0.60 per million output. You can verify every number on the vendors' own pricing pages.
Run a typical support-chatbot workload through that ladder: 10,000 requests a day, each about 1,500 input tokens and 300 output tokens.
| All requests to | Cost per request | Cost per day |
|---|---|---|
| GPT-5.2 | $0.0068 | $68.25 |
| gpt-5-mini | $0.0010 | $9.75 |
| gpt-5-nano | $0.0002 | $1.95 |
Now the routing promise: send the trivial 70% of queries to nano and keep GPT-5.2 for the hard 30%. Blended cost lands around $21.84 a day, a 68% cut, with the flagship reserved for the conversations that actually need it. That is the whole pitch, and on this workload shape the pitch is correct.
The catch is that everything depends on that 70% number. If only 20% of your traffic is genuinely trivial, the same setup saves about 19%. Routing tools love to quote the first case. Your traffic decides which case you live in, which is why you should measure before installing anything (more on that below).
Three ways to route, and what each one actually decides
Every router on the market in 2026 makes one of three kinds of decisions, and it matters a lot which one you are buying.
Static rules. You decide, in code or config, that summarization goes to the cheap model and legal analysis goes to the expensive one. No classifier, no vendor. This is what most teams actually should start with, and interestingly, it is where several routing startups landed after their fancier systems disappointed them. It fails when your task labels lie: "evaluate the tests in this repo and improve them" is trivial on a static site and brutal on the Linux kernel.
Classifier routers. A small model or heuristic scores each prompt for complexity and picks a tier. LiteLLM's Auto Router is the clearest example: it classifies every request into SIMPLE, MEDIUM, COMPLEX, or REASONING and maps each tier to a model you choose, with a sub-millisecond heuristic classifier or a small LLM doing the scoring. You keep one model name in your clients and the gateway picks the tier underneath.
Per-task evals and market signals. Instead of classifying the prompt, classify the task and trust aggregate behavior. OpenRouter's updated Auto Router (model string openrouter/auto) assigns each prompt one of about 30 fine-grained task types, such as code debugging or multi-step agent planning, then routes to whichever models the OpenRouter community actually spends the most on for that task type over the trailing week. Their launch post calls it "wisdom of the market": when developers migrate a workload to a new model, the router follows within days, with no retraining. You set a cost tier from low to max and can restrict the candidate pool with wildcards.
There is also a fourth thing often confused with routing: provider routing, where the model stays fixed and the gateway just picks the cheapest or fastest provider serving it, with fallbacks when one is down. OpenRouter does this by default across every model, LiteLLM and Cloudflare AI Gateway do it, and it is unambiguously good infrastructure. It does not have the failure modes of model switching. If someone sells you "routing", ask which of the four they mean.
The tool landscape in September 2026
The past twelve months rearranged this space, so ignore any comparison written before mid-2026. Current state, verified:
OpenRouter is the default starting point. It is a hosted aggregator: one API for hundreds of models, provider load balancing, transparent fallbacks, and the Auto Router on top. It adds no markup to per-token prices; it earns a 5.5% fee (minimum $0.80) when you buy credits, and a 5% fee on bring-your-own-key usage above a monthly allowance, per its FAQ. The Auto Router itself costs nothing extra. Model variants like :nitro (fastest providers) and :floor (cheapest) give you one-line provider routing without any classifier at all.
LiteLLM is the self-hosted option and the one with the best public evidence. In August 2026 it published numbers from a real deployment: 272,876 requests across 450+ users over nearly four months cost $11,736 through the Auto Router versus a computed $23,985 if every request had gone to the flagship tier, a 51% saving, with 95% of requests never needing the flagship at all. Read the full breakdown: the methodology is disclosed, the counterfactual is priced with the same token counts, and the authors flag their own assumptions. That is rarer than it should be. A separate Terminal-Bench run matched Claude Opus 5's task solve rate at 27% lower cost.
Not Diamond is alive and doubled down. Its routing SDK returns a model recommendation per request and you call the model with your own keys; tradeoff modes let you optimize for quality, cost, or latency. In August 2026 it launched Not Diamond Code, a router aimed specifically at long-running coding agents, and it lists on AWS Marketplace. Reported pay-as-you-go pricing is around $0.05 per million tokens routed (vendor-reported; the public pricing page does not show a number).
Martian, which marketed itself as the inventor of the commercial LLM router, has pivoted. The company now presents itself as an interpretability research lab; it still operates a 200+ model gateway, but the router-as-product is no longer the front door. If you find a 2024 or 2025 article recommending the Martian Model Router as a going concern, that article is stale.
RouteLLM, the LMSYS framework behind most of the "cut costs 85%" claims floating around, is research canon with a shelf life. Its GitHub repository has not been updated since August 2024. The paper's numbers (roughly 85% cost reduction while keeping about 95% of GPT-4 quality) were measured against the 2024 model landscape, where the strong model was GPT-4 and the weak model was dramatically worse. The framework is fine to study, but as production infrastructure in 2026 it is unmaintained.
Gateways around the edges. Portkey, a popular open-source gateway, was acquired by Palo Alto Networks (closed May 29, 2026) and folded into the Prisma AIRS security platform. Vercel AI Gateway charges no markup on tokens; Cloudflare AI Gateway offers visual dynamic routing with fallbacks and unified billing with Workers AI. These are provider-routing plays; they do not pick models for you.
The math that kills naive routing: the cache
Here is the part most routing pitches skip, and it is the reason the Manifest team deleted their own router in 2026, after launching it in March, deprecating it in June, and shutting it down September 1. They explained the whole postmortem on their blog, and their findings are worth internalizing.
Prompt caches bill a repeated prompt prefix at a small fraction of the base input price: 10% on the Claude API and 10% of standard input across OpenAI's GPT-5 family. In agentic workloads, that prefix is almost the entire cost. The median coding-agent step in a widely cited 2026 trace study read about 126,000 prefix tokens and emitted only about 250 output tokens: over 99% of input was context the agent had already sent.
Now watch a router operate on that shape. Staying on Opus 5 with a warm cache: 126,000 tokens at 10% of $5 per million is about $0.063, plus new tokens and output, roughly $0.074 for the step. Routing that step to the "cheaper" Claude Sonnet 5, whose cache is cold because you just switched models: about 127,000 tokens at the full $3 per million input rate, roughly $0.38 for the step. Sonnet is cheaper on list price. The routed step cost five times more.
The general rule falls out of the arithmetic: a cache read costs 0.1x the base input price, a cold route costs 1.0x, so switching mid-session only pays if the target model's uncached input price is more than 10x cheaper than the model you are leaving. Within one vendor's family that bar is almost never cleared. Haiku at $1 does not beat a warm Opus cache at $5. A jump to gpt-5-nano ($0.05) or DeepSeek flash ($0.15) can clear it, if you then stay there long enough to warm the new cache.
That is why OpenRouter's own Auto Router announcement contains a candid sentence about increased costs when the input cache is rebuilt, and implements session stickiness to keep multi-turn conversations on one model. Manifest put it more bluntly: a cache-aware router ends up not switching, which means it does its job by not doing it.

When routing is worth it, in one paragraph each
Route when your workload is short and stateless. Support chat, classification, extraction, single-shot Q&A: small prompts, no shared prefix, cheap models genuinely indistinguishable on the task. This is where the 50%+ savings live and where the LiteLLM production numbers came from.
Route across a large gap, then stay put. The 10x rule means useful routing pools are small and sharply differentiated: a flagship and a sub-$0.50 model, not a flagship and a mid-tier. Pin sessions, route at conversation start rather than mid-stream, and let the cache do the rest.
Route providers aggressively, models reluctantly. Fallbacks, price sorting, and throughput sorting are pure wins. Model switching is where the failure modes live: cache loss, behavior shifts mid-task, and classification errors that send a subtle request to a model that fumbles it.
Do not route quality-sensitive agent loops per step. Manifest's second finding was behavioral: engineers reported that jumping between models during a working session lowered the quality of the overall work, separate from any cost effect. For agent workloads, choose a model at session start and commit. If you are building in that world, our pieces on how to pick the right AI model and on AI agents versus AI automation cover the selection side without a router in the middle.
Measure first: replay a fixed prompt set before you buy
Any router worth installing can be evaluated before it touches production, and the method is cheap. Collect 100 to 300 real prompts from your actual traffic (the distribution matters more than the count). Replay the same set through your current model and through the two or three candidates your router would use. For every request, log four things: cost from the API usage fields, wall-clock latency, quality against whatever checkable signal you have (unit tests for code, schema validity for extraction, a rubric score for prose), and failures (refusals, empty outputs, format errors).
Two comparisons fall out. First, the ceiling: cost of the cheap model on everything versus the flagship on everything. Second, the realistic blended number: what share of your requests the cheap model answered at acceptable quality. If that share is under 40%, a router saves you little and a fixed cheap-first, escalate-on-failure rule might beat it. If it is 70% or more, routing (or even a hard-coded default to the cheap model) is one of the highest-ROI changes you can make to an LLM product.
That replay also future-proofs you: model prices moved repeatedly through 2026, and a 20-minute re-run of the same prompt set tells you whether your routing map is still right. This kind of ongoing cost review is the same discipline we recommend for the hidden costs of AI automation: the invoice only shows you what you measured.
FAQ
What is the best AI model router for cost and quality? For hosted use, OpenRouter's Auto Router: no extra fee, market-informed picks, cost tiers. For self-hosting, LiteLLM's Auto Router, which has the strongest published production savings (51% in a 2026 deployment) and runs inside your own VPC. For coding agents specifically, Not Diamond Code is the dedicated option. Start with provider routing in any case; it is pure upside.
Can I route prompts to different LLMs automatically?
Yes: send a special model string (openrouter/auto on OpenRouter, an auto-router model group in LiteLLM) and the platform classifies each prompt and picks a model. The caveats above still apply: keep sessions sticky, watch cache rebuild costs, and replay a fixed prompt set first to confirm the savings on your traffic rather than the vendor's.
Pricing and plan details are as published by the vendor around September 2026 and can change - confirm on the official site.




