Vital signs
Why it's on the table
On the table, Model Router (Rt) is seat 4 of 58, in the Intelligence family. It is an emerging element — the job is real and here to stay, but the leaderboard still changes quarterly. Choose for this quarter, hold loosely, and watch the changelog. It is optional: plenty of companies run without it — until a specific trigger (scale, regulation, cost, or customers) makes it essential for them. It sits in the lowest paid band — lunch money against the hours it returns.
Model Router: the top 5 — v2026.Q3
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
1OpenRouterOpenRouter, Inc.
Pass-through token pricing (0% markup) · 5.5% fee on credit purchases ($0.80 min; 5% via crypto) · BYOK free to 1M req/mo, then 5%Best for Any team that wants every model behind one key and one invoice today, with zero infrastructure to run.
The category's runaway aggregator: 25T tokens/week (up 5x in six months, on pace for a quadrillion/year), 8M+ developers, 400+ models across 60+ providers, and a $113M Series B at ~$1.3B (May 2026) with CapitalG, NVIDIA, ServiceNow, MongoDB, Snowflake and Databricks all on the round. Edge routing adds ~25ms; pricing is pure pass-through, so the only cost is the credit fee. Sacra pegs annualized platform revenue at ~$50M by Mar 2026 — proof the ~5% take rate model works at scale.
Watch That 5.5% compounds painfully at volume — heavy spenders graduate to direct provider contracts or BYOK (which itself costs 5% past 1M req/mo). Fully managed only: your entire model traffic transits a single startup, and critics argue frontier-lab consolidation erodes the aggregator's reason to exist.
$113M Series B led by CapitalG at ~$1.3B; 25T tokens/week, 8M devs, 400+ models (May 26–28, 2026) [src] · Valuation 2.4x in a year ($547M Jun 2025 → $1.3B May 2026); 100T tokens/month (TechCrunch, May 26, 2026) [src] · ~$50M annualized revenue by Mar 2026 on ~5% take; ~25ms added edge latency (Sacra) [src]2LiteLLMBerriAI (YC W23; MIT OSS)
Open source free forever (self-host) · Enterprise custom annual, sized to request capacity — never per-token · 30-day trialBest for Teams that need the router inside their own VPC — keys, logs, and spend controls that never leave your infrastructure.
The open-source gateway standard: 53.8k GitHub stars, 1,000+ contributors, 240M+ Docker pulls, and testimonials from Netflix, NVIDIA, Okta and Stripe. Claims 140+ providers and ~1,900 models behind one OpenAI-compatible API with virtual keys, team budgets, fallbacks and Prometheus metrics all in the free tier. The 2026 Rust-core gateway answered the performance critics: 0.66ms added at p99 (vendor benchmark, ~2,800 RPS).
Watch You run it — upgrades, scaling, and a famously sprawling codebase (2.5k open PRs) are your problem. Performance numbers are vendor-published; rivals (Bifrost) built entire marketing campaigns on LiteLLM's Python-era overhead. Enterprise pricing is quote-only and opaque.
3Vercel AI GatewayVercel
0% markup, 0% platform fee — provider list price via credits · BYOK free · free tier (subset of models) · metered add-ons (team-wide ZDR/allowlist $0.10/1k req)Best for TypeScript/AI SDK teams and anyone who refuses to pay a take rate on tokens at all.
The price disruptor: GA since Aug 21, 2025 with sub-20ms latency, hundreds of models, automatic failover — and genuinely zero markup, including on BYOK, where OpenRouter charges 5%+. Its monthly production index (a real telemetry set: open-weight models hit 29% of gateway tokens on <4% of spend, Jul 2026) shows serious traffic already flows through it. Vercel subsidizes the gateway to sell the surrounding cloud, which is exactly why it's cheap.
Watch Youngest of the leaders; observability and governance arrive as metered add-ons (trace drains, tag writes, query fees) that quietly rebuild a bill. Loss-leader economics could change when strategy does. Deepest value only lands if you're inside the Vercel/AI SDK ecosystem.
4Cloudflare AI GatewayCloudflare
Core features free (analytics, caching, rate limits, dynamic routing) · unified billing +5% on credit purchases · Workers Paid ($5/mo) lifts log caps to 10M/gatewayBest for Teams already on Cloudflare who want a free, edge-speed control plane — caching, rate limiting, DLP — in front of any model API.
The strongest free tier in the category, run on the network a fifth of the web already transits. The Aug 27, 2025 refresh added dynamic routing (A/B tests, per-user limits, model chaining), Secrets Store key management, DLP scanning in the AI firewall, and 350+ models across major providers — with optional unified billing at a 5% credit fee only if you want one invoice.
Watch Routing intelligence is thinner than dedicated routers — no learned model selection, and provider translation coverage trails OpenRouter's catalog. Observability depth is Cloudflare-grade metrics, not LLM-native evals. Least compelling if you're not otherwise a Cloudflare shop.
5PortkeyPortkey → Palo Alto Networks (closed May 29, 2026)
Gateway fully open source (Mar 2026) · Developer free (10k logs/mo) · Production $49/mo (100k logs, +$9/100k) · Enterprise custom (VPC, SOC2/HIPAA)Best for Enterprises that treat the gateway as a governance layer — guardrails, RBAC, budgets, prompt management, audit — not just a switchboard.
The governance-first gateway at real scale: 1T+ tokens and 120M+ requests daily across 24,000+ organizations, $180M+ in managed AI spend (Mar 2026), with the full enterprise gateway open-sourced that same month. Palo Alto Networks acquired it (closed May 29, 2026) to anchor Prisma AIRS — the clearest possible signal that routing is becoming the enforcement point for AI security.
Watch The acquisition cuts both ways: the roadmap now serves Prisma AIRS, and the standalone product's independence is over. Buying it increasingly means buying Palo Alto. Pre-acquisition Portkey was also sub-scale commercially ($15M Series A) relative to its traffic.
1T+ tokens/day, 120M+ requests/day, 24,000+ orgs; gateway fully open-sourced (Mar 24, 2026) [src] · Palo Alto Networks completed acquisition May 29, 2026; Portkey becomes core of Prisma AIRS (terms undisclosed) [src] · Production tier $49/mo for 100k logs; enterprise adds VPC deploy, SSO, custom guardrails (pricing page, Aug 2026) [src]
Model Router: the top 8 compared
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
| Tool | Take rate / fees | Hosting | Open source | Catalog | Added latency | Routing depth | Enterprise |
|---|---|---|---|---|---|---|---|
| OpenRouter | 0% markup · 5.5% credit fee · BYOK 5% >1M req | Managed edge | No | 400+ models · 60+ providers | ~25 ms | Fallbacks · price/throughput sorts · ZDR routing | Workspaces · spend mgmt · ZDR |
| LiteLLM | Free · enterprise custom (not per-token) | Self-host (+ cloud) | Yes (MIT) | 140+ providers · ~1,900 models (claimed) | 0.66 ms p99 (vendor) | Fallbacks · load-balance · budgets · caching | SSO · SCIM · audit · air-gap |
| Vercel AI Gateway | 0% markup · 0% BYOK · metered add-ons | Managed CDN | No | 100s of models | sub-20 ms | Failover · provider ordering · per-request ZDR | Invoiced billing · allowlists |
| Cloudflare AI Gateway | Core free · +5% unified billing | Managed edge | No | 350+ models | edge (unpublished) | Dynamic routing · caching · rate limits | DLP · Secrets Store · Logpush |
| Portkey | OSS free · $49/mo prod · ent custom | Both | Yes (gateway) | All major providers | low (edge deploys) | Guardrails · load-balance · canary · prompt mgmt | RBAC · SSO · VPC · SOC2/HIPAA |
| Requesty | 5% all-in markup | Managed | Partial | 600+ models | <14 ms failover | Smart routing · semantic cache · guardrails | SSO · spend controls (EU angle) |
| Helicone AI Gateway | OSS free · cloud 0% markup (beta) | Both | Yes (Rust) | 100+ models | light (Rust, vendor) | Cheapest-provider routing · fallbacks · caching | Observability-native · SLA on cloud |
| Bifrost | OSS free · enterprise custom | Self-host | Yes (Apache-2.0) | 1,000+ models · 23+ providers | <15 µs @5k RPS (vendor) | Adaptive load-balance · semantic cache · MCP | Cluster mode · governance · budgets |
How to choose your model router
- If you just hit two models in production and have no infra team
- OpenRouter. One key, one invoice, ~25ms overhead; treat the 5.5% credit fee as rent until your spend justifies direct contracts.
- If keys, prompts, or logs cannot leave your VPC — or procurement says self-host
- LiteLLM open source. It's the de facto standard (Netflix, NVIDIA testimonials), and the Rust core removed the old performance excuse.
- If gateway fees offend you and you're in the TypeScript/AI SDK world
- Vercel AI Gateway — 0% markup and free BYOK is the best sticker price in the category; just watch the metered observability add-ons.
- If you're already behind Cloudflare and mostly need caching, rate limits, and a kill switch
- Cloudflare AI Gateway free tier first; add unified billing (+5%) only if the single invoice is worth it.
- If the gateway is your AI governance and security enforcement point (regulated enterprise, agent fleets)
- Portkey under Palo Alto — or self-host its now-open-source gateway if the acquisition roadmap worries you; Kong if you're already a Kong shop.
Model Router: the whole field
20 more tools tracked in this category, including 5 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.
| Tool | Maker | What it is | Entry | Status |
|---|---|---|---|---|
| Requesty | Requesty (Amsterdam) | EU-flavored OpenRouter alternative: 600+ models, 5% all-in markup, <14ms failover, 225B+ tokens/day claimed; $3M seed led by 20VC (Sep 2025) | 5% markup | active |
| Helicone AI Gateway | Helicone (YC W23) | Open-source Rust gateway from the observability company; cloud version with 0%-markup passthrough billing launched to beta Sep 2025 | free + tokens | active |
| Bifrost | Maxim AI | Go gateway marketing itself on speed — '50x faster than LiteLLM', <15µs overhead at 5k RPS (vendor benchmark), 1,000+ models, MCP built in | free + tokens | active |
| Kong AI Gateway | Kong Inc. | The API-gateway incumbent's AI layer: semantic routing/caching, prompt firewalling, multi-LLM plugins on Kong Gateway/Konnect | OSS free · Konnect tiers | active |
| Microsoft Foundry Model Router | Microsoft | Hyperscaler-native learned router, GA Nov 2025 (version 2025-11-18): picks among 28 models incl. GPT-5.x, Claude, DeepSeek, Grok inside Azure | Azure usage | active |
| Bedrock Intelligent Prompt Routing | AWS | Per-prompt routing between models of a family inside Bedrock; AWS-only answer to the router question | AWS usage | active |
| LLM Gateway | LLMGateway (OSS) | Open-source OpenRouter alternative; 0% fee with own keys, 5% on credits — publishes the category's best fee-comparison research | free + tokens | active |
| TrueFoundry AI Gateway | TrueFoundry | Enterprise self-hosted gateway (K8s-native); prolific comparison-content marketer aiming at LiteLLM/Portkey buyers | custom | active |
| Eden AI | Eden AI | Unified API beyond LLMs (vision, OCR, speech); 5.5% fee on credits — broader but shallower than LLM-first routers | credits + 5.5% | active |
| Not Diamond | Not Diamond | Trained per-prompt quality router (pick the best model, not just the cheapest); shipped a coding-agent router in 2026 — niche but alive | free tier + usage | active |
| Arch | Katanemo | Open-source agent-native proxy with its own Arch-Router model for preference-based routing | free + tokens | active |
| Apache APISIX AI Gateway | Apache Software Foundation | AI plugins (ai-proxy, rate-limiting, prompt guard) on the APISIX API gateway; solid if APISIX is already your edge | free | active |
| Higress | Alibaba (OSS) | Envoy-based cloud-native AI gateway, big in China deployments; token-aware routing and model failover | free | active |
| new-api | QuantumNous (OSS) | Popular Chinese-community multi-provider aggregator/reseller panel in the One API lineage | free + tokens | active |
| RouteLLM | LMSYS (research OSS) | The academic cost-quality router that popularized learned routing (2024); repo largely quiet since — ideas absorbed by commercial routers | free | fading |
| Glama Gateway | Glama | Low-fee gateway attached to an MCP-centric workspace; small but liked by indie agent builders | usage-based | active |
| Martian | Martian (withmartian) | Invented the 'model router' pitch ($9M, NEA/Prosus 2023; Accenture 2024) — has pivoted to AI interpretability research; router no longer the product | — | fading |
| Unify | Unify AI | Benchmark-driven LLM router (a16z-backed, 2024) — router retired; company pivoted to 'AI teammates' workflow agents by 2026 | — | dead |
| TensorZero | TensorZero | Rust gateway + LLMOps stack; archived its repo and shut down Jun 12, 2026, returning remaining capital — the category's starkest OSS casualty | — | dead |
| MLflow AI Gateway | Databricks (OSS) | Deployments/gateway module inside MLflow; maintained but overshadowed — Databricks invests via Mosaic AI Gateway instead | free | fading |
Model Router: the category in numbers
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
- OpenRouter raised a $113M Series B at ~$1.3B led by CapitalG (May 26–28, 2026) — the category's first unicorn; weekly tokens 5T → 25T in six months [src]
- Consolidation wave: Palo Alto Networks acquired Portkey (closed May 29, 2026) into Prisma AIRS; TensorZero shut down Jun 12, 2026; Martian and Unify both pivoted away from routing [src]
- Take rates converged: 0% token markup is now table stakes; money is made on ~5–5.5% credit/platform fees (OpenRouter 5.5%, Cloudflare 5%, Eden 5.5%) — and Vercel's 0%/0% undercuts them all (fee survey updated Aug 5, 2026) [src]
- Hyperscalers ship routing natively: Microsoft Foundry Model Router GA Nov 2025 now routes across 28 models including Claude, DeepSeek and Grok — squeezing standalone routers from above [src]
- Routers are the arbitrage layer: open-weight models hit 29% of Vercel AI Gateway tokens on under 4% of spend; Anthropic took 61% of spend on 32% of tokens (Jul 2026 production index) [src]
- Portkey open-sourced its full enterprise gateway (Mar 24, 2026) at 1T+ tokens/day across 24,000+ orgs — governance features are commoditizing fast [src]
Model Router: method & sources
Ranking criteria: effective take rate at real volume, catalog breadth, added latency, governance depth, and survivability — this category had a brutal 2026 (one shutdown, one acquisition, two pivots), so 'will it exist next year' is a first-order criterion. Conflicts resolved: (1) OpenRouter's ~$50M annualized revenue is Sacra's estimate, not company-disclosed — treated as directional; (2) a Medium post claiming Martian 'nearing $1.3B valuation' is unsourced and contradicted by its pivot to interpretability research — rejected; (3) Bifrost's '40x/50x faster than LiteLLM' and LiteLLM's '0.66ms p99' are both vendor benchmarks predating each other's latest rewrites — reported as claims, not facts; (4) model-count claims (400+, 600+, 1,892, 1,000+) are vendor-reported and drift weekly. Catalog counts and uptime figures are single-source (vendor pages). Adjacent elements: inference clouds serving open weights (Together, Fireworks, Groq, DeepInfra) are providers routers route TO → Ow; GPT-5-style in-product auto-routing lives inside the model app → Fm; observability-first tools (Langfuse, Helicone-as-monitor, Braintrust) → Ev; prompt/response filtering as a product → Gd. Cloud-desk aggregators reselling coding-agent subscriptions are out of scope (→ Ca). Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: rt.json.
All sources (24)
- https://openrouter.ai/blog/announcements/series-b/
- https://techcrunch.com/2026/05/26/openrouter-more-than-doubles-valuation-to-1-3b-in-a-year/
- https://sacra.com/c/openrouter/
- https://openrouter.ai/docs/faq
- https://www.litellm.ai/
- https://www.litellm.ai/pricing
- https://github.com/BerriAI/litellm
- https://vercel.com/blog/ai-gateway-is-now-generally-available
- https://vercel.com/docs/ai-gateway/pricing
- https://vercel.com/blog/ai-gateway-production-index-july-2026
- https://developers.cloudflare.com/ai-gateway/reference/pricing/
- https://blog.cloudflare.com/ai-gateway-aug-2025-refresh
- https://portkey.ai/pricing
- https://www.paloaltonetworks.com/company/press/2026/palo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents
- https://www.globenewswire.com/news-release/2026/03/24/3261574/0/en/portkey-s-gateway-is-now-fully-open-source-processing-over-1-trillion-tokens-every-day.html
- https://www.requesty.ai/
- https://www.requesty.ai/blog/requesty-raises-3m
- https://llmgateway.io/blog/ai-gateway-fees-compared
- https://github.com/maximhq/bifrost
- https://learn.microsoft.com/en-us/azure/foundry/foundry-models/whats-new-model-router
- https://byteiota.com/tensorzero-shuts-down-what-oss-llmops-cant-survive/
- https://www.helicone.ai/blog/ptb-gateway-launch
- https://unify.ai
- https://withmartian.com
Our take
A router is optional until you have two models in production. Then it's plumbing you'll wish you'd laid earlier.
Combines with
This is element 4 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.
Explore the full table →