Home / Blog / OpenRouter Rankings
ENGINEERING BLOG · 2026.07.27

OpenRouter Rankings July 2026:
Who's Actually Winning the AI Model Race

If you are still picking an LLM based on a benchmark chart from two months ago, you are already behind. OpenRouter — the largest neutral LLM routing marketplace — publishes something more honest than any vendor announcement: real, paid, production token volume. This article breaks down data through July 25, 2026: Xiaomi's Mimo V2.5 taking the top spot, Chinese labs crossing 46% market share, and the barbell split between volume leaders and quality leaders. Bottom line up front: the useful question is not who is #1 this week, but which side of the barbell your workload belongs on.

01

July's board moves faster than June's OpenRouter rankings, with fiercer competition inside the Chinese camp. If you still default to "US big three = safe choice," expect these hidden costs:

  • Confusing volume #1 with quality #1: OpenRouter ranks token usage, not capability. Mimo V2.5 at 1.4T tokens/day does not mean it fits your hardest agentic planning — Claude still leads spend on difficult reasoning.
  • Ignoring a ~35x price gap: DeepSeek V4 Flash lists around $0.05–$0.14 per million input tokens; GPT-5.5 sits near $5.00. The leaderboard reflects price sensitivity as much as raw ability.
  • Model charts only, app charts never: Hermes Agent alone holds ~45% of tracked app token share; roleplay and companion apps move serious volume that enterprise coverage barely mentions.
  • Single-vendor lock-in: The "model of the month" crown keeps rotating — MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July. See our OpenRouter integration guide for fallback architecture.

In one line: capability and popularity are diverging — Chinese open-weight models bought volume with price; US closed-frontier labs defend pricing power on hard tasks.

02

Model token figures below are as of July 25, 2026 (rankings shift daily — verify at openrouter.ai/rankings before citing).

Top 12 models by daily token volume (2026-07-25)
Rank Model Provider Daily tokens 30-day total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (new entry)
10Ling 3.0 FlashInclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Seven of the top ten model slots belong to Chinese labs; the US side is mainly Nemotron 3 Ultra, Claude, and Gemini. DeepSeek remains the most stable #1 provider by share (16–18%), but the monthly volume champion keeps rotating inside the Chinese camp.

Provider token share (7-day blended estimates)
Provider Origin Share (approx.)
DeepSeekChina16%–18%
XiaomiChina8%–18% (Mimo V2.5 spike)
AnthropicUS10%–15%
TencentChina8%–13%
GoogleUS8%–13%
Chinese labs combined~46% (was <2% a year ago)
US big three combined~30%–36% (was ~70%)
Pricing and positioning (July 2026)
Model Input/M Output/M Positioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.28Best value; agentic coding default
Nemotron 3 Ultra$0.42 (free tier exists)$2.61US open-weight, NVIDIA stack
GLM 5.2$0.45$3.31Closest open match to Opus planning
Kimi K3~$3~$151.4TB open weights, closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed frontier; top July benchmarks

03

Every OpenRouter recap should lead with this caveat: these rankings measure token volume, not capability. A cheap model behind one high-traffic consumer app can outrank a genuinely stronger model teams reserve for the hardest 10% of work.

Spend-by-task-category tells a different story: general chat 35.7%, agentic workflows 30.4%, code 26.5%, data work 7.5%. Drill into Classification and Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each, with GPT-5.5 third at 11.6% — the cheap open models dominating volume charts barely register here.

The market is bifurcating: cheap Chinese open-weight models absorb high-volume, error-tolerant workloads (chat, creative writing, roleplay, routine coding), while closed frontier models command pricing power on hard, low-error-tolerance work. Anthropic's July 24 Claude Opus 5 launch — FrontierBench v0.1 at 43.3% vs GPT-5.6 Sol's 37.5% at Opus-tier $5/$25 per million tokens — is the clearest proof point. See Claude Opus 5 and Kimi K3 controversy coverage and Kimi K3 deep dive for context.

04

The app leaderboard (openrouter.ai/apps) shows what those models are actually doing:

  • Hermes Agent (Nous Research's self-improving agent) holds ~45% of tracked app token share;
  • Coding agents fill most of the top 10: Kilo Code (~13%), OpenClaw (~9%), Claude Code (~6%), Cline (~1.7%);
  • The Cline → Roo Code → Kilo Code fork chain shows the youngest fork has overtaken its ancestors — first-mover advantage in open-source dev tooling does not last;
  • Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) move serious volume — OpenRouter × a16z's State of AI report finds creative roleplay accounts for more than half of open-model usage.
Top 10 apps by token share (approx.)
Rank App Category Share
1Hermes AgentPersonal agent / CLI~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7LemonadeCompanion / gaming~2.1%
8ISEKAI ZERORoleplay~2.0%
9Janitor AIRoleplay~1.8%
10ClineCoding agent (IDE)~1.7%

Six steps to build task-tiered routing from July's barbell pattern:

  1. Inventory tasks and error tolerance: Split workloads into chat/creative, agent orchestration, code, and hard reasoning; tag acceptable error rates and latency ceilings per tier.
  2. Validate on both volume and spend charts: Default high-volume work to DeepSeek V4 Flash and GLM 5.2; reserve Claude Opus 5 / GPT-5.6 for steps where cheaper models actually fail.
  3. Use OpenRouter as a prototyping sandbox: One API key across hundreds of models for fast A/B tests; measure latency and quality on your real tasks, not leaderboard rank.
  4. Implement hybrid routing to cut cost: Start coding flows on DeepSeek V4 Flash, escalate only on failure; see DeepSeek V4 GA pricing and migration guide.
  5. Add vendor safety to your scorecard: This week's OpenAI sandbox escape disclosure and proposed US "AI Kill Switch Act" make vendor safety track record a formal selection criterion for agentic deployments.
  6. Prepare stable compute for agent production: Local Claude Code / Cursor agents plus cloud API calls need 24/7 uptime; evaluate bare-metal Mac mini vs virtualized stacks for Hypervisor overhead and Metal toolchain compatibility.

05

August outlook based on July trajectory:

  • Chinese open-weight combined share likely climbs toward or past 50% unless a major US provider cuts price;
  • The monthly volume champion keeps rotating — watch August release cadence from Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot;
  • Anthropic may ship a cheaper volume-focused tier; Opus 5 is already their fourth flagship in under two months;
  • Kimi K3's 1.4TB weights will likely see community quantization within 2–4 weeks;
  • Security and governance become real selection criteria for enterprise procurement.
  • Mimo V2.5 daily volume: ~1.4 trillion tokens/day as of July 25; 31.2T over 30 days.
  • US–China price gap: DeepSeek V4 Flash ~$0.05–0.14/M input vs GPT-5.5 ~$5/M — roughly 35x.
  • Chinese share migration: From under 2% to ~46% in 12 months — one of the steepest share shifts in AI.
  • Claude Opus 5 benchmark: FrontierBench v0.1 at 43.3%; pricing $5/$25 per million tokens (fast tier $10/$50).

FAQ

What does OpenRouter rank by?

Real paid token volume — production routing choices developers pay for — not benchmark scores.

Which model topped OpenRouter in July 2026?

Xiaomi's Mimo V2.5 at roughly 1.4 trillion tokens per day as of July 25, 2026.

What share do Chinese models hold on OpenRouter?

Roughly 46% of identified token volume, up from under 2% a year ago.

Does high usage rank mean better quality?

Not necessarily. Route high-volume tolerant work to cheap open models; reserve premium closed models for hard tasks.

Primary sources below; rankings shift daily — verify current figures before citing.

OpenRouter official rankings

OpenRouter Apps leaderboard

OpenRouter Blog: DeepSeek V4 Adoption

OpenRouter × a16z State of AI report

The story to remember from July: capability and popularity are diverging. For teams running Claude Code, Kilo Code, or custom agents 24/7, API routing alone does not solve local toolchain, Metal acceleration, or uptime — virtualized Mac environments add Hypervisor overhead and iOS CI compatibility risk. For production environments that need zero-overhead native compute, stable iOS CI/CD, and 24/7 AI agent automation, ZUKCLOUD bare-metal Mac mini cloud nodes are usually the better fit: dedicated Apple Silicon hardware, no hypervisor tax, always-on, flexible daily/weekly/monthly billing. See our bare-metal architecture manifesto for the full argument.