If you are still picking an LLM based on a benchmark chart from two months ago, you are already behind. OpenRouter — the largest neutral LLM routing marketplace — publishes something more honest than any vendor announcement: real, paid, production token volume. This article breaks down data through July 25, 2026: Xiaomi's Mimo V2.5 taking the top spot, Chinese labs crossing 46% market share, and the barbell split between volume leaders and quality leaders. Bottom line up front: the useful question is not who is #1 this week, but which side of the barbell your workload belongs on.
01 Still using last year's mental model? Four selection pain points
July's board moves faster than June's OpenRouter rankings, with fiercer competition inside the Chinese camp. If you still default to "US big three = safe choice," expect these hidden costs:
- Confusing volume #1 with quality #1: OpenRouter ranks token usage, not capability. Mimo V2.5 at 1.4T tokens/day does not mean it fits your hardest agentic planning — Claude still leads spend on difficult reasoning.
- Ignoring a ~35x price gap: DeepSeek V4 Flash lists around $0.05–$0.14 per million input tokens; GPT-5.5 sits near $5.00. The leaderboard reflects price sensitivity as much as raw ability.
- Model charts only, app charts never: Hermes Agent alone holds ~45% of tracked app token share; roleplay and companion apps move serious volume that enterprise coverage barely mentions.
- Single-vendor lock-in: The "model of the month" crown keeps rotating — MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July. See our OpenRouter integration guide for fallback architecture.
In one line: capability and popularity are diverging — Chinese open-weight models bought volume with price; US closed-frontier labs defend pricing power on hard tasks.
02 OpenRouter July 2026 leaderboard: top models and provider share
Model token figures below are as of July 25, 2026 (rankings shift daily — verify at openrouter.ai/rankings before citing).
| Rank | Model | Provider | Daily tokens | 30-day total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot AI | 157.6B | 1.6T (new entry) |
| 10 | Ling 3.0 Flash | InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
Seven of the top ten model slots belong to Chinese labs; the US side is mainly Nemotron 3 Ultra, Claude, and Gemini. DeepSeek remains the most stable #1 provider by share (16–18%), but the monthly volume champion keeps rotating inside the Chinese camp.
| Provider | Origin | Share (approx.) |
|---|---|---|
| DeepSeek | China | 16%–18% |
| Xiaomi | China | 8%–18% (Mimo V2.5 spike) |
| Anthropic | US | 10%–15% |
| Tencent | China | 8%–13% |
| US | 8%–13% | |
| Chinese labs combined | — | ~46% (was <2% a year ago) |
| US big three combined | — | ~30%–36% (was ~70%) |
| Model | Input/M | Output/M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | Best value; agentic coding default |
| Nemotron 3 Ultra | $0.42 (free tier exists) | $2.61 | US open-weight, NVIDIA stack |
| GLM 5.2 | $0.45 | $3.31 | Closest open match to Opus planning |
| Kimi K3 | ~$3 | ~$15 | 1.4TB open weights, closed-tier capability |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | Closed frontier; top July benchmarks |
03 Cheapest LLM API for coding? Usage rank is not a quality signal
Every OpenRouter recap should lead with this caveat: these rankings measure token volume, not capability. A cheap model behind one high-traffic consumer app can outrank a genuinely stronger model teams reserve for the hardest 10% of work.
Spend-by-task-category tells a different story: general chat 35.7%, agentic workflows 30.4%, code 26.5%, data work 7.5%. Drill into Classification and Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each, with GPT-5.5 third at 11.6% — the cheap open models dominating volume charts barely register here.
The market is bifurcating: cheap Chinese open-weight models absorb high-volume, error-tolerant workloads (chat, creative writing, roleplay, routine coding), while closed frontier models command pricing power on hard, low-error-tolerance work. Anthropic's July 24 Claude Opus 5 launch — FrontierBench v0.1 at 43.3% vs GPT-5.6 Sol's 37.5% at Opus-tier $5/$25 per million tokens — is the clearest proof point. See Claude Opus 5 and Kimi K3 controversy coverage and Kimi K3 deep dive for context.
04 What is actually getting used: app leaderboard and 6-step routing guide
The app leaderboard (openrouter.ai/apps) shows what those models are actually doing:
- Hermes Agent (Nous Research's self-improving agent) holds ~45% of tracked app token share;
- Coding agents fill most of the top 10: Kilo Code (~13%), OpenClaw (~9%), Claude Code (~6%), Cline (~1.7%);
- The Cline → Roo Code → Kilo Code fork chain shows the youngest fork has overtaken its ancestors — first-mover advantage in open-source dev tooling does not last;
- Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) move serious volume — OpenRouter × a16z's State of AI report finds creative roleplay accounts for more than half of open-model usage.
| Rank | App | Category | Share |
|---|---|---|---|
| 1 | Hermes Agent | Personal agent / CLI | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 6 | pi | Agent | ~3.3% |
| 7 | Lemonade | Companion / gaming | ~2.1% |
| 8 | ISEKAI ZERO | Roleplay | ~2.0% |
| 9 | Janitor AI | Roleplay | ~1.8% |
| 10 | Cline | Coding agent (IDE) | ~1.7% |
Six steps to build task-tiered routing from July's barbell pattern:
- Inventory tasks and error tolerance: Split workloads into chat/creative, agent orchestration, code, and hard reasoning; tag acceptable error rates and latency ceilings per tier.
- Validate on both volume and spend charts: Default high-volume work to DeepSeek V4 Flash and GLM 5.2; reserve Claude Opus 5 / GPT-5.6 for steps where cheaper models actually fail.
- Use OpenRouter as a prototyping sandbox: One API key across hundreds of models for fast A/B tests; measure latency and quality on your real tasks, not leaderboard rank.
- Implement hybrid routing to cut cost: Start coding flows on DeepSeek V4 Flash, escalate only on failure; see DeepSeek V4 GA pricing and migration guide.
- Add vendor safety to your scorecard: This week's OpenAI sandbox escape disclosure and proposed US "AI Kill Switch Act" make vendor safety track record a formal selection criterion for agentic deployments.
- Prepare stable compute for agent production: Local Claude Code / Cursor agents plus cloud API calls need 24/7 uptime; evaluate bare-metal Mac mini vs virtualized stacks for Hypervisor overhead and Metal toolchain compatibility.
05 August outlook, citable data, and FAQ
August outlook based on July trajectory:
- Chinese open-weight combined share likely climbs toward or past 50% unless a major US provider cuts price;
- The monthly volume champion keeps rotating — watch August release cadence from Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot;
- Anthropic may ship a cheaper volume-focused tier; Opus 5 is already their fourth flagship in under two months;
- Kimi K3's 1.4TB weights will likely see community quantization within 2–4 weeks;
- Security and governance become real selection criteria for enterprise procurement.
- Mimo V2.5 daily volume: ~1.4 trillion tokens/day as of July 25; 31.2T over 30 days.
- US–China price gap: DeepSeek V4 Flash ~$0.05–0.14/M input vs GPT-5.5 ~$5/M — roughly 35x.
- Chinese share migration: From under 2% to ~46% in 12 months — one of the steepest share shifts in AI.
- Claude Opus 5 benchmark: FrontierBench v0.1 at 43.3%; pricing $5/$25 per million tokens (fast tier $10/$50).
FAQ
What does OpenRouter rank by?
Real paid token volume — production routing choices developers pay for — not benchmark scores.
Which model topped OpenRouter in July 2026?
Xiaomi's Mimo V2.5 at roughly 1.4 trillion tokens per day as of July 25, 2026.
What share do Chinese models hold on OpenRouter?
Roughly 46% of identified token volume, up from under 2% a year ago.
Does high usage rank mean better quality?
Not necessarily. Route high-volume tolerant work to cheap open models; reserve premium closed models for hard tasks.
Primary sources below; rankings shift daily — verify current figures before citing.
OpenRouter Blog: DeepSeek V4 Adoption
OpenRouter × a16z State of AI report
The story to remember from July: capability and popularity are diverging. For teams running Claude Code, Kilo Code, or custom agents 24/7, API routing alone does not solve local toolchain, Metal acceleration, or uptime — virtualized Mac environments add Hypervisor overhead and iOS CI compatibility risk. For production environments that need zero-overhead native compute, stable iOS CI/CD, and 24/7 AI agent automation, ZUKCLOUD bare-metal Mac mini cloud nodes are usually the better fit: dedicated Apple Silicon hardware, no hypervisor tax, always-on, flexible daily/weekly/monthly billing. See our bare-metal architecture manifesto for the full argument.