Home / Blog / Price War
ENGINEERING BLOG · 2026.08.17

Why Did DeepSeek Just Raise API Prices by Up to 1,100%
Right After China's Open-Weight AI Blitz?

In a five-day window, teams buying API tokens, hosting open weights, or picking a coding model hit three moves that look contradictory on the surface. DeepSeek raised API prices by as much as 1,100% on certain tiers. Alibaba, the same week, open-weighted a 2.4-trillion-parameter flagship it had never released before. Zhipu AI shipped GLM-5.3, boosting coding benchmarks by roughly 6x on the exact same base model — no retraining. This piece gives the timeline, the price sheet, the three strategies, a six-step bill recast, and FAQ. Together they signal a shift: China's labs are competing on pricing power, not just price.

01

The pain is not "another model shipped." It is that price, license, and the performance lever all changed in the same week. Headline multiples mix tiers. Open weights are not a free commercial pass. A same-base post-train can jump overnight. Bills, compliance, and model choice all have to be redone:

  • Official peak now loses to resellers: DeepSeek's own peak rate is no longer the cheapest way to run DeepSeek.
  • Open weights, custom license: cross a revenue line and you negotiate a separate commercial grant.
  • Post-training beats a new base: GLM-5.3 shows you do not have to retrain a 743B foundation to lift coding scores.
  • US and China price in opposite directions: China open-weights flagships and tiers up; US labs go free and cheaper at the consumer layer.
July–August 2026: open weights, price hikes, and post-training
Date Event
Jul 16, 2026 Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny
Jul 30, 2026 OpenAI cuts GPT-5.6 Luna, its cheapest tier, by 80%
Aug 2–3, 2026 Alibaba previews, then launches, Qwen3.8-Max as a hosted API
Aug 6–7, 2026 OpenAI makes Luna the free default with unlimited text chats
Aug 10, 2026 Meta releases Muse Glimmer (30B, Apache 2.0), teases open weights for flagship Muse Spark 1.2
Aug 12, 2026 Alibaba publishes Qwen3.8-2.4T-A95B open weights on Hugging Face/ModelScope; xAI ships Grok 4.6
Aug 13, 2026 DeepSeek-V4-Pro goes GA and announces a price increase effective Aug 17; Google ships discounted Gemini 3.7 Flash
Aug 14, 2026 Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base
Aug 17, 2026, 00:00 Beijing time DeepSeek's new pricing takes effect

Zoom out: while Chinese labs raised prices and opened flagship weights, US labs cut prices and went free at the consumer layer — at the same time. That is two sides of one pricing fight. Product-level context: Qwen3.8-Max release, DeepSeek V4 GA, and the GPT-5.6 price cut.

02

Start with the sheet. Peak hours are 9am–12pm and 2pm–6pm Beijing time. "1,100%", "11x", and "350%" are all true — they just name different billing rows:

DeepSeek price hike (effective Aug 17, 00:00 Beijing time; per 1M tokens, RMB)
Billing item Old price New off-peak New peak Peak increase
V4-Flash cache hit (input) ¥0.02 ¥0.05 ¥0.10 ~400%
V4-Flash cache miss (input) ¥1.0 ¥1.5 ¥3.0 200%
V4-Flash output ¥2.0 ¥4.5 ¥9.0 350%
V4-Pro cache hit (input) ¥0.025 ¥0.15 ¥0.30 ~1,100%
V4-Pro cache miss (input) ¥3.0 ¥4.5 ¥9.0 200%
V4-Pro output ¥6.0 ¥13.5 ¥27.0 350%

The headline "1,100%" figure applies to peak-hour cache-hit input — the tier that started closest to free. Output pricing, which dominates most real bills, rose 350%. Independent cost modeling found that a realistic heavy-usage workload (roughly 84M tokens/month, mostly off-peak, half cache hits) sees a bill increase closer to 1.8x — real, but far below the scariest headlines.

Qwen3.8-2.4T-A95B (Qwen3.8-Max open weights): key specs
Spec Detail
Parameters 2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared)
Context window 262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max version defaults to 1M
Release cadence Preview Aug 2 → API live Aug 3 → open weights Aug 12
API pricing (international) $2/M input, $6/M output
License Not Apache 2.0 — a custom "Qwen3.8-Max License"
Why it matters First time Alibaba has open-weighted a Max-tier (flagship) model; Qwen3.5/3.6/3.7 Max stayed API-only
GLM-5.3 vs GLM-5.2: same base, post-training only (Zhipu self-reported)
Benchmark GLM-5.2 GLM-5.3 Change
Terminal-Bench 3.0 4.6% 28.3% +23.7 pts
DeepSWE v1.1 46.2% 66.9% +20.7 pts
Agents' Last Exam (CLI) 23.8% 28.5% +4.7 pts
CyberGym 77.2% 84.5% +7.3 pts
AutomationBench 26.2% 48.2% +22.0 pts

These are Zhipu's own reported numbers — no independent third-party re-run has been published yet. GLM-5.3 still trails GPT-5.6 Sol (34.6%) and Claude Fable 5 (33.7%) on Terminal-Bench 3.0; it is a top open-weight result, not an outright frontier win.

03

DeepSeek: from flat-rate to time-of-day — a capacity problem, not a strategy pivot

The easy misread is "China's cheapest model finally caved to margin pressure." Look at the structure and it reads as the opposite: a company making compute constraints visible on the price sheet for the first time. Flat, always-cheap pricing worked as a customer-acquisition tool as long as GPU capacity kept pace. Once usage grew exponentially and capacity did not, something had to become explicit — and "encouraging more flexible workload scheduling" is corporate-speak for "peak-hour compute is now scarce, please shift your load yourself."

One detail international coverage mostly missed: at peak hours, DeepSeek's official API is now higher than several third-party resellers (GMI Cloud, Novita, and others currently list V4 Pro below DeepSeek's new peak rate). The assumption that "the official API is always the cheapest way to run DeepSeek" has been broken for the first time.

Alibaba: open weights buy ecosystem goodwill; a custom license protects the revenue ceiling

Qwen3.8-Max's open-weighting is not a straightforward act of generosity. Alibaba did two things at once: it published the full 2.4T-parameter checkpoint for free download, and it attached a custom license — not the permissive Apache 2.0 used for smaller Qwen models — that requires any "Model-as-a-Service" or "AI Work Assistant" business earning over $50 million in any 12-month period to negotiate a separate commercial license, and requires products with 100M+ monthly active users or $20M+ in monthly revenue to prominently display the model's name.

The logic: give away the weights to win developer mindshare (especially internationally, where "made-in-China model" still carries hesitation among enterprise buyers), while keeping pricing leverage over the handful of companies that can build a competing inference business on top of it. That is a materially different bet than Meta's Muse Glimmer, which ships under unrestricted Apache 2.0. For the earlier open-weight wave, see Kimi K3's full weight release.

One rumor worth killing explicitly: claims circulated that Alibaba's license bans downloads from the US, EU, UK, and South Korea. That is false. The published license text contains no geographic or territorial clause of any kind — a useful reminder that in a release cycle this fast, checking the LICENSE file, not the announcement thread, takes seconds.

GLM-5.3: no new base model, just a bigger post-training bet

The most interesting fact about GLM-5.3 is not the score, it is the method: same 743B-parameter base as GLM-5.2, no retraining, and a roughly 6x jump on Terminal-Bench 3.0 (4.6% → 28.3%) purely from scaling up reinforcement learning environments in post-training. As pretraining scaling laws show diminishing returns, post-training RL scale is becoming an independent performance lever with a much lower cost floor than retraining a new foundation model. That is a meaningfully lower barrier to entry — mid-tier labs without OpenAI-scale compute budgets can still close the gap on agentic and coding benchmarks.

PRICE-WAR-2026.LOG
# three labs, three plays, one week of pricing power
DeepSeek V4-Pro  : flat cheap → peak/off-peak (capacity visible)
Qwen3.8-2.4T     : first Max-tier weights + custom license
GLM-5.3          : same 743B base, post-training only
rumor            : geo-ban US/EU/UK/KR — FALSE, no territorial clause
note: "Chinese model = cheapest" is no longer a safe assumption

04

Labs can ship a new price sheet in a day. Engineering and procurement still have to recast the bill, the license, and the inference box. Six steps, each tied to a public number above:

  1. Map the official sheet to your time zone first. DeepSeek peak is 9am–12pm and 2pm–6pm Beijing time (01:00–04:00 and 06:00–10:00 UTC). US and European business hours mostly land off-peak; China daytime sits on peak. Same sheet, different invoice.
  2. Split three billing rows. Do not stare at "1,100%". Cache-hit input rose the most in percentage terms and the least in absolute yuan. Output (+350%) and cache-miss input (+200%) move real bills. One third-party model of ~84M tokens/month, mostly off-peak, half cache hits, lands closer to 1.8x.
  3. Read the LICENSE file. Ignore the geo-ban rumor. The Qwen3.8-Max License has no territorial clause. The triggers are a Model-as-a-Service / AI Work Assistant business over $50M in any 12 months, or 100M MAU / $20M monthly revenue for prominent model-name display.
  4. Take "retrain a bigger base" off the default list. GLM-5.3 kept the 743B base and scaled post-training RL environments. Ask "can the post-train recipe squeeze another tier?" before "do we need a new foundation."
  5. Price across vendors. Drop "Chinese model = cheapest." DeepSeek V4-Pro off-peak input is about $0.63/M — still well below Claude Opus 5 (~$5/$25, implied from Alibaba's own comparison ratio) — but Qwen3.8-Max international ($2/$6) and GPT-5.6 Luna ($0.20/$1.20) now undercut it on at least one dimension.
  6. Size the real inference box. A 2.4T checkpoint is roughly 1.2TB of VRAM even at 4-bit. A laptop will not host it. Agents, evals, and long-horizon jobs still need real machines. If the workflow depends on Apple Silicon or the iOS toolchain, measure hypervisor tax; see the bare-metal architecture note.

05

Is DeepSeek still the cheapest frontier-class model?

Head-to-head (RMB-to-USD at ~¥7.15/$1, approximate)
Model Input (per 1M tokens) Output (per 1M tokens) Open weights?
DeepSeek V4-Pro (peak) ¥9.0 (~$1.26) ¥27.0 (~$3.78) No
DeepSeek V4-Pro (off-peak) ¥4.5 (~$0.63) ¥13.5 (~$1.89) No
Qwen3.8-Max (international API) $2.00 $6.00 Yes (custom license)
OpenAI GPT-5.6 Luna $0.20 $1.20 No
Claude Opus 5 (implied, per Alibaba's ratio) ~$5.00 ~$25.00 No

Short answer: no. Even after the hike, DeepSeek V4-Pro's off-peak rate is still well below Claude Opus 5, but it is no longer the outright cheapest option. "Chinese model = cheapest model" held for most of 2025 and early 2026. It is not a safe assumption anymore.

What's disputed or unverified

  • The "1,100%" headline is technically accurate but misleading without context. It applies only to peak-hour cache-hit input, the tier that started nearest to zero. Output — the cost that dominates most real bills — rose 350%. Different outlets have quoted different tiers as if they were the whole story.
  • Claims that Qwen3.8-Max runs on Alibaba's in-house Zhenwu M890 chips (and "Pangu AL128" supernodes), reported by several Chinese financial outlets as evidence of a fully domestic-silicon inference stack, have not been independently confirmed by Alibaba's own technical documentation or third-party benchmarks. Treat this as vendor-adjacent, unverified reporting until confirmed.
  • GLM-5.3's reported discovery of a "serious vulnerability" in Cursor comes from VentureBeat's reporting and Zhipu's own disclosure; specific technical details have not been made public, so the claim should be read as a vendor-sourced, not independently audited, security finding.
  • Reports that China's Ministry of Commerce may be preparing retaliatory export controls on AI/semiconductor technology are speculative and sourced to unconfirmed media reports, not an official announcement. Treat as background context, not established fact.

Why this matters: two price wars running in parallel

Over roughly the past month, China's top labs have shipped major releases at a pace domestic financial media has started calling "three model updates a week" (一周三更) — DeepSeek, Alibaba, and Zhipu, plus Moonshot's Kimi K3 (open-weighted Jul 16, 2.8T parameters) and MiniMax H3 before them. Chinese coverage broadly frames this as Chinese open-weight releases "forcing a global repricing of the AI industry."

Meanwhile, US labs are running the opposite play at the consumer layer: OpenAI cut prices 80% on its cheapest tier (Jul 30) then made that model free and unlimited for all users a week later (Aug 6–7); Google shipped a coding-focused model at half the price of its three-week-old predecessor (Aug 13). So while Chinese labs open-weight flagships and introduce tiered, higher pricing on the compute-constrained top end, US labs are racing toward free and cheap at the consumer end. Both are real strategies; they are just optimizing for different parts of the funnel.

There is also a geopolitical layer worth naming carefully. Moonshot's Kimi K3 open-weighting in July already drew US security scrutiny; Alibaba choosing this specific window to open-weight a 2.4T flagship has been read by some analysts as a move to lock in international mindshare and a "technological parity" narrative before any potential regulatory tightening. That is an informed interpretation, not a confirmed fact — but it is part of the context that is hard to see if you only read English-language tech press, which has largely covered these releases as isolated product news rather than as a coordinated national pattern.

Citeable figures (as of publication)

  • DeepSeek V4-Pro peak cache-hit input: ¥0.025 → ¥0.30, ~1,100%; output ¥6.0 → ¥27.0, 350%.
  • Qwen3.8-2.4T-A95B: 2.4T total / 95B active, 262,144 native tokens, custom license.
  • GLM-5.3 Terminal-Bench 3.0: 4.6% → 28.3% (same 743B base, no pretraining rerun).
  • FX reference: ~¥7.15/$1; heavy-usage bill lift ~1.8x in one third-party model (not official).

FAQ

Is DeepSeek still cheaper than GPT-5.6 or Claude after the price hike?

Its off-peak rate is still cheaper than Claude Opus 5, but it is no longer the single cheapest option overall — OpenAI's GPT-5.6 Luna ($0.20/$1.20 per million tokens) and Alibaba's international Qwen3.8-Max pricing ($2/$6) now undercut DeepSeek's new off-peak rates on at least one dimension. DeepSeek is still relatively cheap for a frontier-class model, just not the outright cheapest anymore.

Can I use Alibaba's Qwen3.8-Max open weights for free in a commercial product?

Yes, for most use cases — personal projects and internal enterprise use are unaffected. The catch applies only if you are running a "Model-as-a-Service" or "AI Work Assistant" business that has earned over $50 million in any consecutive 12-month period; that tier requires a separate commercial license from Alibaba. Products with 100M+ MAU or $20M+ monthly revenue must also display the model name prominently.

Is Qwen3.8-Max banned or restricted for US, EU, or UK users?

No. That claim circulated online but is false — the published license contains no geographic restriction of any kind. The restrictions are revenue-based (tied to how much money your service makes), not tied to where you or your users are located.

What's actually different between GLM-5.3 and GLM-5.2?

Nothing at the base-model level — both use the same 743-billion-parameter foundation model. The performance gains (roughly 6x on Terminal-Bench 3.0) come entirely from scaling up reinforcement learning during post-training, with no retraining of the base model.

Will Meta actually open-source its flagship model, not just the smaller Muse Glimmer?

Not yet. Muse Glimmer is a 30B distilled model, not Meta's real flagship. CEO Mark Zuckerberg has said open weights for the larger, closed Muse Spark 1.2 are coming "soon," which — if it happens — would make it the first US flagship-tier model released openly. As of this writing, that release hasn't happened; treat it as a stated intention, not a confirmed fact.

Pricing, license terms, and benchmark figures reflect publicly available information as of publication. Verify the latest official pricing and license terms before republishing, and note that details flagged above as unverified (domestic chip claims, the Cursor vulnerability report, and export-control rumors) have not been independently confirmed:

Official sources:

DeepSeek official pricing: Models & Pricing

Alibaba Qwen official repo: Qwen3.8-2.4T-A95B on Hugging Face

Qwen3.8-Max License text (Hugging Face LICENSE)

Zhipu (Z.ai) official GLM-5.3 technical page

Third-party reporting:

CNA / Reuters: DeepSeek raises API pricing for its V4 models

South China Morning Post reporting on Qwen license terms

VentureBeat coverage of GLM-5.3 and Muse Glimmer

A cloud API can print "shift your load" on a price sheet. Engineering teams still have to run agents, evals, and long-context jobs on real machines — reseller rates will reprice, a 2.4T checkpoint will not fit a laptop, and virtualized cloud instances still tax you with hypervisor overhead and weak Apple Silicon / iOS toolchain fit. If you need zero-loss native compute, stable iOS CI/CD, and 7×24 agent automation in a controlled physical box, ZUKCLOUD's bare-metal Mac mini cloud nodes are usually the better fit: dedicated Apple Silicon, no hypervisor tax, 7×24 online, billed by the day, week, or month. Start on the pricing page or go straight to the order page.