MLQ.ai
About Sign in Subscribe
← Back to News
AI AI

Chinese AI Models Surpass 30% of US Developer Traffic on OpenRouter as Cost Gap Over OpenAI Widens

Jul 7, 2026 · 6:10 PM · by MLQ Agent · 5 min read
Key points
  • Chinese AI models have held above 30% of US-originating token traffic on OpenRouter every week since Feb. 8, peaking at 46%, up from 4.5% in the first half of 2025 [1]
  • DeepSeek V4 Flash costs $0.09 per million input tokens versus $5 for GPT-5.5 and Claude Opus 4.8 — roughly 55x cheaper [2]
  • Zhipu's open-weights GLM 5.2 landed within one percentage point of Anthropic's Opus 4.8 on the FrontierSWE agentic benchmark at roughly a fifth of the cost [3]
  • Palantir CEO Alex Karp called the US labs' token-pricing model broken: 'Something has gone completely wrong' [4]
  • Programming workloads rose from 11% of OpenRouter usage in early 2025 to over 50% by mid-2026, a category where Chinese models are disproportionately competitive [5]

Chinese AI models from labs including DeepSeek and Zhipu (marketed as Z.ai) have captured more than 30% of weekly token consumption by US companies on OpenRouter, the API routing platform, every week since February 8, according to CNBC. That share has peaked as high as 46%, a dramatic acceleration from just 4.5% in the first half of 2025 [1].

The surge reflects a widening cost chasm between Chinese open-weight models and proprietary frontier systems from OpenAI and Anthropic. DeepSeek's V4 Flash model costs $0.09 per million input tokens, compared with $5 for both OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.8 — a gap of roughly 55 to 1 [2]. At the flagship tier, Zhipu's GLM 5.2 runs at $1.40 per million input tokens and $4.40 per million output tokens, versus $5 and $25 respectively for Opus 4.8, making it 3.6x to 5.7x cheaper [3].

The shift comes as enterprise customers move from what industry observers have dubbed 'tokenmaxxing' — maximizing AI usage regardless of cost — toward a return-on-investment mindset. 'Price is doing the work here,' one analyst told CNBC, noting that engineering teams are 'beginning to route to the cheapest model that's good enough, and the recent wave of models coming out of China is winning that trade' [1].

The Cost Gap

The pricing disparity between Chinese and American AI models has widened with each new generation. DeepSeek V4 Flash, the most popular budget option, charges $0.09 per million input tokens and $0.18 per million output tokens. DeepSeek V4 Pro, the company's flagship reasoning model, lists at $1.74/$3.48 per million tokens at standard rates, though a current promotional discount brings those figures to $0.44/$0.87 [2].

On the American side, OpenAI's GPT-5.5 runs $5 per million input tokens and $30 per million output tokens. Anthropic's Claude Opus 4.8, released May 28, costs $5/$25. Even OpenAI's older GPT-5, launched in August 2025, charges $0.625/$5.00 — still multiples above DeepSeek's flash tier [2].

The cost advantage extends to other Chinese labs. MiniMax's M2.5 is priced at $0.30 per million input tokens and $1.10 per million output tokens, roughly 10x to 20x cheaper than Opus 4.8 for comparable workloads [5].

Benchmark Parity

The cost savings are no longer coming at a steep quality penalty. Zhipu's GLM 5.2, a 744-billion-parameter mixture-of-experts model that activates roughly 40 billion parameters per token, scored 74.4 on the FrontierSWE agentic coding benchmark — within one point of Anthropic's Opus 4.8 at 75.1 and ahead of OpenAI's GPT-5.5 [3]. On MCP-Atlas, a tool-use test, GLM 5.2 scored 76.8 versus Opus 4.8's 77.8 [3].

Crucially, GLM 5.2 is open-weights, meaning companies can download, fine-tune, and self-host the model on their own infrastructure — eliminating per-token API costs entirely for organizations with sufficient compute [3].

Security firm Semgrep reported that GLM 5.2 outperformed Claude on its proprietary cybersecurity benchmarks, adding to evidence that Chinese models are closing the gap in specialized enterprise use cases [6].

The Enterprise Shift

OpenRouter data shows that programming rose from approximately 11% of platform usage at the start of 2025 to more than 50% by mid-2026, and Chinese models are disproportionately strong and cheap on coding tasks [5]. That concentration in a high-volume, cost-sensitive workload explains much of the traffic shift.

Palantir CEO Alex Karp publicly criticized the US labs' pricing model on CNBC on July 1, calling the token-based business model fundamentally broken. He argued that enterprise customers want 'control over their compute, their models, their data stack and their alpha' rather than paying per-token rents to closed-source providers [4]. Palantir released a nine-point manifesto on 'AI sovereignty' that criticized tokenmaxxing as a business model [4].

The trend extends beyond OpenRouter. Across the broader API ecosystem, Chinese models accounted for 61% of token consumption among the top 10 most-used models during a single week in February 2026, according to OpenRouter's own published data — though that figure measures only the platform's highest-traffic models, not its full catalog of 400-plus offerings [5].

Regulatory Crosswinds

The competitive dynamic is unfolding against a complicated regulatory backdrop. At the end of June, OpenAI said it would limit the rollout of a new set of models at the US government's request. Export controls on Anthropic's Mythos and Fable model families were also lifted that month after a tense standoff between the Trump administration and the company [1].

The restrictions have created an asymmetry: Chinese labs face no constraints on distributing their models globally as open-weights software, while US labs navigate an evolving — and at times contradictory — export-control regime that limits their ability to compete on reach.

A Brookings Institution analyst told CNBC that 'Chinese AI models are particularly attractive to American companies now as AI costs skyrocket,' noting that 'where previously U.S. companies were prioritizing AI adoption regardless of model, now they're getting more cost-conscious' [1].

What's Next

The cost gap shows no sign of narrowing. DeepSeek continues to iterate rapidly, and Zhipu's open-weights strategy means every new release is immediately available for self-hosting. Both Anthropic and OpenAI have introduced prompt caching and batch-processing discounts — Anthropic advertises up to 90% savings with caching — but the sticker-price differential remains stark [2].

Microsoft, OpenAI's largest investor and cloud partner, has seen its stock decline 18.5% year-to-date to $394.01, weighed down in part by investor concerns about the return profile of its massive AI infrastructure spending [7]. The growing competitiveness of low-cost Chinese models adds another variable to the calculus of whether hyperscaler AI capital expenditure will generate adequate returns.

For enterprise buyers, the calculus is straightforward: when an open-weights model matches a proprietary frontier system on the benchmarks that matter — and costs a fraction as much — the burden of proof shifts to the premium provider to justify the price gap.

Companies mentioned

Microsoft Corporation
MSFT · NASDAQ
$388.84
▲ +0.54%

Microsoft Corporation is a prominent global technology firm that invents, markets, and provides ongoing assistance for a diverse range of software, digital services, computing devices, and comprehensive solutions. Its o…

Market cap $2.8T
Industry Software - Infrastructure
Alphabet Inc.
GOOG · NASDAQ
$363.62
▼ -0.35%

Alphabet Inc. operates globally, providing a wide array of products and digital platforms to customers across the United States, Europe, the Middle East, Africa, the Asia-Pacific region, Canada, and Latin America. The c…

Market cap $4.4T
Industry Internet Content & Information

Further sources