FREE CONSULTATION
PROGRAMMATIC CPM$4.21â–²1.2%RETAIL MEDIA$148Bâ–²3.4%CTV INVENTORY86%â–¼0.8%AD-TECH INDEX2,914â–²0.6%CREATOR EARNINGS$31Bâ–²5.1%SEARCH SPEND$92Bâ–²1.9%COOKIE COVERAGE32%â–¼4.0%SOCIAL AD ROI3.8xâ–²0.3xPROGRAMMATIC CPM$4.21â–²1.2%RETAIL MEDIA$148Bâ–²3.4%CTV INVENTORY86%â–¼0.8%AD-TECH INDEX2,914â–²0.6%CREATOR EARNINGS$31Bâ–²5.1%SEARCH SPEND$92Bâ–²1.9%COOKIE COVERAGE32%â–¼4.0%SOCIAL AD ROI3.8xâ–²0.3x
Last updated: Tuesday, August 04, 2026

Chinese AI Models Are Taking on US AI Giants

Laptop displaying GLM Coding Plan with a Chinese flag on a dark red desk

Alibaba launched a 2.4-trillion-parameter model on Monday. DeepSeek’s newest runs within a point of OpenAI’s budget model at 40% less per task. One startup moved 100% of its traffic off Claude and says it will save millions.

Published: Monday, 3 August 2026 | BrandClickX News Desk

Disclosure, please read first. This article is about competitive pressure on US AI labs, principally OpenAI and Anthropic. Anthropic makes Claude, which was used to write this article. That is a direct conflict of interest. We have included the material unfavourable to Anthropic including a customer moving entirely off its models, and criticism of its conduct and readers should weigh the analysis accordingly.

Summary

Chinese open-weight AI models have taken more than 30% of US-routed developer tokens every week since February 2026, peaking near 46%, against roughly 11% in the previous year. Alibaba launched Qwen 3.8-Max on Monday 3 August, a 2.4-trillion-parameter model positioned against the best US systems. Chinese models cost 60% to 90% less than comparable US offerings.

Key Takeaways

  • Chinese models have exceeded 30% of US-routed tokens weekly since February, peaking near 46%
  • The prior twelve-month average was about 11%
  • Chinese open models cost 60–90% less than leading US models
  • Citi puts leading Chinese pricing as low as $0.18 per million tokens against roughly $4 for US frontier models
  • Alibaba launched Qwen 3.8-Max, 2.4 trillion parameters, on 3 August
  • DeepSeek V4 Flash runs within a point of GPT-5.6 Luna at 40% less per task
  • Lindy moved 100% of its traffic from Claude to DeepSeek in June
  • The US has no open-weights model competitive with the Chinese leaders
  • Data risk depends on deployment mode, not on who trained the weights

The Adoption Numbers

This is not speculative. The traffic has already moved.

MetricFigureSource
Chinese share of US-routed tokens, weekly since 8 Feb 2026>30%OpenRouter, via CNBC
Peak share~46%OpenRouter, via CNBC
Prior twelve-month average~11%OpenRouter, via CNBC
Chinese model cost vs leading US models60–90% cheaperOpenRouter, via CNBC
Leading Chinese model pricingas low as $0.18 per million tokensCiti
Top US frontier model pricingroughly $4 per million tokensCiti

That is a price ratio of more than twenty to one at the extremes.

The single clearest case: Z.ai’s GLM 5.2, released in late June, recorded the fastest adoption of any model Vercel tracked in 2026. In its first full week, daily token volume grew about 27x and customer count about 80x, according to Harpreet Arora, Vercel’s head of agentic infrastructure.

His explanation was blunt: “Price is doing the work here.” He added that when a task does not need the best model, teams are routing it to the cheapest one that is good enough and Chinese models are winning that trade.

What Launched This Week

Alibaba released Qwen 3.8-Max on Monday 2.4 trillion parameters, positioned directly against the top models from Anthropic and OpenAI.

It followed DeepSeek V4 Flash 0731 by days. According to independent benchmarks from Artificial Analysis, that model performs within a single point of OpenAI’s budget-oriented GPT-5.6 Luna while costing 40% less per task.

The parameter count matters as much as the benchmark. At 284 billion parameters, DeepSeek V4 Flash is small enough to run on relatively modest enterprise servers and workstations — meaning companies can host it themselves rather than paying per token at all.

Other models in the same wave include Moonshot’s Kimi K3 and Z.ai’s GLM 5.2.

The uncomfortable comparison: America’s most capable open-weights model, Inkling, is just under a billion parameters, and by The Register’s assessment cannot keep pace with DeepSeek’s current releases. On open weights specifically, the US is not competitive.

Alibaba’s Qwen family passed Meta’s Llama as the most downloaded model family on Hugging Face back in September 2025.

Why Enterprises Are Switching

Because AI budgets are being exhausted faster than anyone planned for.

Uber CEO Dara Khosrowshahi described it on the Invest Like the Best podcast: “We blew through our AI budget in a quarter, you know, for the whole year, essentially. And it is forcing us to adjust.” Reporting indicates Uber exhausted its entire 2026 AI budget in four months after employees rapidly adopted AI coding tools, prompting usage limits.

That pattern is repeating across corporate America, and it has changed how model selection works. Cost is now an architectural variable rather than a procurement footnote.

The most direct example involves Anthropic. In June, AI automation startup Lindy moved 100% of its traffic from Anthropic’s Claude models to DeepSeek, saying the switch would save it millions.

Lindy’s chief executive Flo Crivello put the broader position starkly: “The open-source scene right now is absolutely dominated by the Chinese. It’s not even close.” He said every founder he knows working in AI has either switched or is considering it.

The Structural Reasons, Not Just Price

CSIS identified two mechanisms that explain why this gap may persist rather than close.

Open weights create price competition that closed models do not face. With a closed US model, the owner controls the API, serving infrastructure, safety layer, compliance system, tools and uptime guarantees and largely sets the price. With open weights, many providers can host the same model and compete on inference efficiency. A model hostable by anyone faces structural downward price pressure.

Chinese labs have lower training costs, in CSIS’s assessment, partly because they do not carry the same margin for error that US companies do.

Neither of those is a temporary condition.

The Trust Problem the US Created

CSIS identified a self-inflicted factor, and it is worth reporting precisely.

On 12 June 2026, the US government suspended foreign access to Anthropic’s Fable and Mythos models, which led to Anthropic withdrawing them entirely. The Department of Commerce lifted those controls on 30 June, and Anthropic restored access on 1 July.

The suspension lasted under three weeks. CSIS’s argument is that the damage is not measured in days:

If foreign firms believe US model access can be withdrawn quickly and unilaterally, they will diversify. Some will choose Chinese open-weight models. Some will choose local sovereign models. Some will use multiple providers specifically to avoid dependence on Washington.

The strategic point CSIS makes is that US compute advantage without a global user base does not deliver the leverage Washington expects.

The Criticism of US Labs

The Register’s framing is unflattering to Anthropic and OpenAI, and we are including it rather than omitting it.

Its account is that Chinese developers are attacking on both price and performance, while US labs respond by raising concerns about the origins and safety of Chinese models. The implication that safety warnings from commercial rivals serve a competitive purpose alongside a security one is a fair criticism to record, whatever its merits.

The counter-argument deserves equal space. Data governance concerns about Chinese models are not invented. The direct API path for many of these models runs through China, raising genuine data-residency and compliance questions, and the issue has drawn congressional scrutiny.

The Nuance Most Coverage Misses

Security analysts converge on a single point: the data risk is set by deployment mode, not by who trained the weights.

Because these models are open-weight, a US or European company can download them and run them inside its own jurisdiction capturing most of the price advantage without sending any data to China.

The practical constraints are real: running large open-weight models requires high-performance GPUs subject to US export controls, plus staff able to patch quickly. Many teams therefore use a Western provider that serves the open weights, which delivers most of the cost saving without operating the hardware directly.

That distinction reframes the entire debate. The question for an enterprise is not “American or Chinese model” but “which weights, hosted where.” Coverage that collapses those into one decision misstates the choice.

Timeline

DateDevelopment
Sept 2025Qwen overtakes Llama as most-downloaded model family on Hugging Face
8 Feb 2026Chinese share of US-routed tokens passes 30% and stays there
June 2026Z.ai releases GLM 5.2; Lindy moves 100% off Claude to DeepSeek
12 Jun 2026US suspends foreign access to Fable and Mythos
30 Jun 2026Commerce lifts controls; access restored 1 July
7 Jul 2026CNBC publishes OpenRouter traffic data
31 Jul 2026DeepSeek V4 Flash 0731 released
3 Aug 2026Alibaba launches Qwen 3.8-Max

Expert Analysis

Three things are true simultaneously, and most coverage picks one.

The US still leads at the frontier. On the hardest reasoning, the longest agentic runs and the most reliable execution, American models retain an edge. Alibaba’s launch is positioned as “toe-to-toe,” which is a claim rather than a settled benchmark result.

The frontier is not where most workloads live. The majority of enterprise AI tasks do not require the best available model. Once a cheaper model is good enough, the price differential decides and Vercel’s adoption data shows exactly that happening.

The pricing pressure is structural. Open weights invite competing hosts, competing hosts compete on price, and no closed model can match that dynamic. This is not a promotional discount that expires.

The commercial squeeze is real for US labs. Both OpenAI and Anthropic have filed confidential IPO paperwork. Public-market scrutiny of profitability arrives at precisely the moment their pricing is under attack from below. Raising prices to demonstrate margin accelerates the defection; cutting prices delays profitability. That is a genuine bind.

What would change the picture: a US open-weights model competitive with DeepSeek and Qwen, which does not currently exist; a demonstrated capability gap large enough that “good enough” stops being good enough; or regulatory action restricting Chinese model use in US enterprises, which has been discussed but not enacted.

Frequently Asked Questions

How much US AI traffic goes to Chinese models?

More than 30% of US-routed tokens every week since 8 February 2026, peaking near 46%, according to OpenRouter data reported by CNBC. The average over the previous twelve months was about 11%.

How much cheaper are Chinese AI models?

Between 60% and 90% cheaper than leading OpenAI and Anthropic models, according to OpenRouter. Citi put leading Chinese pricing as low as 18 cents per million tokens against roughly $4 for top US frontier models.

What is Qwen 3.8-Max?

A 2.4-trillion-parameter model launched by Alibaba on 3 August 2026, positioned against the best models from Anthropic and OpenAI. Alibaba’s Qwen family overtook Meta’s Llama as the most downloaded on Hugging Face in September 2025.

Are Chinese models as good as US models?

Not at the very top. US models retain an edge on the hardest reasoning and longest agentic tasks. For most enterprise workloads, independent benchmarks show Chinese models close enough that price becomes the deciding factor.

Is it safe to use Chinese AI models?

Security analysts say the risk depends on deployment mode rather than who trained the model. Using a Chinese-hosted API raises data-residency concerns; running the same open weights in your own US or EU environment does not.

Which companies have switched?

AI automation startup Lindy moved 100% of its traffic from Anthropic’s Claude to DeepSeek in June, saying it would save millions. Reporting has named other US developers routing workloads to Chinese models, though not all have confirmed publicly.

Why are US models so expensive?

Closed models mean a single owner controls the API, infrastructure and pricing. Open-weight models can be hosted by many providers competing on inference cost, which creates structural downward price pressure that closed models do not face.

Did US export controls make this worse?

CSIS argues so. The suspension of foreign access to Anthropic’s Fable and Mythos models on 12 June 2026, though reversed by 1 July, signalled that US model access could be withdrawn unilaterally encouraging foreign firms to diversify.

Conclusion

The traffic data settles the descriptive question. Chinese open-weight models went from roughly a tenth of US developer tokens to nearly half at peak, in under a year, and the movement is driven by price rather than politics.

What it does not settle is whether this is a squeeze or a rout. US labs still lead where difficulty is highest. But that ground is narrow, most workloads do not sit on it, and the pricing pressure comes from a structural property of open weights rather than a promotional strategy.

The uncomfortable detail for Washington is that a three-week suspension of foreign access to American models, since reversed, may have done more to encourage diversification than any Chinese release. Capability advantage means little if customers conclude it can be switched off.

 | Chinese AI Models Are Taking on US AI Giants

Vikas Verma

Vikas Verma is an Editorial Contributor at BrandClickX, covering industry news, agency developments, and commerce trends shaping modern business growth.
Vikas@brandclickx.com

Scroll to Top