What You Should Know

In July, Moonshot AI's Kimi K3 climbed to third place globally on the Artificial Analysis leaderboard, trailing only Claude Fable 5 and GPT-5.6. The model is open-weight, priced at roughly a third to a fifth of its US frontier rivals, and reportedly cost a fraction of what OpenAI or Anthropic spend to train comparable systems. The gap that separated US and Chinese frontier models, once measured in six to twelve months, is now closer to weeks. On OpenRouter, the largest open AI-model marketplace, Chinese-developed models now hold eight of the top ten places by weekly token volume, and Anthropic is the only US lab still in the top ten. On Hugging Face, Alibaba's Qwen family has spawned more than 160,000 derivative models, 2.6 times Meta's footprint, and adds roughly 200 more every day.

file

Every other release seemingly harkens flashbacks of the original "DeepSeek moment" of January last year. This begs one to ask if this is bearish for the hundreds of billions of dollars in AI infrastructure spending committed by US hyperscalers. If a Chinese lab can approach frontier capability at a fraction of the price, why would demand for premium compute and premium models hold up? That instinct conflates two separate questions: whether Chinese open-weight competition threatens pricing power at the model layer, and whether it threatens aggregate demand for the compute underneath it.

Key Takeaways

  1. China's open-weight models have closed the capability gap with US frontier labs to a matter of weeks, priced at a third to a fifth of comparable US models. The cost advantage is structural rather than just distillation: cheap electricity, architectural efficiency, and cluster-level engineering all contribute, and China is effectively exporting cheap power in the form of tokens. Open-weight distribution is a deliberate strategy to commoditize the model layer and capture downstream monetization.
  2. Falling token prices historically expand total usage rather than shrink it (Jevons paradox), so aggregate compute demand is not the layer under threat. The actual casualty is frontier labs' ability to charge a premium for capability.
  3. The model layer has effectively polarized into a duopoly where the US commands the closed frontier tier, while China dominates the high-volume, mid-to-high-tier open segment.
  4. The primary capture of sustainable enterprise value remains in the US-aligned supply chain, concentrated heavily in upstream physical bottlenecks (semiconductors, advanced packaging, power, and thermal management) and deeply embedded enterprise workflow applications.

I. How Close Is The Gap, Really?

The benchmark gap between proprietary American frontier models and open-weight Chinese releases has reached historical lows. Today, Chinese open-weight models reliably deliver 90% to 95% of frontier performance across coding, reasoning, and mathematics benchmarks. Kimi K3 and GLM5.3 are currently tied for third place on Artificial Analysis's Intelligence Index, right behind Fable 5 and GPT-5.6. What was once a gap of years and then months, is now seemingly only weeks.

file

This capability convergence has transformed global inference traffic distribution. On OpenRouter, Chinese-developed models hold eight of the top ten places by weekly token volume. OpenAI and Google have dropped out of the top ten entirely. On Hugging Face, Alibaba's Qwen family has spawned over 160,000 derivative models, many times Meta's entire open-source footprint. A derivative is a developer who has already committed engineering time to a specific architecture, which is a stickier form of dependency than an API contract and accumulates without a renewal date. The two sources differ in that one tells you where AI demand is going to, and the other where AI development (and dependency) is happening.

Throughout this improvement, Chinese models have remained considerably cheaper than their US counterparts and with token demand projections universally forecasting an exponential increase in the next few years driven by the continued proliferation and development of agentic workflows, demand for these cheaper models will only increase. Goldman Sachs forecasts global token consumption growing more than 24-fold by 2030, reaching approximately 120 quadrillion tokens per month. As token demand increases, infrastructure operators route tasks by cost-per-token relative to the quality threshold the task actually requires. Frontier models...

Already a subscriber? Click here to log in.

Subscribe to Enjoy
Full Access to Our Services
Unlimited Chart & Data Access

Comprehensive data at your service
with key indicators for investment insights

Exclusive Reports & Insights

Exclusive flash reports
on key events and data

Powerful Toolbox & Features

Create your own charts and analysis
including performance backtesting

Insightful Community & Engagement

Hub of professionals to engage
in meaningful discussions and insights

Get answers from MM AI.

    • How have Chinese open-weight models narrowed the capability gap with US frontier labs?

      💡Chinese open-weight models have narrowed the capability gap with US frontier labs to a matter of weeks by consistently delivering 90% to 95% of frontier performance across coding, reasoning, and mathematics benchmarks. This rapid convergence is attributed to several structural advantages, including cheaper electricity, architectural efficiency like sparse Mixture-of-Experts (MoE) designs, and advanced cluster-level engineering. For example, Moonshot AI's Kimi K3 climbed to third place globally on the Artificial Analysis leaderboard, just behind Claude Fable 5 and GPT-5.6, while being priced at a third to a fifth of its US rivals.

    • What key advantages do Chinese AI models have over US counterparts in terms of cost?

      💡Chinese AI models possess key cost advantages over US counterparts primarily due to cheaper electricity, architectural efficiency, and sophisticated cluster-level engineering. Electricity in China's western computing hubs is roughly 40% to 60% cheaper than in the US, ranging from ¥0.28 to ¥0.35/kWh (3.9¢ to 4.9¢/kWh), which is the single largest recurring cost for inference at scale. Additionally, Chinese models often utilize sparse Mixture-of-Experts (MoE) architectures, activating only 3% to 5% of parameters per token, and employ cluster-level engineering to achieve system-level scale, compensating for individual chip limitations.

    • How does China's 'open-weight' distribution strategy influence the AI model market?

      💡China's 'open-weight' distribution strategy is a deliberate tactic to commoditize the AI model layer and capture downstream monetization, effectively eroding the pricing power of US frontier labs like OpenAI and Anthropic. By allowing the global developer community to download, fine-tune, and stress-test their models for free, Chinese labs create a sticky dependency, as seen with Alibaba's Qwen family spawning over 160,000 derivative models on Hugging Face. This strategy aims to route users and developers towards their monetized layers, such as cloud compute rental (Alibaba Cloud, Tencent Cloud) and advertising (Douyin, Baidu), rather than direct model sales.

    • Why does the Jevons paradox suggest increased compute demand despite falling token prices?

      💡The Jevons paradox suggests that increased compute demand will occur despite falling token prices because as the cost per unit of intelligence decreases, new and previously uneconomical use cases become viable, expanding total usage. Historically, this expansion in usage has outpaced the price decline. For instance, agent workflows consuming 10x to 1,000x more tokens become affordable at lower price points. Usage-weighted token prices are down approximately 40% since late June, yet AI spend per employee among heavy-using enterprises rose close to 50% month-on-month in July, illustrating this market expansion.

    • What is China's industrial strategy for AI models, mirroring its past in solar and EVs?

      💡China's industrial strategy for AI models mirrors its past approach in solar panel and electric vehicle manufacturing: identifying a cost advantage in a key input or process, aggressively scaling production with state-directed capital, and then exporting the resulting overcapacity into global markets at prices competitors cannot structurally match. This strategy aims to commoditize the model layer by open-sourcing weights and eroding the pricing power of competitors, thereby becoming the default starting point that routes users and developers toward China's monetized downstream layers and sovereign AI infrastructure exports.

    • Why do falling token prices not indicate an AI bubble or correction for US hyperscalers?

      💡Falling token prices do not indicate an AI bubble or correction for US hyperscalers because the Jevons paradox applies, where cheaper tokens expand the overall market by making new use cases economically viable, thus increasing aggregate compute demand. While model prices are collapsing, capital expenditure commitments keep rising because the expansion in usage outpaces the price decline. Demand for underlying compute is holding strong, as evidenced by H100 GPU rental prices rebounding off June lows while token prices continued to fall, indicating that the market for AI intelligence is widening rather than shrinking.

    • How do GPU rental prices indicate sustained demand for underlying AI compute?

      💡GPU rental prices indicate sustained demand for underlying AI compute by showing a rebound in H100 rental prices off their June lows, while token prices continued to fall. This dynamic suggests that even as the cost of AI intelligence decreases, the demand for the foundational hardware remains robust because cheaper tokens enable more extensive and novel AI applications. Furthermore, B200 GPU pricing has been even firmer, trading above its May peak, reinforcing that the market for physical compute infrastructure is not weakening but rather experiencing sustained high demand.

    • Why are Chinese pure-play AI labs experiencing structural net losses despite growth?

      💡Chinese pure-play AI labs are experiencing structural net losses despite growth due to an aggressive domestic 'involutionary' price war, with API prices for models cut to as low as $0.20 to $0.95 per million tokens. This state-directed export strategy, which mirrors past deployments in other technology sectors, leads to individual market participants being unable to defend profit margins, even for technologically sophisticated services. For example, Zhipu AI reported first-half revenue growth of 404% YoY but still posted a loss of over 2 billion yuan, underscoring the severe profitability challenges.

    • How do US closed-source models maintain premium pricing despite lower token volume?

      💡US closed-source models maintain premium pricing despite lower token volume by offering deterministic reasoning, robust service-level agreements (SLAs), compliance certifications (SOC2, HIPAA, FedRAMP), and enterprise data privacy protections, which open-weight Chinese models struggle to clear in Western corporate environments. Enterprise customers are willing to pay premium rates for these assurances, as evidenced by Anthropic alone accounting for approximately half of all dollar spend on Vercel's AI Gateway while processing roughly 12% of total token volume. This indicates that token volume does not directly translate to revenue capture in the high-value enterprise segment.

    • What investment opportunities are recommended, given the polarization of the AI model layer?

      💡Given the polarization of the AI model layer, recommended investment opportunities are primarily in US-aligned, supply-constrained hardware across Asia, insulated from the model price wars. These include semiconductors, advanced packaging, power, and thermal management. Investors should focus on physical bottlenecks where profits are heavily concentrated due to hard physical supply constraints. The analysis suggests avoiding Chinese LLMs as attractive investment opportunities, as China's rapid AI growth does not guarantee attractive margins, with price competition and US export controls limiting global monetization.

  • The Era of Compute Securitization Has Arrived: Will AI Repeat the 2008 Subprime Crisis? (2026-08-21) South Korea’s Bull Market Return: What the Economy and Leverage Data Tell Us (2026-08-20)

    Live Outlook + AI Supply Chain Hub Track the $1T AI megatrends and the entire supply chain—all in one plan. → Claim It Before Sep 30