Skip to content
SIGPULSE
China AI Watch Issue 2 3 min read raw .md ↗

Why Are US Startups Routing 30%+ of Tokens to Chinese AI Models? The Inside-China Read

Chinese original 2026-08-08 · 「谁在给硅谷换AI底座?8周5款中国大模型正在改写规则」 · translated to English 2026-08-26

In the eight weeks ending August 2026, five Chinese flagship models launched in succession — and for the first time, the essay claims, the five most-called models on OpenRouter were all Chinese, despite 47% of the platform’s users being US developers [unverified]. US enterprises’ token consumption on Chinese models reportedly rose from 4.5% eighteen months ago to a stable 30%+, peaking at 46% [unverified], as companies like Coinbase and DoorDash shifted routine workloads to cheaper, open-weight Chinese models.

The numbers

  • 8 weeks, 5 flagship models
  • OpenRouter: top 5 most-called models all Chinese (July 2026) [unverified]
  • 47% of OpenRouter users are US developers; 6% Chinese
  • US-enterprise token share on Chinese models: 4.5% → 30%+ stable since February 2026, peak 46% [unverified]
  • GLM-5.2: mid-June, topped CodeArena; Kimi K3: July 16, 2.8T parameters; DeepSeek-V4-Flash: late July; Qwen3.8-Max: early August, 2.4T-parameter MoE; Seedance 2.5: early August
  • Stanford 2026 AI Index: US–China top-model performance gap narrowed to 2.7% [unverified]
  • Lindy (25-person startup): API costs down ~90% after migrating from Claude to DeepSeek-V4; bills previously exceeded total payroll
  • Coinbase: AI spending nearly halved; internal survey — 91% of engineers don’t need frontier-level performance
  • Tenstorrent: costs down 5× after switching from Claude
  • Cost comparison: ~$25 (Claude) vs ~$0.18 (DeepSeek) for the same coding workload; some Chinese models cost ~1% of US closed flagships
  • Qwen APP update August 7: free research features, scheduled tasks, PC-controlling office assistant

The take inside China

The story Chinese tech media are telling: a release blitz with no precedent. Eight weeks, five flagships — covering text, code, multimodal, and video — while OpenAI and Anthropic update flagships on a yearly cadence.

The essay’s argument: this isn’t brute-force compute. It’s a different technical bet — MoE architecture plus inference-side optimization plus open weights — instead of one giant closed model recouped via high API prices. The payoff shows in adoption: the named switchers are telling (Lindy, Coinbase, DoorDash, Tenstorrent), and the candid counterpoint gets airtime too: US engineers note gaps remain in extreme deep reasoning. The real pattern, the essay argues, is “layered mixing” — Chinese models for the 80% of routine work, US flagships for the hard 20%. Its pointed question: if China eats 80% of workload volume, can the remaining 20% sustain US closed-model valuations?

Meanwhile, the Qwen APP update on August 7 — free research, scheduled tasks, a PC-controlling office assistant — takes the model layer straight to ordinary users; viral Excel-and-PPT demos on Douyin and Bilibili did marketing no benchmark could.

Why it matters outside

If the cost gap (reportedly up to two orders of magnitude) and the open-weight advantage hold, Western AI pricing and business models face real pressure — and enterprises gain leverage via multi-vendor hedging. The trajectory also complicates the assumption that export controls can hold back Chinese frontier AI. English coverage of the same wave: The Register, Reuters, China Daily.

Sources

Provenance & disclosure. Originally published in Chinese on our WeChat channel on 2026-08-08 (“谁在给硅谷换AI底座?8周5款中国大模型正在改写规则”); drafted with AI assistance under human editorial direction. Translated to English on 2026-08-26 (AI-assisted, human-reviewed). The OpenRouter share figures, named-company quotes, and the Stanford 2.7% figure are the essay’s citations, kept with [unverified] tags; the release wave itself is corroborated by the English sources above. The essay carries a pro-China framing; that framing is preserved as translated opinion, not asserted as fact. This is translated commentary — not a SigPulse measurement. Our first-party measurements live in the dispatches and the /data/ ledger.

FAQ — Direct Answers

Which Chinese AI models launched in the summer 2026 wave?
Per our translated essay: GLM-5.2 (mid-June, topped CodeArena), Kimi K3 (July 16, 2.8T parameters), DeepSeek-V4-Flash (late July), Qwen3.8-Max (early August, 2.4T-parameter MoE), and ByteDance's Seedance 2.5 (early August). Reuters and The Register independently cover the same August wave.
What share of usage do Chinese models have on OpenRouter?
Per the essay citing OpenRouter data [unverified]: in July 2026 the five most-called models on OpenRouter were all Chinese, and US enterprises' token consumption on Chinese models had risen from 4.5% eighteen months earlier to a stable 30%+, peaking at 46%.
Why are US companies switching to Chinese AI models?
The translated take: cost and openness. The essay cites ~$25 (Claude) vs ~$0.18 (DeepSeek-V4) for the same coding workload, and names Lindy (API costs down ~90%), Coinbase (AI spend nearly halved; 91% of engineers don't need frontier performance), and Tenstorrent (5× cost cut). The pattern it calls 'layered mixing': Chinese models for routine 80% of work, US flagships for the hard 20%.