Skip to content
SIGPULSE
China Decoder Economy & Markets Decoder 12 11 min read raw .md ↗

The Meter and the Free Ride: How China's Free AI Started Billing

Synthesized from 4 Chinese originals (2026-08-15 – 2026-08-14) · adapted to English 2026-09-07

Within one week in August 2026, four things happened in the Chinese AI economy. On the evening of August 13, DeepSeek announced that from August 17, 00:00 Beijing time, its API would switch to peak/off-peak pricing — weekday daytime at roughly double the night rate, with cached input repriced upward by up to 12× — while the consumer app and web versions stayed free; the topic drew 9.11 million views on the tech hot-search lists [unverified — platform metrics as relayed by our column]. In the same quarter, Zhipu’s API prices had already risen cumulatively by more than 80 percent, and ByteDance’s Doubao was selling paid consumer tiers at 68, 200 and 500 yuan a month. On the night of August 16→17, as the new prices took effect, the Zhihu question drew 5.86 million views [unverified — same relay], its top-voted comment reading “I thought a big hike meant doubling — I didn’t expect ten times,” while the people actually paying were mostly not posting: they were rewriting their plans overnight. And 48 hours after DeepSeek open-sourced its agent framework on August 13, Tencent QQ announced an official plugin that installs a free, memory-keeping DeepSeek-powered bot into any group chat with a QR-code scan.

Our Chinese-language column ran these as separate dispatches. Read together they are not four stories but one machine: free intelligence stopped being a flat, unlimited condition and acquired a billing structure — a meter on one face, a managed free ride on the other — and this piece takes that machine apart. The first dispatch was also translated in full in watch Issue 1.

The gears

The mechanism has a name imported from the power grid: time-of-use pricing — 峰谷分时, peak-and-valley time-slot billing. For a year before the switch, DeepSeek had sold API access at 20–25 percent of list prices (2到2.5折 — a 4–5× discount against list), and had used it to pull enterprises, agent developers and AI startups onto its platform. The cost of that subsidy showed up as congestion: on weekday daytime, when all of corporate China calls at once, GPU queues and rate-limiting hit the paying customers.

The meter is the reply. Peak windows — weekdays 9:00–12:00 and 14:00–18:00, 7 of 24 hours (29.2% of the day) — are billed at exactly 2× the off-peak rate: V4-Flash output at 9 yuan per million tokens peak against 4.5 off-peak (9/4.5 = 2.0); V4-Pro at 27 against 13.5 (2.0). Demand then sorts itself by elasticity. Loads that can move — batch inference, offline data processing — get half price by running at night. Loads that cannot move — customer-service bots, office assistants, real-time recommendation — sit on the daytime peak, where business hours and price hours fully overlap. And one price line moved not 2× but an order of magnitude: cache-hit input, the stored reuse of intermediate results that had been priced near zero as a developer subsidy, rose by up to 12× (12× = +1,100%; Western coverage measured about 1,114% on its anchor tier).

The machine’s second face is the free-ride channel, and the same week supplied its cleanest example: a free bot inside QQ, the platform where group culture already lives. Free did not disappear when the meter was installed; it was relocated. The meter prices demand that will not move; the free channel is spent where users are acquired — the consumer app, and now the group chat. Both draw on the same GPU pool.

Instance one: the meter switches on

Both sides of August 13 need reporting. The working side: off-peak batch workloads now run at 4.5 yuan per million tokens of Flash output — half the peak price — which for nightly data-processing jobs is a discount, not a hike; V4-Pro, launched the same day with a 1M-token context window, shipped alongside the repricing, so the goods were upgraded and repriced together; and the pro-metering argument inside the developer community is physical — compute genuinely has peaks, time-of-use pricing matches cost to time, and the grid has run this playbook for decades. The failing side: for businesses whose traffic peaks when humans are awake, the same change is a doubling of the largest cost line, with no workaround; cached input, the piece of the bill heavy users had stopped thinking about, rose up to 12×, which is where the “ten times, not two times” reaction came from; and API budgets stopped being forecastable in a way a flat price never was. The question our column logged and could not answer: of “relieving congestion” versus “raising revenue,” what the split is — only DeepSeek holds that number.

Instance two: the walls that followed

The second dispatch framed the repricing as the first of five walls — billing, cost, geopolitics, legal judgment, and the consumer bill — and its load-bearing facts verify. Zhipu had already moved twice in the first quarter: February 12, with the GLM-5 launch, API usage fees rose 67–100 percent (TrendForce), and on March 16 GLM-5-Turbo rose another 20 percent (Yicai) — a cumulative increase Sina Finance put at more than 80 percent for the year to date, and our column relayed, from the CEO’s own figure, as 83 percent. The step percentages and the cumulative number are different bases and do not compose linearly; the anchor is Sina’s cumulative figure. The working column here is striking: Zhipu’s API call volume rose rather than fell through the hikes — supply short of demand — and its stock jumped 13 percent on the March increase [unverified — market reaction as relayed]. The consumer ladder was already built: Kimi memberships since September 2025 at 49/99/199/699 yuan a month (top-to-base 699/49 = 14.3×), Doubao’s June 2026 tiers at 68/200/500 yuan a month (500/68 = 7.35×; annual plans 688/2,048/5,088 yuan, i.e. the top annual tier equals 424 yuan a month against 500 paid monthly — a 15% annual discount).

The other walls, compressed to what is on the record. On July 21, U.S. Treasury Secretary Scott Bessent said officials had found the “watermarks” of American large language models inside Chinese AI systems, called it unacceptable, and said sanctions were on the table (CNBC); the White House later named Moonshot AI in distillation allegations [unverified — relayed via Yahoo Finance]. Around 200 American startups signed a joint letter opposing sanctions on Chinese open models, their own products built on those models [unverified — relayed from our July 28 dispatch]. Inside China, developer communities were discussing domestic open models attaching regional use terms — a community signal, not a published policy, as the column itself flagged. On distillation, the technical record our column pointed to: the technique buys outputs, not weights or source code, and has been standard practice for a decade; training runs take months, so a model released one month after a competitor’s cannot be a copy of it in the arithmetic sense. The failing column of this wall is the wrapper startup: two years of business plans written on the premise “API stays cheap” now repriced overnight, with the exit options being pass-through to users or self-hosting — GPUs, operations and tuning, each a real cost.

Instance three: the overnight rewrites

The third dispatch watched the meter take effect at midnight and noticed who was loud and who was silent. The loudest commenters, by and large, had nothing at stake — the app stayed free, and the most common rumor of the week (“DeepSeek is charging users now”) was false in the direct sense. The people paying were rewriting plans, and the dispatch laid out the three paths with their arithmetic. Path one, shift the hours: any workload that tolerates a night run drops from 9 to 4.5 yuan per million tokens (Flash) or 27 to 13.5 (Pro) — a 50% cut. Path two, downgrade at peak: run Flash at 9 yuan during the day and keep Pro for off-peak complexity — at peak, Pro-to-Flash is 27/9 = 3×. Path three, self-host: the open weights are free to download, so the API bill can be traded for a one-time GPU and operations bill. The working side of this instance: for shiftable workloads the change is genuinely a discount, and path three exists at all only because the weights are open. The failing side: a customer-service bot cannot move to 2 a.m. — its peak is humans being awake; downgrading trades answer quality at exactly the hours quality matters most; and self-hosting’s entry cost, for a three-to-five-person team, was the column’s “not a pivot but a verdict” — its characterization, relayed as such.

Instance four: the free-ride channel

The fourth dispatch is the machine’s other face. On August 13 DeepSeek released DeepSeek Harness (dsh), an MIT-licensed agent framework whose design principle is “everything is a plugin.” The official QQ plugin appeared on GitHub the evening of August 14 [unverified — timing as relayed by ITHome], and on August 15 QQ announced it on Weibo: install, start, scan a QR code; no application form, no keys, no developer certification. Each private chat and each group gets an independent memory that survives restarts; models can be switched mid-conversation with context retained; and the bot answers only when @-mentioned — in a 500-member group, an always-replying bot is unusable, and the @-only mode is the design detail that makes group deployment viable. Two days from framework release to official platform integration is the timing fact; that it reflected coordination rather than coincidence is our column’s reading, relayed as such.

Both columns of this instance. The working side: a group owner — a study group, a project team, a game guild — gets a 24/7 assistant with persistent memory at zero marginal price; for developers, the integration threshold collapsed from server-plus-keys-plus-protocol to a weekend project; and the AI entry point moved from an app you open into a chat you already have. The failing side is the same fact read from the ledger: the channel is free because it sits on the customer-acquisition side of the bill, and the acquisition side draws on the same GPUs the meter now prices by the hour. The dispatch’s closing detail cuts both ways — a bot with permanent per-group memory is a “living archive” of everything the group has ever said: convenient, and total.

The ladder underneath

The backdrop is arithmetic. A year of 20–25%-of-list API pricing (a 4–5× discount) bought share and produced peak-hour queuing for paying customers; the meter is the subsidy’s balance sheet closing. The consumer ladder was built before the API meter turned: Kimi since September 2025 (four tiers, 49 to 699 yuan), Doubao since June 2026 (three tiers, 68 to 500), Zhipu consumer plans from 19 yuan [unverified — relayed from our column’s source list]. Our column’s projection for the consumer side is not a paywall but the sequence video streaming ran: free quotas tightening gradually, advanced features moving into membership, advertising entering the free tier — a projection, on the record as one. The open questions it logged at the end of the week: which vendor follows DeepSeek onto time-of-use pricing — and whether any vendor now moves the other way and picks up the discount banner the price-war leader dropped.

What outsiders usually get wrong

Four corrections, all load-bearing. First: “DeepSeek started charging its users” — false as stated; the consumer app and web versions remained free, and the repricing hit API callers only, a fact stated in the announcement and visible on the official pricing page. Second: “prices rose tenfold” — the output prices rose 2× at peak against off-peak (9/4.5 and 27/13.5, both exactly 2.0); the tenfold experience is real but specific: cache-hit input, repriced from near-free by up to 12× (+1,100%; about 1,114% on Quartz’s anchor tier). Third: “Chinese model APIs keep getting cheaper” — the direction reversed in 2026: Zhipu up more than 80 percent cumulatively by March with call volume rising, Doubao and Kimi operating paid consumer ladders, and DeepSeek, the firm that set the discount benchmark at 20–25% of list, installing the meter itself. Fourth: “peak/off-peak is a surcharge” — it is a restructure that prices time, not only level: for any workload that can run at night, the same tokens cost half (4.5 against 9; 13.5 against 27), and the increase lands only where demand will not move.

Sources

Provenance & disclosure. This piece synthesizes four Chinese-language originals from our WeChat channel — “DeepSeek涨价背后,一个时代结束了” (2026-08-14, also translated in watch Issue 1), “涨价只是开头:大模型的围墙、审判和你的账单” (2026-08-17), “DeepSeek涨价今晨生效:你不用多花一分钱,但有人连夜改了方案” (2026-08-17) and “腾讯官宣:3步把DeepSeek装进QQ,每个群一个独立记忆” (2026-08-15) — drafted with AI assistance under human editorial direction and adapted to English 2026-09-07. Verification: DeepSeek pricing and timing against Reuters, Quartz and the official pricing page; Zhipu’s cumulative increase against Sina Finance with steps against TrendForce and Yicai; Doubao tiers against Securities Times and 21jingji; Kimi tiers against 21jingji; Bessent’s July 21 remarks against CNBC; the QQ integration against ITHome, PingWest and the official plugin repository. The 9.11M hot-search and 5.86M Zhihu view counts, the ~200-startup letter, the Moonshot allegation, the Zhipu stock move, Zhipu’s 19-yuan consumer tier, and the plugin’s GitHub timing are relayed from Chinese reports or our earlier dispatch and marked [unverified]. Ratios recomputed (9/4.5 = 2.0; 27/13.5 = 2.0; 27/9 = 3.0; 12× = +1,100%; 500/68 = 7.35×; 699/49 = 14.3×; 5,088/12 = 424 vs 500 = 15% annual discount; 7/24 = 29.2%; 20–25% of list = 4–5× under). This is reported synthesis — not a SigPulse measurement. Our first-party measurements live in the dispatches and the /data/ ledger.

FAQ — Direct Answers

Did ordinary DeepSeek users have to start paying in August 2026?
No. The consumer app and web chat remained free. The 2026-08-13 announcement, effective 2026-08-17 00:00 Beijing time, repriced only the API — the wholesale interface used by enterprises and developers. The consumer app staying free was stated in the announcement and matches the official pricing page.
What is peak/off-peak (feng-gu) API pricing?
Time-of-use pricing imported from the electric grid. On weekdays, 9:00-12:00 and 14:00-18:00 (7 of 24 hours, 29.2% of the day) are billed at peak rates; all other hours at half. V4-Flash output: 9 yuan per million tokens peak, 4.5 off-peak (9/4.5 = 2x). V4-Pro: 27 and 13.5 (2x). Shiftable batch workloads can run at half price; businesses whose traffic peaks during the day pay double.
Where does the '10x price hike' figure come from?
Cache-hit input, not output. Cached input — reuse of stored intermediate results, previously priced near zero — rose by up to 12x (Chinese coverage) or about 1,114% on Quartz's anchor tier; 12x = +1,100%. Heavy cache users experienced their bills roughly tenfold, which produced the top-voted comment on Zhihu. Output prices rose 2x at peak, not 10x.
What is the QQ bot free-ride channel?
On 2026-08-15, two days after DeepSeek open-sourced its agent framework DeepSeek Harness (dsh, MIT license), Tencent QQ announced an official plugin connecting dsh-powered assistants to private chats and group chats: a three-step setup (install, start, scan a QR code — no keys, no developer certification), independent per-chat and per-group memory that survives restarts, mid-chat model switching, and an @-mention-only response mode. Free to the group owner.