Skip to content
SIGPULSE
China AI Watch Issue 4 3 min read raw .md ↗

Who Is Making Money With Your Voice? The ¥5 Cloning Black Market. The Take Inside China

Chinese original 2026-08-18 · 「比AI换脸更隐蔽:谁在偷你的声音赚钱?5块钱就能克隆」 · translated to English 2026-08-30

Ji Guanlin — the voice of Zhen Huan in Empresses in the Palace and Judy in Zootopia’s Chinese dub — recently heard “her own” voice performing dialogue in a production she never worked on. Our WeChat column’s piece (2026-08-18, “More hidden than AI face-swaps: who is making money with your voice? ¥5 buys a clone”) opens there and descends into the supply chain underneath. This entry translates it, with the pricing and the court record pinned to sources.

The numbers

  • Black-market pricing (Beijing News undercover, 2026-03): cloning software ¥1; generation services ¥5–100; sample requirement 3–15 seconds
  • Reporter’s self-test: 3 minutes of training → a model of her own voice she could not distinguish from herself
  • Voice actor Li Longbin, on hand-tuned clones: similarity above 90% — “sometimes even I have to listen carefully”
  • Landmark verdict: ¥250,000 damages, Beijing Internet Court, 2024-04-23 (Civil Code Art. 1023)
  • Enforcement ledger, per the take: ~1 year for voice actor Ye Qing’s lawyer to identify a first infringer; voice actor Shen Anyu, three years and counting [unverified — as reported by China Youth Daily, June]

The take inside China

The market teardown. The chain is short: raw material (any voice ever left online — short-video clips, streams, dubbed drama) → ¥1 software or ¥5 services → output monetized in self-media narration, audiobooks, ads. The reporter’s three-minute self-clone is the piece’s exhibit A that “AI voices are obviously fake” is a retired heuristic. Gao Xiaosong has issued two statements this year denying he operates or authorized any AI-synthesized voice.

The splicing dodge. The black market’s real shield, per voice actor Li Longbin’s wording (“after hand-tuning”): infringers avoid 1:1 copies, splicing several people’s voiceprints into something highly similar yet not identical — similar enough that listeners default to you, dissimilar enough to survive a lawsuit.

The industry’s biggest collective pushback. Ji Guanlin’s March statement against unauthorized voice harvesting drew reposts from dozens of leading voice actors (Bian Jiang, Zhang Jie and others); 729 Voice Studio’s roster issued statements through the spring; in August, voice actor Sanshi publicly denied voicing an AI comic-drama ad and demanded takedown — was mocked by the producer as “chasing clout”, with public opinion overwhelmingly on his side and the ad still running. By mid-August, coverage’s headline had escalated to “AI dubbing infringement storm; multiple parties call for legislation.”

The asymmetry ledger. The law exists (Art. 1023; a winning ¥250K precedent) but reads as one side’s pricing problem: ¥5 and 3 seconds to steal; a year-plus, self-funded forensics and uncertain recognition standards to chase. The take’s summary line: the spread between infringement cost and enforcement cost is the black market’s profit margin — and the people who spent a lifetime perfecting a craft voice turn out to be its least-protected practitioners.

What to do. Individuals: guard your sample — no custom voice packs in unknown apps, no read-aloud “voice tests”. The structural prescriptions from lawyer Zhang Yanlai: watermarks and metadata in generated audio, platform-side filtering, professional voiceprint registration. And the industry’s long hedge: when fake voices flood, verifiably human performance becomes the premium label — if the label can be kept honest.

What the Chinese take left out

Platform-side numbers (how many infringing listings the e-commerce sites actually removed after the undercover report), and any adjudicated application of the 2024 precedent to a hand-tuned spliced voice — the exact dodge the piece identifies has, so far, no cited test case.

Why it matters outside

Voice is the biometric that everyone publishes for free — every voice message and meeting recording is training material at 3–15 seconds per identity. The Chinese case gives English readers the full arc in one market: pricing collapse (¥1/¥5), a precedent (¥250K under portrait-rights-by-reference), an enforcement-cost wall, and the recognizability-standards gap that lets tuned clones through. For anyone designing voice-authentication or consent systems, that ledger is the requirements document.

Sources

Provenance & disclosure. Originally published in Chinese on our WeChat channel on 2026-08-18 (“比AI换脸更隐蔽:谁在偷你的声音赚钱?5块钱就能克隆”); drafted with AI assistance under human editorial direction. Translated to English on 2026-08-30 (AI-assisted, human-reviewed). Black-market pricing and the reporter’s self-test were checked against Beijing News’ undercover report; the 2024 verdict against Cailianshe’s coverage and Art. 1023 against standard legal commentary; individual enforcement timelines remain as-cited [unverified]. Names romanized. This is translated commentary — not a SigPulse measurement. Our first-party measurements live in the dispatches and the /data/ ledger.

FAQ — Direct Answers

How cheap has voice cloning become in China?
Per Beijing News' March 2026 undercover reporting: cloning software sells for ¥1, per-use generation services run ¥5–100, and 3–15 seconds of audio suffices for a one-to-one replica — the reporter trained a model on her own recording for three minutes and could no longer distinguish the output from her own voice. Independent testing relayed in coverage put per-50-character cloning under ¥1 with 3-second samples reaching 98%+ similarity.
What did the landmark court case decide?
On April 23, 2024, the Beijing Internet Court ruled in China's first AI-generated voice personality-rights case: a voice actor's recordings had been used without consent to build a text-to-speech product. Under Civil Code Article 1023 (voice protected by reference to portrait rights), the court awarded ¥250,000 and set the standard — AI-generated voices with recognizability to the general or relevant public fall within personality-rights protection. Copyright in recordings does not automatically confer voice rights.
Why is enforcement still so slow?
The economics, per the take: stealing a voice costs ¥5 and 3 seconds of sample; pursuing one infringer took voice actor Ye Qing's lawyer roughly a year just to identify the first perpetrator, followed by self-funded voiceprint forensics and uncertain case acceptance. There is also a standards gap — some precedents accept recognizability to the 'relevant public', others demand the 'general public', leaving non-celebrity voice actors with high-similarity but legally unrecognizable voices.
How do you protect your own voice?
The practical rules from the take: don't record 'custom voice packs' in unfamiliar apps, and don't join 'read this passage to test your voice' campaigns — you don't know whose training set your 15 seconds feed. The lawyers' structural proposals run further: mandatory watermarks/metadata in AI-generated audio, platform-side deepfake filtering, and a voiceprint registration system for professionals.