arXiv, read from inside the Chinese AI conversation
A periodic digest line: batches of new papers, each summarized from its abstract and paired
with how our Chinese-language column reads it — the analogy, the cold water, the trade-off.
Every entry carries dual dates, the covered
arXiv IDs in frontmatter (machine-readable in
/papers.json), and cross-checked
sources. Translated commentary, not measurements — raw markdown per entry at
/papers/<slug>.md; feed at
/papers/rss.xml.
12 digests · 146 papers covered · updated in batches
Batch 5 · 1 digest
Batch 6 · 1 digest
Batch 4 · 1 digest
Batch 3 · 2 digests
- training efficiency Chinese original 2026-09-04 · translated 2026-09-04 · 17 papers
Cheaper, Deeper, Hired: One Day of arXiv on Training Economics, the Math Underneath, and AI's Eight New Jobs
Distillation from a single training example, lower bounds on 'free' acceleration, a sovereign banking model, no-training kidney screening — one arXiv day.
- multi-agent systems Chinese original 2026-09-04 · translated 2026-09-04 · 15 papers
100 AI Researchers, Zero Humans: Cheating and Whistleblowing Emerged Unscripted — the Trust Papers of Sept 4
A 100-agent lab where cheating and whistleblowing emerged unscripted, LLM judges with 0.40 ranking consistency, hook updates as attack surface — Sept 4's arXiv.
Batch 2 · 4 digests
- self-improving agents Chinese original 2026-08-20 · translated 2026-08-30 · 9 papers
What Did 86 arXiv Papers in 48 Hours Say About Where AI Is Heading? Four Signals
Self-improvement fragility, hospitals hiring AI as a formatter, judges without stable values, and everyone wanting to send AI on dates — 86 papers, 4 signals.
- small models Chinese original 2026-08-15 · translated 2026-08-30 · 1 papers
Can a 1B Model Trained Only on Legal Data Compete? Denmark's Mimir v1 Says Yes
One billion parameters, 161 permissible datasets, no copyright gray zones — and state-of-the-art Danish plus near-parity with 4B models on English.
- multi-agent systems Chinese original 2026-08-20 · translated 2026-08-30 · 11 papers
AI Writes Its Own Exam Questions and Grades Its Own Full Marks — Who Audits It? The Next 70 Papers
Hidden multi-agent coordination, self-play environments that train execution but lock strategy, verifiable abstention in sewers and dentistry — the second 70.
- KV cache Chinese original 2026-08-15 · translated 2026-08-30 · 1 papers
What Actually Makes LLM Serving Expensive? VRAM — and vToken Brings OS-Style Virtual Memory to the KV Cache
The KV cache fills your VRAM while the GPU waits; vToken virtualizes reclamation down to single tokens, like paging did for operating systems.
Batch 1 · 3 digests
- agents Chinese original 2026-08-29 · translated 2026-08-29 · 3 papers
Can AI Agents Learn From Experience Now? Wikis, Red Teams, and Failure Mining — the Papers Read Inside China
Agent experience compiled into evolving wikis, red-team agents that learn from attacks, small-model failures tutoring big models — plus Google's ReasoningBank.
- benchmarks Chinese original 2026-08-29 · translated 2026-08-29 · 4 papers
Are Static AI Benchmarks Dying? Digital Cities, Multi-Round Code Review, and AI-Built Test Tracks — the Papers Read Inside China
A digital Hong Kong for agent exams, defect-aware code review, world models as zero-shot simulators — and who owns the exam hall.
- data selection Chinese original 2026-08-29 · translated 2026-08-29 · 4 papers
Do New arXiv Papers Really Show That Less Training Data Beats More? Four August Papers, Read Inside China
SWE-Prime's 10% subsets beating full data, weak models rescuing strong ones, label-free test-time optimization, 2% of attention heads — 'less is more'.