---
title: "Are Static AI Benchmarks Dying? Digital Cities, Multi-Round Code Review, and AI-Built Test Tracks — the Papers Read Inside China"
date: 2026-08-29
originalDate: 2026-08-29
originalTitle: "AI的路考来了：谁出题，谁监考，谁说了算"
issue: "Batch 1"
description: "A digital Hong Kong for agent exams, defect-aware code review, world models as zero-shot simulators — and who owns the exam hall."
tags:
  - "benchmarks"
  - "evaluation"
  - "world models"
  - "agents"
  - "auditability"
papers:
  - id: "2608.27456"
    title: "UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City"
    url: "https://arxiv.org/abs/2608.27456"
  - id: "2608.27442"
    title: "From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench"
    url: "https://arxiv.org/abs/2608.27442"
  - id: "2608.27406"
    title: "CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators"
    url: "https://arxiv.org/abs/2608.27406"
  - id: "2608.27427"
    title: "Persona-Execution Separation for Governed Agent Auditing"
    url: "https://arxiv.org/abs/2608.27427"
faq:
  - q: "What is UrbanGround?"
    a: "Per the abstract (verified against the arXiv page at translation): the first physically constrained sandbox replicating a real city — built from territory-wide 3D geospatial data of Hong Kong — supporting first-person closed-loop agent interaction plus an interactive navigation map. Its headline finding: modern multimodal agents recognize scenes and reason short-range decently, but over long explorations their local abilities 'do not compose into sustained goal-directed behavior' — errors accumulate without correction."
  - q: "Why does benchmark dynamism concentrate power?"
    a: "The Chinese take's argument: rebuilding a digital city or a multi-round review simulation costs far more than shipping a static question bank, so 'who can build road-test sites' shrinks to a few institutions — unless the environments are open-sourced, in which case openness itself becomes the power question. Simulation has always worked this way: flight simulators, driving schools, war games — whoever controls the environment defines 'qualified.'"
  - q: "What is persona-execution separation?"
    a: "An architecture where an agent's persona and its execution live in different trust domains, bridged by governed contracts, with execution fully auditable. The take reads it as the second system growing alongside evaluation: once exams get realistic, a supervision layer becomes procurement language."
---
An exam is changing its format. The old AI test was a written exam: fixed question banks, single-turn answers, passable by memorization — and once the bank leaks, everyone scores 100 and the score means nothing. A late-August arXiv batch shows the exam hall moving wholesale to the road test. Our Chinese-language column's diagnosis ran under "AI's road test is here: who writes the questions, who proctors, who decides." This digest translates it; paper claims follow their abstracts (UrbanGround verified line-by-line at translation).

## The claims

- **UrbanGround (2608.27456):** a physically constrained Hong Kong replica built from citywide 3D geospatial data; first-person closed-loop interaction; evaluation across scene grounding, navigation, and robustness. Finding: agents' local skills don't compose into sustained goal-directed behavior over long horizons.
- **MCR-Bench (2608.27442):** the first defect-state-aware benchmark — code review restored to what it actually is, a multi-round interactive dynamic decision process between author and reviewer, not a one-shot "find the bug."
- **CLAP (2608.27406):** a cross-embodiment, action-conditioned video world model trained on internet-scale human-and-robot video that works as a zero-shot physical simulator — the exam venue itself can now be model-generated, at least for physical-action tasks.
- **PES (2608.27427):** persona and execution in separate trust domains, bridged by governed contracts, execution auditable end-to-end — governance as architecture rather than policy layer.

## How the Chinese column reads it

First result: scores regain signal value — water squeezed out. Then three ripples, each with the take's boundary condition. **A rising barrier to entry** — road-test sites are expensive, so evaluation capability concentrates into few hands; open-sourcing the sandboxes partially offsets the concentration, which makes openness itself a power question. **Audit becomes the new necessity** — PES-style auditable execution is the second system growing next to the exam hall. And the subtlest ripple: **the exam hall starts being built by AI** — CLAP-style world models serve as both training ground and test track, and when the builder's engine and the examinee's engine share a substrate, "independence" needs re-pouring.

The historical slot the take assigns this to: whoever controls the simulation defines "qualified" — driving schools, flight sims, war games; AI just joined an old list. Hence its re-aiming of the buyer's question: not "what did it score" but "where was it tested, and who owns the exam hall."

## Cold water

The take's own skeptic, quoted straight: don't romanticize evaluation tech into a "game of thrones" — a sandbox is still a sandbox; however fine the digital Hong Kong, it cannot test long-tail risk in real society. The criticism is admitted, then turned: that gap *is* the system's next missing piece.

## What to watch

Whether dynamic-evaluation cost forces a two-tier market (an honest signal for those who can pay, static comfort for the rest); whether "auditable execution" clauses migrate from papers to procurement documents; and the arrival of the first major benchmark built by a world model rather than by humans — the moment the referee's and racer's engines provably share parts.

## Sources

- [UrbanGround (arXiv:2608.27456)](https://arxiv.org/abs/2608.27456)
- [MCR-Bench (arXiv:2608.27442)](https://arxiv.org/abs/2608.27442)
- [CLAP (arXiv:2608.27406)](https://arxiv.org/abs/2608.27406)
- [Persona-Execution Separation (arXiv:2608.27427)](https://arxiv.org/abs/2608.27427)

> **Provenance & disclosure.** Originally published in Chinese on our WeChat channel on 2026-08-29 ("AI的路考来了：谁出题，谁监考，谁说了算"); drafted with AI assistance under human editorial direction. Translated to English on 2026-08-29 (AI-assisted, human-reviewed). UrbanGround's claims were verified against its arXiv abstract at translation; the other three follow the abstracts as listed by the column (paper titles shortened where the column's citation was informal). This is translated commentary — not a SigPulse measurement. Our first-party measurements live in the [dispatches](/posts/) and the [/data/ ledger](/data/).
