---
title: "What Can't Be Tokenized? 31 arXiv Papers on Foundations and Frontiers"
date: 2026-09-12
originalDate: 2026-09-12
originalTitle: "从鞅到张量：16篇统计学习理论新论文 / 该检索还是该路由？知识栈重构的6篇论文 / 语音补课、拓扑开路：多模态的9道新题"
issue: "Batch 6"
description: "An eleven-year conjecture falls, attention gets statistical-mechanical phase diagrams, knowledge placement becomes token economics. 31 papers."
tags:
  - "learning theory"
  - "knowledge management"
  - "RAG"
  - "model routing"
  - "multimodal"
  - "arXiv digest"
papers:
  - id: "2609.11918"
    title: "General Quantification of Covariate and Concept Shifts"
    url: "https://arxiv.org/abs/2609.11918"
  - id: "2609.11845"
    title: "$β$-Skewed Maximal Spanning Forests"
    url: "https://arxiv.org/abs/2609.11845"
  - id: "2609.11807"
    title: "Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead"
    url: "https://arxiv.org/abs/2609.11807"
  - id: "2609.11795"
    title: "Small-Ball Marginals Do Not Control Restricted Eigenvalues by Euclidean Gaussian Width"
    url: "https://arxiv.org/abs/2609.11795"
  - id: "2609.11740"
    title: "From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization"
    url: "https://arxiv.org/abs/2609.11740"
  - id: "2609.11712"
    title: "Generalization Analysis of Distributed Kernel-based Robust Gradient Descent Algorithms"
    url: "https://arxiv.org/abs/2609.11712"
  - id: "2609.11606"
    title: "Identifiability of Nonnegative Tensor Decompositions via Positive Scattering"
    url: "https://arxiv.org/abs/2609.11606"
  - id: "2609.11557"
    title: "Martingale central limit theorems in $p$-Wasserstein distance"
    url: "https://arxiv.org/abs/2609.11557"
  - id: "2609.10976"
    title: "Phases in a class of associative memories via hidden neurons"
    url: "https://arxiv.org/abs/2609.10976"
  - id: "2609.10928"
    title: "AUC Maximization from Biased Positive-unlabeled Data with Confidence"
    url: "https://arxiv.org/abs/2609.10928"
  - id: "2609.10904"
    title: "Agnostic Model-Assisted Estimation with Machine Learning for Survey Data"
    url: "https://arxiv.org/abs/2609.10904"
  - id: "2609.10886"
    title: "Relatively Smart II: Tractable or Semi-Supervised Instance-Optimal Learning"
    url: "https://arxiv.org/abs/2609.10886"
  - id: "2609.10879"
    title: "Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry"
    url: "https://arxiv.org/abs/2609.10879"
  - id: "2609.10767"
    title: "Weighted Empirical Risk Minimization for Machine Learning under Long-Range Dependence: Exact Pathwise Rates and Learning-Error Geometry"
    url: "https://arxiv.org/abs/2609.10767"
  - id: "2609.10729"
    title: "A Quantum-Inspired Dequantization Method for Diagonally Weighted Matrix Functions: Application to Learning with Optimized Random Features"
    url: "https://arxiv.org/abs/2609.10729"
  - id: "2609.10534"
    title: "Likelihood-free inference with nuisance parameters through normalizing flows"
    url: "https://arxiv.org/abs/2609.10534"
  - id: "2609.11859"
    title: "From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge"
    url: "https://arxiv.org/abs/2609.11859"
  - id: "2609.11656"
    title: "Learnware and AI Model Management System"
    url: "https://arxiv.org/abs/2609.11656"
  - id: "2609.11569"
    title: "Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)"
    url: "https://arxiv.org/abs/2609.11569"
  - id: "2609.11414"
    title: "SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations"
    url: "https://arxiv.org/abs/2609.11414"
  - id: "2609.11393"
    title: "Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning"
    url: "https://arxiv.org/abs/2609.11393"
  - id: "2609.11390"
    title: "VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents"
    url: "https://arxiv.org/abs/2609.11390"
  - id: "2609.11900"
    title: "MindTopo: Can Foundation Models Reason in Topological Space?"
    url: "https://arxiv.org/abs/2609.11900"
  - id: "2609.11899"
    title: "Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding"
    url: "https://arxiv.org/abs/2609.11899"
  - id: "2609.11892"
    title: "Nuha-Speech: Building General-Purpose Arabic Speech-LLMs"
    url: "https://arxiv.org/abs/2609.11892"
  - id: "2609.11864"
    title: "RetroThinker: Enabling Retrospective Thinking in Speech LLMs"
    url: "https://arxiv.org/abs/2609.11864"
  - id: "2609.11772"
    title: "Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural Understanding"
    url: "https://arxiv.org/abs/2609.11772"
  - id: "2609.11762"
    title: "Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs"
    url: "https://arxiv.org/abs/2609.11762"
  - id: "2609.11708"
    title: "Language-Augmented Semantic Priors for B-Spline Surface Fitting"
    url: "https://arxiv.org/abs/2609.11708"
  - id: "2609.11499"
    title: "Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs"
    url: "https://arxiv.org/abs/2609.11499"
  - id: "2609.11412"
    title: "X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation"
    url: "https://arxiv.org/abs/2609.11412"
faq:
  - q: "What eleven-year-old conjecture fell?"
    a: "At COLT 2015, Banerjee, Chen and Sivakumar asked whether a uniform small-ball condition on the rows of a random design matrix forces a restricted-eigenvalue lower bound governed by the Euclidean Gaussian width. Small-Ball Marginals Do Not Control Restricted Eigenvalues (arXiv:2609.11795) answers the natural distribution-free formulation in the negative — for every sample size it constructs counterexamples. In theory's value system, refuting a conjecture is worth as much as proving one: it deletes the proof routes the field might otherwise burn a decade on."
  - q: "What does attention have to do with phase transitions?"
    a: "Phases in a class of associative memories via hidden neurons (arXiv:2609.10976) studies Hopfield-style associative memory whose higher-order and exponential extensions turn the retrieval update into softmax attention — the same operation a transformer performs. Using hidden neurons as the order parameter, it unifies the polynomial and exponential capacity regimes in one statistical-mechanical analysis, with replica-symmetric phase diagrams. In 2026, attention has its own phase-transition maps."
  - q: "Where does a model's knowledge actually live?"
    a: "From Parameters to Answers (arXiv:2609.11859) intervenes on hidden states layer by layer across Qwen, Llama and Gemma families, measuring how much the answer depends on query-routing information versus target knowledge at each depth — a mechanistic map of internal retrieval with a stable layered structure. The practical echo runs through the whole batch: knowledge placement — parameters, document RAG, knowledge graphs, or routing between models — is increasingly a cost question, measured in tokens per point of accuracy, not an architectural creed."
  - q: "Why test reasoning in topological space?"
    a: "MindTopo (arXiv:2609.11900) asks whether foundation models can reason about continuity, separation, order, enclosure and knots — properties invariant under continuous deformation that cognitive science treats as the foundation of spatial understanding. Metrics and shapes can be tokenized; topological invariants largely cannot. Baseline performance is weak, but the value is making the failure explicit — and separating 'cannot see it' from 'cannot reason about it' is exactly what such benchmarks must now do."
---
Third digest of the week's 81-paper series: 31 papers on foundations and frontiers — statistical learning theory (16, the largest single class of the batch), the LLM knowledge stack (6), and multimodal research splitting into two speeds (9). The theory block is the one place this series runs against its own news rhythm: no paper here chases a hot model, and each one moves a boundary that will still matter when this cycle's models are retired. (All papers are arXiv preprints, not peer-reviewed; claims follow their abstracts.)

## Theory's hard answers

The window this batch opens: old questions getting definitive answers. The headline refutation — a COLT 2015 question on small-ball conditions and restricted eigenvalues, closed in the negative after eleven years (arXiv:2609.11795) — is worth exactly as much as a new theorem, because it deletes the proof routes the field might have burned (see FAQ). Around it, a cluster of long-gestating problems land: martingale central limit theorems extended to p-Wasserstein distances and ℓr norms (2609.11557), nonnegative tensor decompositions made identifiable via positive scattering — nonnegativity itself contributing structure beyond dimension and independence (2609.11606), and a general quantification of covariate and concept shifts that unifies both error bounds under entropic optimal transport as γ*-shifts, with estimators that concentrate — moving learning bounds from idealized to sample-estimable (2609.11918).

The rest of the toolbox: Relatively Smart II (2609.10886) competes against every error guarantee certifiable from unlabeled data — certification itself as a learning objective; orthogonal multi-index models analyzed beyond small initialization, with incremental learning and neuron competition (2609.10879); exact pathwise rates for weighted ERM under long-range dependence (2609.10767); latent-factor inference from good starts with implicit regularization (2609.11740); robust kernel-gradient-descent generalization for the distributed setting (2609.11712); AUC maximization from biased positive-unlabeled data (2609.10928); model-assisted survey estimation, method-agnostic by design (2609.10904); a quantum-inspired dequantization for diagonally weighted matrix-function learning (2609.10729); near-optimal RL with multi-step transition lookahead (2609.11807); and the batch's most cinematic piece — associative-memory phases unified through hidden neurons (2609.10976), where the retrieval update is softmax attention and the capacity limits come back as phase diagrams (see FAQ). The column's honest caveat: theory's correct use is not "deployable next week" but "the map you find when an engineering line hits a wall two years from now."

## The knowledge stack becomes an accounting question

Six papers cover the whole knowledge chain, and the shared conclusion is economic: where knowledge lives is now a cost question. Inside parameters, From Parameters to Answers (arXiv:2609.11859) measures the layer-by-layer dependency of answers on routing information versus target knowledge (see FAQ). Outside, VikingRAG (2609.11390) attacks the token bill of structure-aware retrieval — the competition is no longer "who retrieves more accurately" but "who pays fewer tokens per point of accuracy" — while EXYGEN (2609.11569) scales knowledge-graph dialogue via text-to-SPARQL from automatically built schemas, the higher-ceiling, higher-cost route. At decision time, SWRouter (2609.11414) catches a failure mode the routing literature had skipped: a router that performs well on single-turn queries degrades as conversation context drifts — its answer to "which model fits best" changes with the dialogue, so the router must track it online. Beyond Confidence (2609.11393) squeezes test-time adaptation with stability signals past predictive entropy. And Learnware (2609.11656) argues model pools are still just storage systems — the field needs genuine model management, the way files needed DBMSs: retrievable, composable, traceable.

## Multimodal at two speeds

Speech and vision are moving at visibly different velocities. Speech LLMs are in the harvesting phase: Nuha-Speech (arXiv:2609.11892) lays down the full Arabic stack — 1.5M+ speech-QA samples, training recipe, evaluation — the data-moat playbook that worked for Chinese and other mid-size languages; X-AuT (2609.11412) compresses audio encoders progressively with cross-scale distillation after finding that naive block deletion corrupts downstream embeddings; component-aware differential privacy (2609.11762) discovers that acoustic-encoder and LM-layer gradients have heterogeneous sensitivity, so per-parameter clipping budgets fail and must be split by component; RetroThinker (2609.11864) asks whether a streaming speech model can revise reasoning it has already spoken aloud — accuracy-versus-latency made an adjustable spectrum; and a Whisper-based multilingual video transcription study (2609.11772) shows the tooling's maturity while the column keeps its caveat: subtitles are not understanding.

The visual side is still breaking ground on what tokenization cannot buy. MindTopo (arXiv:2609.11900) tests topological reasoning — with weak results that matter precisely because they are explicit (see FAQ). Language-augmented B-spline fitting (2609.11708) replaces hand-tuned geometric initialization with semantic priors — the language-geometry interface deepening from describing pictures to constraining shapes. Recursive Code World Models (2609.11499) rebuilds complex 3D worlds from a single reference image via recursively generated scene programs — world models shifting from predicting the next frame to writing the world as executable code. And Caption-once, Frames-on-Demand (2609.11899) manages budget-constrained long-video understanding by keeping temporal memory as one offline caption pass and recalling frames from the cloud on demand — language as the proxy for visual memory, exchanged back when needed.

## The takeaway

Foundations had the batch's biggest single class, and its health signal is uniform: a field confident enough to close its own conjectures, price its own knowledge, and name what its tokenizers cannot reach. The two frontiers to watch from here: which architecture first admits the topological/geometric priors that don't tokenize, and whether model management grows from vision to substrate as the model count keeps compounding.

> **Provenance & disclosure.** Originally published in Chinese on our WeChat channel on 2026-09-12 as three parts of a nine-part "arXiv 81" series ("从鞅到张量：16篇统计学习理论新论文"; "该检索还是该路由？知识栈重构的6篇论文"; "语音补课、拓扑开路：多模态的9道新题"); drafted with AI assistance under human editorial direction. Translated and consolidated to English on 2026-09-12 (AI-assisted, human-reviewed). All 31 cited arXiv IDs were resolved via the arXiv API on 2026-09-12 and titles matched the claims; the COLT 2015 attribution, the entropic-optimal-transport/γ* mechanism, and the softmax-attention retrieval update were verified against the abstracts; the 1.5M+ Arabic-sample figure follows the abstract as cited. Preprint caveat inherited. Translated commentary — not a SigPulse measurement; first-party numbers live in the [dispatches](/posts/) and the [/data/ ledger](/data/).
