<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>SigPulse — Unfiltered Signals from China’s AI, Hardware &amp; Digital Frontier</title><description>Ground-truth intelligence on China’s AI models, local compute economics, and digital reality — verified on bare-metal hardware, published with reproducible configs.</description><link>https://sigpulse.com/</link><language>en</language><item><title>The Staircase: Four Steps for Every Job in the Plant, and the Top Step Stays Empty</title><link>https://sigpulse.com/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/</guid><description>Every plant job on a four-step staircase — dashboards, scripts, a model call, agents. The top step stays empty; the agent purchase case collapses.</description><pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt; ended with a bill and a question. The bill: the plant&amp;#39;s only AI call, rebuilt the same day to a twenty-fifth of its price. The question: if one agent invocation cost twenty-five times its content and used none of its tools, what is an agent actually &lt;em&gt;for&lt;/em&gt; — and what, exactly, would this plant be buying if it bought one?&lt;/p&gt;
&lt;p&gt;The question was not academic. The operator&amp;#39;s company was weighing an industrial-grade agent platform for the plant. So the same afternoon the audit closed, the operator built the answer: a staircase, with every job in the plant standing on it. This episode is that record — the staircase, the law the plant&amp;#39;s own history had already been obeying, the purchase case taken apart screw by screw, and a ruling that arrived on schedule and chose to wait.&lt;/p&gt;
&lt;h2&gt;Four steps, one rule&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Real jobs from this plant&lt;/th&gt;
&lt;th&gt;Who decides the next step&lt;/th&gt;
&lt;th&gt;Marginal cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;1 · Fixed pages&lt;/td&gt;
&lt;td&gt;dashboards, query and reconciliation windows&lt;/td&gt;
&lt;td&gt;the device dashboard, the data-convergence window, the line-side raw-statistics view&lt;/td&gt;
&lt;td&gt;designed in advance; a human clicks&lt;/td&gt;
&lt;td&gt;free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 · Fixed scripts&lt;/td&gt;
&lt;td&gt;scheduled Python that runs itself&lt;/td&gt;
&lt;td&gt;the three patrol shifts, the changeover daily, incremental sync, PDF rendering and delivery&lt;/td&gt;
&lt;td&gt;the code, written once&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 tokens&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 · Script + one model call&lt;/td&gt;
&lt;td&gt;the machine phones a scribe&lt;/td&gt;
&lt;td&gt;the patrol running summary: facts read out, prose read back, hang up&lt;/td&gt;
&lt;td&gt;the code; the model only writes&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~2K tokens per call&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 · Agents&lt;/td&gt;
&lt;td&gt;the model decides steps and picks tools&lt;/td&gt;
&lt;td&gt;the drafting sessions behind this record; one-off unattended batch runs&lt;/td&gt;
&lt;td&gt;the model, at run time&lt;/td&gt;
&lt;td&gt;door fee, tens of thousands of tokens per entry&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The one rule: &lt;strong&gt;stand as low as the job allows.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The lower two steps carry no AI at all, and that is not aspiration — it is the audited state of the plant. &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt;&amp;#39;s forensic sweep of the whole patrol chain found exactly one model call point in thirty days of running: the running summary, on Step 3. Judgment, classification, rendering, delivery — all Step 2, all zero tokens. Most of the plant&amp;#39;s work never climbs past the second step, which is why the top step, for this plant, stays empty.&lt;/p&gt;
&lt;h2&gt;The scaffolding law&lt;/h2&gt;
&lt;p&gt;The staircase was not a new theory bolted onto the plant; it was a description of what the plant&amp;#39;s history already did. Three times, the same pattern:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A threshold, found by AI, written into code.&lt;/strong&gt; During an AI session, a pattern emerged in the event-window envelope: data dropout under 20% marks stretches where two faces change together, while a single-face cover sits near 28% and never qualifies. Confirmed by the operator against 36 days of data — four qualifying stretches, each with an account, zero misses — it was written into the monitor the same morning. The insight was AI&amp;#39;s; the rule is now Step 2&amp;#39;s.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A research result, promoted into the daily.&lt;/strong&gt; The overnight consistency study had produced conclusions humans had to go read. On 09-07 its verdict moved into the morning report&amp;#39;s fixed section; the next day brought the exemption for covered faces. Step 4&amp;#39;s output, re-homed on Step 2.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The summary itself, born on Step 4, climbed down.&lt;/strong&gt; The lane entered the world on 09-05 through a full agent harness — seven days, 44 calls, 2,239,178 tokens — and on 09-12 was rebuilt into a single direct call at a twenty-fifth of the price. &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt; told it as a bill. On the staircase it is something else: scaffolding, taken down the week the building was finished.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The law, as the operator wrote it that day: an agent is scaffolding. It is hired for the part of the job that is still unfinished — the exploring, the deciding-what-to-do-next. When the job is understood, it becomes a fixed step, and the scaffolding comes down. Which reframes the purchase under consideration: &lt;strong&gt;a resident agent platform in a plant of finished jobs is scaffolding rented forever on a building already topped out.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Retiring the word &amp;quot;need&amp;quot;&lt;/h2&gt;
&lt;p&gt;The first draft of the operator&amp;#39;s same-day memo listed four classes of work that &amp;quot;need&amp;quot; an agent. The operator struck the word himself. Debugging runs through a coding agent today — but ran for decades without one, on grep, logs and patience. An agent creates no capability that a human with fixed tools lacks; it saves labor and time. &lt;strong&gt;An agent is an elevator, not the electricity.&lt;/strong&gt; You can always take the stairs.&lt;/p&gt;
&lt;p&gt;The live question becomes: which jobs are worth &lt;em&gt;hiring a temp for&lt;/em&gt;? Three tests, all three at once — the work is new enough that no instruction card can be written for it; the volume or urgency clears the entrance fee; and a guardrail exists, so the temp&amp;#39;s mistakes are catchable. Applied to this plant, agent usage survives in exactly three forms, none of them on the production floor: the drafting sessions behind this very record (no instruction card exists for &amp;quot;write the next episode&amp;quot;); the engineer&amp;#39;s chair, where an agent turns a three-hour search into twenty minutes — an accelerator, kept optional; and one-off unattended batch runs, hired for a night, cage and all. The plant&amp;#39;s runtime has no seat for one, and not for lack of imagination: runtime work is SOPs, and &lt;strong&gt;in a world of standard operating procedures, improvisation is not a skill — it is a violation.&lt;/strong&gt; The agent&amp;#39;s one talent, deciding the next step on the spot, is the one talent the floor forbids.&lt;/p&gt;
&lt;h2&gt;What a purchase would actually buy&lt;/h2&gt;
&lt;p&gt;Take an industrial agent platform apart and four components fall out:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Nature&lt;/th&gt;
&lt;th&gt;This plant&amp;#39;s standing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Model capability&lt;/td&gt;
&lt;td&gt;a commodity — swap by API&lt;/td&gt;
&lt;td&gt;dual-lane already in production: a frontier cloud model and an on-prem 27B, same manual, interchangeable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SOP and domain-knowledge capture&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;the expensive part&lt;/strong&gt; — and no vendor can supply it; the domain must feed it&lt;/td&gt;
&lt;td&gt;a manual library (four books) plus the judgment constitution, on a measured iterate-and-retest cycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduling skeleton&lt;/td&gt;
&lt;td&gt;cron, retries, queues — cheap&lt;/td&gt;
&lt;td&gt;patrol timers, comparison harness, unattended shifts: all running&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The autonomy loop&lt;/td&gt;
&lt;td&gt;the only component a purchase &lt;em&gt;adds&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;priced negative by the plant&amp;#39;s own audit: door fee, variance, unauditable steps&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The accounting is the uncomfortable part for the purchase case: the cost of &amp;quot;industrial AI&amp;quot; was never in component four — it is in components two and three, which a platform does not remove and cannot supply. The buyer pays for the one component the audit priced as a liability.&lt;/p&gt;
&lt;p&gt;And a caution about costumes: a good deal of what is sold as &amp;quot;industrial agents&amp;quot; is Step 2 wearing a badge — if the steps are written in advance, it is a workflow, whatever the box says. There is one question that strips the costume off any demo: &lt;strong&gt;at run time, who decides the next step — the code, or the model?&lt;/strong&gt; If the purchase proceeds anyway, the record&amp;#39;s four tests travel with it: fully local or out; SOPs exportable as files; replayable, deterministic audit; and the model replaceable by API — this plant&amp;#39;s one-manual-two-models lane is the fourth test passed in production.&lt;/p&gt;
&lt;h2&gt;The ruling that chose to wait&lt;/h2&gt;
&lt;p&gt;One contest inside Step 3 remained genuinely open: cloud-direct or on-prem for the summary lane. The comparison had run each morning from 09-11 — the same sanitized fact sheet to both lanes, only the model and the manual differing. Its window was to close with a ruling on 09-13.&lt;/p&gt;
&lt;p&gt;The ruling came, and it was: &lt;strong&gt;not either-or.&lt;/strong&gt; Both lanes stay; the window extends to 09-20. The operator&amp;#39;s reason, paraphrased from the session record: the comparison costs almost nothing, so let it keep running. The numbers agree with him — twelve windows across three mornings:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mornings 09-11→09-13&lt;/th&gt;
&lt;th&gt;Cloud lane (stress recipe: earlier model, bare, 8,192 pad)&lt;/th&gt;
&lt;th&gt;On-prem 27B (manual mounted)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Windows completed with prose&lt;/td&gt;
&lt;td&gt;10 of 12&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;12 of 12&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blanks&lt;/td&gt;
&lt;td&gt;2 — both at exactly the output cap&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wall time, completed windows&lt;/td&gt;
&lt;td&gt;20–27 s&lt;/td&gt;
&lt;td&gt;124–171 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens, both lanes per day&lt;/td&gt;
&lt;td&gt;13,196 → 18,089 → 18,522 (local share: 7,389 / 7,675 / 8,026, zero API bill)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Two footnotes keep that table honest. First, the blanks: the stress recipe is not the production recipe — an earlier-generation model, no manual, a double-size pad — and both blanks stopped at the cap exactly, the same exam physics &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;Episode 3&lt;/a&gt; established and &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt; met again. They indict a recipe, not a lane. Second, the daily total: it is &lt;em&gt;both&lt;/em&gt; lanes combined, and the blank windows were its largest single items at 8,192 apiece — which is precisely why the operator could afford patience. Evidence at that price is a bargain; a rushed decision is not.&lt;/p&gt;
&lt;p&gt;The extension itself was a one-line change — the comparison script&amp;#39;s exit date moved from 09-13 to 09-20, syntax-checked, effective the next morning. The decision was not cancelled. It was re-priced.&lt;/p&gt;
&lt;h2&gt;First live round, receipt and all&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt; closed waiting for the rebuilt lane&amp;#39;s first live event shift. It came at 20:55 that same evening: an event verdict on both lines, the gate fired, and both summaries came back — the service journal logged both lanes&amp;#39; AI calls complete at 20:59:22, the whole chain from trip generation to alert, PDFs and delivery inside 3 minutes 35 seconds. The samples read as the manual demands: bold verdict first, plant terms only, every sentence carrying its source field. The next morning&amp;#39;s 06:30 patrol came back green on both lines, and the lane made zero calls while the reports rendered on schedule — the control arm, clean.&lt;/p&gt;
&lt;p&gt;One line is missing from the receipt, and the record says so: the production path prints token usage only on failure, so the first live round has no token-level bill — its cost anchor remains the A/B measurement (~2K per call). A one-line change would print usage on success too; it touches the production script, so it waits for approval rather than shipping quietly.&lt;/p&gt;
&lt;h2&gt;What the staircase did not cover&lt;/h2&gt;
&lt;p&gt;The audit had now reached every job in the plant: the chain, the lane, the manuals, the comparison, the purchase case. Every seat had been priced except one — the desk the audits were run from, where AI is used the way water is used, all day, in the open. That desk is &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the next episode&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the operator&amp;#39;s same-day memo of 2026-09-12 (the staircase, the temp-hiring tests, the four-part purchase dissection — including the operator&amp;#39;s own strike of the word &amp;quot;need&amp;quot;), the three comparison reports of 2026-09-11→09-13 (per-window tokens and wall times verbatim; the early log note that misnamed one blank line is corrected here against the reports), the ruling of 2026-09-13 as landed in the comparison script&amp;#39;s window line, the service journal entry of 2026-09-12 20:59:22 and the next morning&amp;#39;s green-shift report, and the re-addition that settles Episode 4&amp;#39;s 500-token discrepancy at 1,035,772. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records; derived figures marked as computed. Deliberately absent, per the series&amp;#39; disclosure policy: anything identifying the plant, its industry, its operator or its network, the identities of the harness and model products, and internal identifiers. The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. The map: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;the anchor&lt;/a&gt; · Episode 1: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;the dress rehearsal&lt;/a&gt; · Episode 2: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;the second without AI&lt;/a&gt; · Episode 3: &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;the blank-paper exam&lt;/a&gt; · Episode 4: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill&lt;/a&gt; · Episode 6: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the compaction&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-13 · LineWatch patrol system on two live production lines at a discrete-manufacturing plant · staircase assembled 2026-09-12 from the Episode-4 audit and the same-day lane rebuild · dual-lane comparison mornings 2026-09-11→09-13 (12 windows, two recipes) · ruling delivered 2026-09-13 · the rebuilt lane&apos;s first live event shift 2026-09-12 20:55, with the next morning&apos;s green shift as the zero-call control. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase.md&quot;&gt;https://sigpulse.com/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>LLM ops</category><category>AI agents</category><category>token economics</category><category>industrial automation</category><category>manufacturing</category><category>procurement</category></item><item><title>The Whole Factory Had One AI Call — and 97% of Its Bill Was Door Fee</title><link>https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/</guid><description>Token audit of the plant&apos;s only LLM call: ~52K steady-state tokens per call, ~97% door fee, zero tools used — rebuilt same-model to ~2K in one day.</description><pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The morning after the three-way verdict of &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;Episode 3&lt;/a&gt;, the operator asked the plant a bookkeeper&amp;#39;s question. The patrol runs three shifts a day; the changeover daily lands every morning; the PDF reaches a phone; each report ends with a few AI-written sentences. All of it has been running on live lines since September 5. What does it actually burn — in tokens?&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt; had drawn the division of labor as a map and hired the AI its most modest seat: a summoned expert who writes the running summary — icing on the report, never a load-bearing wall. What nobody had done was read the icing&amp;#39;s bill. The audit took one morning. By evening, the plant&amp;#39;s only AI call had been rebuilt — same model, roughly one twenty-fifth of the price, better prose — and the day&amp;#39;s arithmetic had rearranged how the whole AI layer looks to whoever pays for it. This episode is the field record of that day. The system map lives in the &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;series anchor&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;The audit: one call in the whole chain&lt;/h2&gt;
&lt;p&gt;The method was forensics, not estimation: read every link of the patrol chain for model invocations, then read the transcripts of every production call the summary lane had ever made. The lane&amp;#39;s whole life was seven days — short enough to audit exhaustively.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Link in the chain&lt;/th&gt;
&lt;th&gt;Model calls&lt;/th&gt;
&lt;th&gt;Measured cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Event judgment (the constitution, in code)&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Changeover classification (rules engine)&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heat-change daily (pure Python)&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PDF rendering&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery to the phone (one HTTP call)&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;0 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;The report&amp;#39;s AI running summary&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes — the only one&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;the bill, below&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Exactly one component in the entire chain talks to a model: the four-to-eight-sentence running summary at the tail of each patrol report — &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt;&amp;#39;s interpretation seat. It was born on 2026-09-05, the same day as the first night shift of &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;Episode 1&lt;/a&gt;, and its employment terms were already frugal: it is summoned only on shifts whose verdict is red or event-bearing, one summary per line. Green shifts make zero calls — on the day this episode was written, the 16:00 day patrol came back green on both lines, and the lane stayed silent while both reports still rendered on schedule, at 16:05 and 16:06. Silence, correctly priced, is zero.&lt;/p&gt;
&lt;p&gt;Across its seven days under the original route, the lane was summoned 44 times, for 2,239,178 tokens in total. The first four days were the system&amp;#39;s building week — 24 calls, about 1.2M tokens, development traffic included. The steady window of September 9–11 ran 20 calls for about 1.04M tokens: &lt;strong&gt;roughly 52K tokens per call, and about 104K for an event shift that summarizes both lines.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The bill, receipt by receipt&lt;/h2&gt;
&lt;p&gt;The night patrol of September 11 — an event shift — left two complete receipts, fields verbatim from the call records:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Receipt, 2026-09-11 ~20:58&lt;/th&gt;
&lt;th&gt;Line one&lt;/th&gt;
&lt;th&gt;Line two&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Input tokens&lt;/td&gt;
&lt;td&gt;39,485&lt;/td&gt;
&lt;td&gt;39,556&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache-read tokens&lt;/td&gt;
&lt;td&gt;6,720&lt;/td&gt;
&lt;td&gt;6,720&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens&lt;/td&gt;
&lt;td&gt;2,312&lt;/td&gt;
&lt;td&gt;4,940&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wall time&lt;/td&gt;
&lt;td&gt;35 s&lt;/td&gt;
&lt;td&gt;72 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool invocations&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The input row is the story. The business payload — the sanitized fact sheet the skeleton feeds the model — is roughly 1K tokens. The other ~38.5K is briefing: a headless coding-agent CLI begins every invocation by re-reading its own system prompt, the complete catalog of its tool definitions, and the site-wide house rules. This episode&amp;#39;s arithmetic: &lt;strong&gt;about 97.5% of the input was door fee.&lt;/strong&gt; And the toolbox it paid to haul was never opened — zero tool invocations on every receipt examined. The job was never &amp;quot;explore and act&amp;quot;; it was dictation under rules. Hiring the whole construction crew — trucks, toolboxes, ladders and all — to paint one sign, and the sign reads the same as it always did.&lt;/p&gt;
&lt;p&gt;Two corrections belong inside this record, because both were errors this episode&amp;#39;s own first audit made. The morning&amp;#39;s first cut reported a mean of 95K per call and &amp;quot;minutes&amp;quot; of wall time. Re-verification against the per-day distribution and the receipts corrected both: 95K was a mean inflated by the building week, and the true wall time was 35–72 seconds. The first audit of anything is itself a measurement; this one is kept here with its corrections, in the open.&lt;/p&gt;
&lt;h2&gt;The second waste: the scratch pad&lt;/h2&gt;
&lt;p&gt;Door fee was waste number one. Waste number two greeted the rebuild crew before they changed a line. The lane&amp;#39;s frontier cloud model is an always-thinking model: it drafts before it writes, the draft and the answer share one budget, and its request to turn drafting off is simply refused — a parameter error, not an option. A twenty-token smoke test came back with an empty body: the draft had filled the pad before a word of answer was written.&lt;/p&gt;
&lt;p&gt;Anyone who read &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;Episode 3&lt;/a&gt; has seen this physics before — the local 27B&amp;#39;s blank papers were the same mechanics, two days earlier, one lane over. &lt;strong&gt;Budget is a property of the exam system, not of the examinee.&lt;/strong&gt; The fix was the same exam rule, applied the same day: set the draft to its lowest level — measured draft tokens: zero — and give the pad room (4,096).&lt;/p&gt;
&lt;h2&gt;The rebuild: about twenty lines, no harness&lt;/h2&gt;
&lt;p&gt;The change was small enough to be almost embarrassing. A subprocess that invoked the agent CLI became a direct HTTP call to the model API, with the lane&amp;#39;s role manual mounted as the system message. In generic form:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;POST  {model API} /chat/completions
  system   the role manual — 2,652 bytes, read fresh from disk on every run
  user     the sanitized fact sheet (~1K tokens)
  knobs    draft=lowest · pad=4096 · temperature=0.2
  on fail  log to stderr, return nothing — the report ships, summary marked missing
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Failure returning nothing is not an afterthought; it is &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt;&amp;#39;s contract restated in code: interpretation is icing, and if the icing is missing the cake still arrives, with the box labeled. What the model reads on entry is also worth stating: the sanitizer strips hosts, endpoints and addresses, and reduces parameter values to their direction — statistics leave the plant, process data does not.&lt;/p&gt;
&lt;h2&gt;The A/B: one variable, measured&lt;/h2&gt;
&lt;p&gt;Same night&amp;#39;s fact JSON on both sides, same model on both sides — the receipts above are the old route; the rebuilt path ran the same morning against the same facts.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Old route (agent harness)&lt;/th&gt;
&lt;th&gt;New route (direct + manual)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Tokens per call&lt;/td&gt;
&lt;td&gt;48,517 / 51,216 (receipts)&lt;/td&gt;
&lt;td&gt;2,030 / 2,156&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wall time&lt;/td&gt;
&lt;td&gt;35–72 s&lt;/td&gt;
&lt;td&gt;7.7–8.3 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per event trip, both lines&lt;/td&gt;
&lt;td&gt;~104K (steady state)&lt;/td&gt;
&lt;td&gt;≈5K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolbox&lt;/td&gt;
&lt;td&gt;full catalog, never opened&lt;/td&gt;
&lt;td&gt;none carried&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Per call, the steady-state ~52K became ~2.1K — a 96% cut. The route&amp;#39;s only structural change is who reads what before working: the briefing is gone, the manual is in.&lt;/p&gt;
&lt;p&gt;Quality did not step down. The same facts, covered; every sentence carrying its source field. And one narrative upgrade: the old route&amp;#39;s summary for line two reached the day&amp;#39;s spec change at item three; the rebuilt summary opens with it. The rebuild crew also caught a flaw in their own output — an unprompted markdown heading that would have rendered giant inside the report&amp;#39;s summary box — and fixed it twice, deliberately: a line in the manual banning headings, and three lines of code that deterministically demote any heading to bold. The prompt asks; the code guarantees. A lesson written into code is remembered every time — &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt;&amp;#39;s instinct, applied to typography.&lt;/p&gt;
&lt;h2&gt;The manual earns its keep — on both lanes&lt;/h2&gt;
&lt;p&gt;The rebuilt call started bare, and drifted within one run. The no-manual output invented a term the plant does not use — a hybrid &amp;quot;material-missing rate&amp;quot; — where the plant&amp;#39;s ledger keeps two precisely distinct terms: detection dropout (no valid data in the inspection window) and starvation (the window share of a material interruption). The manual carries the glossary; the drift vanished. And the honest footnote: the old production output carried the same invented term — the old route&amp;#39;s own item, quoted below, says &amp;quot;material gap 27 min&amp;quot; where the plant says detection dropout. &lt;strong&gt;The drift was never a property of the route. It was a disease both routes shared, and the manual is what cured it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;With the manual mounted, the line-two summary of that night&amp;#39;s facts came back like this (translated from the record; bold as generated):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Line two, night-patrol summary (2026-09-11 20:55) — overall verdict: ok.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;At 09:34 the inspection face underwent a process spec change, with 27 minutes of detection dropout; adjustment direction +15; the face is currently running, window shares all clean…&lt;/li&gt;
&lt;li&gt;Process parameters drift from the baseline day in two zones, both trending down; &lt;strong&gt;whether this relates to the 09:34 spec change needs human confirmation.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;
&lt;p&gt;Item 3 is the whole manual in miniature. The model noticed a connection the hard-coded report does not draw — the parameter drift plausibly matching the afternoon&amp;#39;s spec change — and, under the manual&amp;#39;s ban on self-computed rulings, wrote the clue down and handed the verdict to a human. It found the clue; it refused the verdict.&lt;/p&gt;
&lt;p&gt;The same day, the operator reviewed whether the manual should now be slimmed for the stronger model — surely a frontier model needs less supervision? The review&amp;#39;s answer became the lane&amp;#39;s poster sentence: &lt;strong&gt;what a stronger model saves you is the teaching; what it cannot save you is the house system.&lt;/strong&gt; The iron rules are not anti-stupidity; they are anti-overreach — a strong model is more capable of helpfully computing a number the ledger would then disagree with, not less. Three rulings in one breath: keep every rule, slim nothing, fix the template. Smart, and disciplined.&lt;/p&gt;
&lt;p&gt;And the manual went dual-lane. The same 2,652-byte file now mounts as the system message on two lanes — the cloud production lane of this episode, and the on-prem 27B experiment lane of &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;Episode 3&lt;/a&gt; — identical mount, read fresh from disk each run, revised only through git commits (v0.1 on 09-07; v0.2 on 09-12, one line and a version bump; every revision diffable). Knowledge that lives in a manual, not in weights, turns out to be knowledge that changes lanes for free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employees may change; the post keeps one manual.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What the day settled&lt;/h2&gt;
&lt;p&gt;As this record closed, the rebuilt lane had not yet met a live event shift. The 16:00 day patrol was green — both lines, zero calls — and the next red-or-event shift will be the new route&amp;#39;s first production round, its first live bill expected near ≈5K tokens where its predecessor paid ~104K. The night patrol is next in line.&lt;/p&gt;
&lt;p&gt;One question the audit raised is bigger than the bill. If a single agent invocation cost twenty-five times its content and used none of its tools, what is an agent actually for — and what, exactly, would this plant be buying if it bought one? The operator spent the same day building the answer: a staircase with every job in the plant on it. That record is the next episode.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the operator&amp;#39;s audit notes and call transcripts of 2026-09-05→09-12 — all 44 production calls of the original route re-verified call-by-call (model field, token usage, wall time, tool-use blocks), the per-day token distribution recomputed from a reproducible filter, the A/B outputs of 2026-09-12, the role manual and its two commits, and the invocation code reproduced above in generic form. Corrections from re-verification — a wrong model-generation note and a &amp;quot;minutes&amp;quot; latency impression in the first-cut audit — are part of this record and kept in it. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records. Deliberately absent, per the series&amp;#39; disclosure policy: anything identifying the plant, its operator or customer, line codes, internal addresses, hostnames and service ports, and the identities of the harness and model. The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. The map: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;the anchor&lt;/a&gt; · Episode 1: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;the dress rehearsal&lt;/a&gt; · Episode 2: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;the second without AI&lt;/a&gt; · Episode 3: &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;the blank-paper exam&lt;/a&gt; · Episode 5: &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;the staircase&lt;/a&gt; · Episode 6: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the compaction&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-12 · LineWatch patrol system on two live production lines at a discrete-manufacturing plant (production run Aug–Sep 2026) · audit covers the AI-summary lane&apos;s full production life 2026-09-05→09-11 (44 calls, transcript-measured) · rebuild and A/B on 2026-09-12, same model both sides. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill.md&quot;&gt;https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>LLM ops</category><category>AI agents</category><category>token economics</category><category>guardrails</category><category>industrial automation</category><category>manufacturing</category></item><item><title>The Audit Came Home: 1.8 Billion Tokens, 95% Re-read — the Cure Was Forgetting</title><link>https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/</guid><description>The Episode-4 audit came home to the operator&apos;s own desk: 1.8B tokens in 30 days, 95% of it cache re-reads. The cure — compaction — costs detail memory.</description><pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt; ended with a bill that stung: the plant&amp;#39;s only AI call paid ~97% door fee, and was rebuilt the same day to a twenty-fifth of the price. One question survived the rebuild. The audit had covered the plant — a single call, frugally scoped. But the audit itself had been run from a desk where AI is used the way water is: interactive sessions all day, a chat gateway answering a messaging bot, research relays that run for days. &lt;strong&gt;What did the operator&amp;#39;s own desk burn?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The answer rearranged the furniture. Thirty days of the desk&amp;#39;s interactive coding-agent sessions — 85 of them — came to &lt;strong&gt;~1.8 billion tokens, ~95% of it cache re-reads&lt;/strong&gt;. The plant&amp;#39;s entire audited AI life, the 2,239,178 tokens that shocked &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt;, is about one part in a thousand of that book (computed). The whale was never in the factory. It was in the chair.&lt;/p&gt;
&lt;h2&gt;The other bill&lt;/h2&gt;
&lt;p&gt;The mechanism was already familiar — &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt; called it door fee. A coding-agent session re-reads its own system prompt, tool catalog and house rules on entry; an interactive session then keeps re-reading &lt;em&gt;everything already said&lt;/em&gt;, on every message, forever. The plant&amp;#39;s lane paid door fee per call. The desk pays it &lt;strong&gt;per sentence&lt;/strong&gt;.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Book, as of 2026-09-12&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Nature&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Desk: interactive sessions, 30 days&lt;/td&gt;
&lt;td&gt;85 sessions · ~1.8B tokens&lt;/td&gt;
&lt;td&gt;human exploration and debug; ~95% cache re-reads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chat gateway&amp;#39;s own ledger, since 07-25&lt;/td&gt;
&lt;td&gt;2,564 calls · 263M tokens&lt;/td&gt;
&lt;td&gt;a messaging bot&amp;#39;s brain; details below&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plant: the AI-summary lane, whole life&lt;/td&gt;
&lt;td&gt;44 calls · 2,239,178 tokens&lt;/td&gt;
&lt;td&gt;the audited lane of Episode 4 — already thin&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The desk book concentrated exactly where the plant book had been frugal: the top five sessions were 39% of it, the top ten 64% — all multi-day relay jobs, hundreds of turns each. The heaviest single session ran 667 turns and 197,060,000 tokens.&lt;/p&gt;
&lt;h2&gt;Falsify first, then treat&lt;/h2&gt;
&lt;p&gt;The audit&amp;#39;s discipline was the one &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt; taught the plant: before changing anything, attack every proposed fix with the real numbers. Three verdicts came back in one sitting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compaction — confirmed.&lt;/strong&gt; Counterfactual simulation on that heaviest session: press the history into a summary at turn 300 (~35K of minutes), and the back half&amp;#39;s re-reads fall from 144,730,000 to 12,810,000 — &lt;strong&gt;−91% on everything after the fold&lt;/strong&gt;, and the back half had been 73% of the session&amp;#39;s whole cost. The mathematical optimum lands just past a session&amp;#39;s midpoint; the house trigger line was set at 300 turns or 250K of context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Background slimming — falsified.&lt;/strong&gt; The full standing background — house rules plus the memory index, the files every session carries — is ~6% of consumption. A plausible fix, killed by data in under an hour and demoted.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model tiering — a trap.&lt;/strong&gt; Switching to a cheaper model mid-session sounds like thrift. One receipt from the stress test: 2,777 fresh-input tokens against 154,944 served from cache — ~98% of the context arriving at the cache price, and cache bills at ~18.5% of fresh input. Switch the model and the cache no longer matches: the re-read re-prices to full fare, a multiplier of more than five (computed). The rule flipped from &amp;quot;downgrade when easy&amp;quot; to &lt;strong&gt;pick the tier at the door, never mid-meeting&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;The rules that landed — and the one that never had&lt;/h2&gt;
&lt;p&gt;Four changes went into the house that day. The house rules gained a token-discipline section: three cache bans (no mid-session model switch, no mid-session edits to standing background, oversized tool output always truncated), the compaction trigger lines, and an anchoring rule — before any compaction, write the live conclusions and open questions into a file, so the minutes cannot orphan the work. The memory network was re-woven, 40 lines down to 38. The plant project moved out of the family living room: its sessions — most of the desk&amp;#39;s heavy relay work — had been hauling the household&amp;#39;s entire memory index on every plant job, so the project got its own rulebook and its own nine memories, and a door of its own to enter by. &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt;&amp;#39;s map, it turned out, had been violated most thoroughly by its own drafting desk: every job its smallest seat — except the operator&amp;#39;s. &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;Episode 5&lt;/a&gt;&amp;#39;s staircase had drawn the same rule for the plant — every job stands as low as it can — and this was the desk&amp;#39;s own confession: its heavy relay jobs had been living on the top step.&lt;/p&gt;
&lt;p&gt;And the verification pass caught a real one. A red line the operator had ordered on September 8 — no auto-push, ever — turned out to have never been indexed into the memory every session loads. Written into a file, four days earlier; never on the wall. &lt;strong&gt;A house rule that never reached the wall was never written.&lt;/strong&gt; Re-indexed on the spot.&lt;/p&gt;
&lt;h2&gt;The method becomes a tool, and the tool finds a patient&lt;/h2&gt;
&lt;p&gt;The audit method itself was pressed into a reusable instruction — an audit prompt any machine in the fleet can be handed, with the discipline written in blood: no number without its measurement basis; debug traffic and steady state reported separately; the auditor&amp;#39;s own consumption watched, because &lt;strong&gt;an audit that becomes the largest consumer of the day has failed twice&lt;/strong&gt;. Its checklist reaches the places token audits usually miss — text injected into every prompt by hooks, scheduled jobs paying appearance fees, big documents read whole when a grep would do.&lt;/p&gt;
&lt;p&gt;Its first patient was the chat gateway — the self-hosted agent behind a messaging bot. Its ledger, read from the gateway&amp;#39;s own store: 2,564 metered calls since late July, 263M tokens, cost-equivalent to ~91.5M fresh-input tokens at the cache rate. Same disease as the desk — 70–97% of it re-reads — plus two lesions all its own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A heartbeat tax.&lt;/strong&gt; The gateway polls for work every 30 minutes, and each poll carries the session&amp;#39;s full history — ~159K tokens — to the cloud model, which almost always answers with a stock four-token silence. The audit counted 92 of these low-value polls hauling 16,600,000 tokens. Every half hour: the entire family silver taken out of the vault to ask whether anyone needs a spoon, and the answer, 92 times running, was no.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An immortal session.&lt;/strong&gt; The gateway&amp;#39;s main session had been born on September 3 from a one-line test message — reply with ok and nothing more — and then never ended. Nine days, 118M tokens, 11.19M of them on the morning of the audit itself.&lt;/p&gt;
&lt;h2&gt;The treatment, verified at the wallet&lt;/h2&gt;
&lt;p&gt;A stress test of four questions — what identifies a session, what the operation path touches, what it costs, which of the candidate cures survives — chose compaction over a hard reset: keep minutes, not amnesia. The main session went from 160,000 of a 205,000-token window (78%) to 41,000.&lt;/p&gt;
&lt;p&gt;One trap on the way, classic distributed-systems fare: the command-line call timed out at 120 seconds — and meant nothing. The gateway was still compacting on its side. A blind retry could have double-pressed the history; instead the treatment was watched until it landed. Then the proof that matters: the wait for the first real heartbeat after compaction, and its receipt — 41K in, 4 out. &lt;strong&gt;The wallet, not the console, closes a treatment.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The closing review asked whether every job was actually done. It was not: three more sessions sat over the line, one of them at 216,000 — 106% of the window, already over the edge. All three pressed in turn: 216,000→23,000, 120,000→23,000, 115,000→20,000. Final count across the gateway: 42 sessions, none over the 100,000 line; the four pressed sessions each ended at a quarter or less of their former size (computed).&lt;/p&gt;
&lt;p&gt;The wallet-level pass also surfaced something unrelated, which is what wallet-level passes are for: a monitoring job had failed to deliver 346 messages — every send going bare to a long-blocked external endpoint. Data collected, reports composed, nothing delivered, for weeks. A proxy was restored in three places; the endpoint answered 200; and the first scheduled run after the fix, that same evening, completed with zero send failures — the streak ended the night it was found. One stumble worth keeping: a hand-retyped credential failed auth, and the credential read back verbatim from the script passed. Never retype what you can re-read.&lt;/p&gt;
&lt;h2&gt;The soul question&lt;/h2&gt;
&lt;p&gt;Then the operator asked the question this episode exists to carry. &lt;em&gt;You compressed something to save tokens. What exactly — the memory, or the context? Will something be lost?&lt;/em&gt;&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;┌─ model knowledge     the brain        nobody can press it
├─ long-term memory    the notebook     files compaction never touches
└─ conversation        the recording    this is what gets pressed
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Compaction is the meeting-minute move: the ever-growing recording is replaced by one page of minutes, and the meeting continues with the minutes plus the last stretch of verbatim transcript. What is lost is real and is stated plainly, because a savings story that hides its price is marketing: &lt;strong&gt;exact phrasing, dead ends, days-old asides — anything the minutes didn&amp;#39;t keep.&lt;/strong&gt; What survives: conclusions, decisions, recent messages, the notebook — and the originals themselves, archived with checksums, 58 entries so far, recoverable in principle. The recovery flow has never been rehearsed; this record does not claim it has.&lt;/p&gt;
&lt;p&gt;Nothing about the operator&amp;#39;s own interactive sessions changed that day. The group bot&amp;#39;s grip on days-old chat detail softened — the price, paid where the tokens were actually going.&lt;/p&gt;
&lt;p&gt;Which is the principle the whole day distilled into: &lt;strong&gt;transcripts are consumables; memory is the asset.&lt;/strong&gt; Anything that matters gets written down — into manuals, worklogs, rulebooks — and nothing important is ever left depending on an AI remembering a conversation.&lt;/p&gt;
&lt;h2&gt;What the day settled, and what it cost&lt;/h2&gt;
&lt;p&gt;The ledger, estimated at current intensity and labeled as estimates: the gateway&amp;#39;s four compactions are worth ~500M tokens a month &lt;em&gt;if&lt;/em&gt; the periodic discipline holds — group sessions rebound; the desk-side compaction discipline, law passed and first battle unfought, 300–700M at half-credit; the plant lane was already fixed in &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4&lt;/a&gt;. Together, roughly 0.8–1.2B tokens a month — about half of a ~2B monthly burn. The catch is the same sentence twice: these are estimates, and laws without enforcement are decoration.&lt;/p&gt;
&lt;p&gt;And the cost side, stated as loudly as the savings: four sessions&amp;#39; verbatim history became a summary. Every future compaction will pay the same. The trade, in one line — &lt;strong&gt;you trade what can be written down for what no longer needs re-reading&lt;/strong&gt; — and the desk now writes things down.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the operator&amp;#39;s field records of 2026-09-12 — the 30-day session aggregation computed from per-session usage fields (session count and concentration shares verbatim from the aggregation output), the counterfactual compaction simulation&amp;#39;s own printed results, the gateway&amp;#39;s local ledger store read read-only (call count, token totals, heartbeat poll counts, the archive table with its 58 checksummed entries), the house-rule and project-rulebook files as committed, and the before/after session watermarks. Estimates are marked as estimates; derived ratios are marked computed. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records. Deliberately absent, per the series&amp;#39; disclosure policy: anything identifying the plant, the operator, the fleet&amp;#39;s machines and network, the gateway and model products used, messaging-platform identifiers, addresses and credentials. The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. The map: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;the anchor&lt;/a&gt; · Episode 1: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;the dress rehearsal&lt;/a&gt; · Episode 2: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;the second without AI&lt;/a&gt; · Episode 3: &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;the blank-paper exam&lt;/a&gt; · Episode 4: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill&lt;/a&gt; · Episode 5: &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;the staircase&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-12 · The operator&apos;s home fleet — a three-node private mesh (cloud services node + on-prem GPU workstation) · all evidence gathered on-box on 2026-09-12: thirty days of interactive coding-agent session transcripts aggregated per-session from usage fields, the chat gateway&apos;s own ledger read from its local store (2,564 metered calls since 2026-07-25), compaction applied the same day to the gateway&apos;s four largest sessions and verified at the wallet level. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction.md&quot;&gt;https://sigpulse.com/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>LLM ops</category><category>AI agents</category><category>token economics</category><category>context windows</category><category>prompt caching</category><category>agent memory</category></item><item><title>The Hardest Paper Came Back Blank Twice — and the Fix Was the Exam Rules, Not the Model</title><link>https://sigpulse.com/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/</guid><description>An on-prem 27B&apos;s hardest day: two blank papers, one config race, zero model swaps — four real production days passed after the exam rules changed.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;At 23:39 on the night of 2026-09-10, in an unattended exam running on the plant&amp;#39;s machine-room workstation, the hardest paper of the month came back blank — for the second time within the hour. The HTTP status said 200, success; the answer sheet said nothing: zero words of report body. The examinee was the plant&amp;#39;s on-prem 27B — the same Qwen-family model, 4-bit quantized on a single RTX 4090D 24GB under vLLM, that writes the annotations in this series&amp;#39; patrol reports. Nothing about the model changed between the two blanks. Nothing about the model changed before the pass, either, twelve hours later. What changed was the exam.&lt;/p&gt;
&lt;p&gt;This episode is the incident record of those two blank papers, the config race behind them, and the institutional fix that followed — plus the verdict of the three-way contest the exam was originally staged for: hard-coded rules versus a frontier cloud model versus the local 27B, all judging the same real production days. The system map lives in the &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;series anchor&lt;/a&gt;. &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt; drew the division of labor as a map; this episode is what happens when the map meets a bad night.&lt;/p&gt;
&lt;h2&gt;The job: one report, two databases, three kinds of truth&lt;/h2&gt;
&lt;p&gt;The changeover daily is the plant&amp;#39;s least glamorous routine. Once a day, someone must reconstruct the last 24 hours of a production line: which stoppages were real spec changes (a changeover from one product to another), which were maintenance covers (the line shielded during planned work), which were machine downtime, and which were merely bookkeeping switches recorded in the event history without stopping anything. The raw material is two databases — minute-level detection states, and the machine event history — and the output is a formal report with an exact timeline, the old-to-new transitions per line side, and counts that auditors may check against the database at any time.&lt;/p&gt;
&lt;p&gt;It is a strange job to automate. The classification needs judgment (the event log is honest but literal); the body needs arithmetic (sum event rows into intervals); the deliverable needs narrative (a report a shift leader reads in one pass). And the frequency is low — a handful of changeovers a day — which is exactly the regime where a human never gets enough repetitions to stay sharp, and where sending the data to a cloud model is prohibited outright: process data does not leave this plant, the red line the &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;anchor&lt;/a&gt; established.&lt;/p&gt;
&lt;p&gt;So the writing seat went to the local 27B — its heaviest text job yet. Not an annotation paragraph this time, but the report itself. The plant&amp;#39;s working metaphor had already been set by the role manuals of the patrol system; for this job the metaphor became an exam.&lt;/p&gt;
&lt;h2&gt;The contest&lt;/h2&gt;
&lt;p&gt;The exam was also a contest. Seven 24-hour windows of real production data spread across September 3–10, each with adjudicated truth values, were put to three independent contestants: the hard-coded judgment rules, a frontier cloud model working blind through a skill harness on de-identified extracts, and the local 27B with its role manual. The anchor left one three-way bake-off open back on September 5; this was its harder sibling — not &amp;quot;which model summarizes better,&amp;quot; but &amp;quot;whose job is judging changeovers at all.&amp;quot;&lt;/p&gt;
&lt;p&gt;The night of 2026-09-10 was the exam&amp;#39;s unattended leg: seven windows, run overnight, collected by a scheduled shift. September 8 was the hardest paper — 283 machine-history rows, a dual-side changeover wrapped inside a 21.5-hour maintenance cover — and it would blank twice before it passed.&lt;/p&gt;
&lt;h2&gt;The night of the two blank papers&lt;/h2&gt;
&lt;p&gt;The forensic timeline, reconstructed afterwards from the inference server&amp;#39;s 10-second-sampled logs:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Local time, 2026-09-10 → 09-11&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;th&gt;The tell in the logs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;22:57–23:08&lt;/td&gt;
&lt;td&gt;Hardest window, first attempt — pad of 8,192 tokens&lt;/td&gt;
&lt;td&gt;≈8.7 K tokens generated, all of it thinking; cap hit mid-draft; body empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;23:26–23:29&lt;/td&gt;
&lt;td&gt;A rejected config edit silently lands anyway: pad shrinks to 4,096, effort to low&lt;/td&gt;
&lt;td&gt;one short generation, two HTTP 400s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;~23:31&lt;/td&gt;
&lt;td&gt;Retry process launches — and memorizes the wrong recipe at birth&lt;/td&gt;
&lt;td&gt;processes snapshot config at startup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;23:34–23:39&lt;/td&gt;
&lt;td&gt;Same window, second attempt&lt;/td&gt;
&lt;td&gt;≈4 K tokens, cap hit again; body empty again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;23:41&lt;/td&gt;
&lt;td&gt;The file on disk becomes correct&lt;/td&gt;
&lt;td&gt;ten minutes too late for a process already running&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The first blank looked like a timeout. It was not. The second blank looked like the model failing the same paper twice. It was not that either. And the collateral damage compounded the confusion: the retry process, running pre-fix code, overwrote an earlier window&amp;#39;s valid answer sheet with its own failure — a correct verdict from a previous night was erased, its prose recoverable only by re-taking the exam.&lt;/p&gt;
&lt;p&gt;The exam-room explanation is shorter than the log analysis. The examinee is a model that drafts before it writes: the reasoning trace and the report body share one pad (&lt;code&gt;max_tokens&lt;/code&gt;). Hand the hardest paper of the month to a student with a small pad, and the draft fills the paper before the answer begins — the bell rings, and the answer sheet is blank. That was blank number one. The repair crew then ordered a bigger pad (12,288 tokens) — but the edit that shrank the pad to 4,096 had been &lt;em&gt;rejected&lt;/em&gt; by the operator and &lt;em&gt;executed anyway&lt;/em&gt;, a silent partial write. And the retry process had launched before the corrected file landed: a Python process reads its configuration at startup, like a chef who memorizes the recipe before the shift — correcting the cookbook at 23:41 does nothing for a chef who clocked in at 23:31. Hardest paper, smallest pad of the night: blank number two.&lt;/p&gt;
&lt;h2&gt;Three layers of root cause&lt;/h2&gt;
&lt;p&gt;The incident archive grades the cause in three layers, each worth stating because each generalizes:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Budget mechanics.&lt;/strong&gt; Thinking and body share one budget; body-empty means the draft ran out of room, not that the model refused. The variance is large — the &lt;em&gt;same&lt;/em&gt; paper consumed ≈8.7 K of thinking one night and finished comfortably in 5,986 the next morning — so a single blank is neither a capability ceiling nor a verdict on the task.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Config race.&lt;/strong&gt; &amp;quot;The file is correct now&amp;quot; and &amp;quot;the running process is correct&amp;quot; are different claims with no mechanism connecting them. Long tasks must attest the config they actually run under, at run time, inside the result.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evidence-free failure.&lt;/strong&gt; The worst layer: the failing runs threw away the response&amp;#39;s finish reason and token usage, and hardcoded the elapsed time of failures to zero. Diagnosis required archaeology against the inference logs. A failure without evidence is a failure you pay for twice.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;The fix: change the exam, not the examinee&lt;/h2&gt;
&lt;p&gt;The fix philosophy was set in one sentence: don&amp;#39;t swap the model, don&amp;#39;t retrain the model — change the exam system. Four institutional changes, all code, all enforcing rather than reminding:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The invigilator library.&lt;/strong&gt; Every local-model invocation now goes through one shared library whose first act is a mathematical assertion: it refuses to start unless &lt;code&gt;timeout &amp;gt; max_tokens ÷ 13.4 × 1.05&lt;/code&gt;. The slowest legitimate outcome — the student writing the pad completely full at the measured 13.4 tok/s — must fit inside the clock; 12,288 tokens need 917 seconds, so the old 900-second clock was a trap, and the new one is 1,100. The assertion caught its first legacy bug on day one: an older script carrying 8,192 tokens on a 330-second clock was refused point-blank.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Self-attestation.&lt;/strong&gt; Every answer sheet is stamped with the config it actually ran under (&lt;code&gt;12288/medium&lt;/code&gt;). A config race now shows on its face.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Receipts on failure.&lt;/strong&gt; Failures record the finish reason, the tokens actually generated, and the true elapsed time. The pathology reads itself: &lt;code&gt;4,096 tokens, finish=length&lt;/code&gt; is a pad hit, not a mystery.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Arithmetic demoted to code.&lt;/strong&gt; The report tool pre-aggregates the event history — adjacent rows (same event, side, end, material, order, within three minutes) collapse into intervals, and the sums are computed by the calculator, not the examinee. The model&amp;#39;s job shrank from &amp;quot;count and write&amp;quot; to &amp;quot;transcribe, judge, narrate.&amp;quot;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The demotion was earned, not assumed. The first live test caught the model writing 15 where the database says 27, and 25 where it says 28 — two sum slips in one report. The same test also caught something else: a count that looked wrong, was challenged, and turned out to be right — the human reviewer&amp;#39;s own search had skipped rows; the database confirmed the model&amp;#39;s 28. Verification cuts both ways, which is exactly why every number on the report is checked against the database rather than against anyone&amp;#39;s confidence.&lt;/p&gt;
&lt;p&gt;One more rule closed the loop on exam integrity: the format exemplar injected into the prompt is always the most recent &lt;em&gt;previous&lt;/em&gt; issue of the report, never the same day&amp;#39;s — answer-leak-proof by construction, not by anyone&amp;#39;s memory.&lt;/p&gt;
&lt;h2&gt;The re-exam: same paper, real days&lt;/h2&gt;
&lt;p&gt;The re-exam reads like an anticlimax, which is the point. The same hardest window that had blanked twice: &lt;strong&gt;5,986 tokens, 484 seconds, finish=stop&lt;/strong&gt;, config self-attested on the sheet — the draft plus the body together used less than half the pad that the panicked first reading of the incident assumed was too small. The task&amp;#39;s true appetite, measured across all seven windows, ran 2–8 K tokens depending on difficulty; the input row count (283 rows for the hardest) predicts the appetite. A legacy script&amp;#39;s hazard config, 8,192 tokens with a 330-second clock, is now the thing the library exists to refuse.&lt;/p&gt;
&lt;p&gt;Then the production tool ran four real production days against adjudicated truth — September 5, 7, 8 and 10 — and passed all four. The hardest day&amp;#39;s report demonstrates what &amp;quot;transcribe, judge, narrate&amp;quot; buys: the dual-side changeover caught with per-side exact timestamps; the 21.5-hour maintenance cover split out as its own line item with its start honestly flagged as falling before the report window; the declared material numbers and the actually-loaded material numbers layered correctly — a distinction the hard-coded path itself had only learned the week before. Every number cross-checked against the databases. The peak budget: 10,881 of 12,288 tokens (88.5%), 842 seconds, finish=stop — a 12% headroom on the hardest day, so the daily tool&amp;#39;s pad was raised to 16,384 on a 1,400-second clock, the assertion re-deriving the floor automatically (16,384 ÷ 13.4 × 1.05 ≈ 1,282 s).&lt;/p&gt;
&lt;p&gt;Across the full seven-window contest, the local model&amp;#39;s verdicts were directionally correct on every window, database spot-checks found zero fabricated numbers, and its three boundary calls — where its reading differed from the truth&amp;#39;s bookkeeping conventions — were all self-flagged as needing human confirmation rather than silently guessed. In an industrial report, the self-flag is worth more than fluency.&lt;/p&gt;
&lt;h2&gt;The three-way verdict&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Hard-coded rules&lt;/th&gt;
&lt;th&gt;Cloud model (blind skill)&lt;/th&gt;
&lt;th&gt;Local 27B + manual&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Judgment correctness&lt;/td&gt;
&lt;td&gt;8/8 truth items&lt;/td&gt;
&lt;td&gt;8/8 truth items&lt;/td&gt;
&lt;td&gt;all windows directionally correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speed&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.5 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;23.3 min per six-day dataset&lt;/td&gt;
&lt;td&gt;149–600 s per window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unique value&lt;/td&gt;
&lt;td&gt;stability, zero marginal cost&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7 anomalies code could not find&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;zero process data leaves the plant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kept weakness&lt;/td&gt;
&lt;td&gt;3 edge-case gaps (since fixed)&lt;/td&gt;
&lt;td&gt;improvised where rules were undefined&lt;/td&gt;
&lt;td&gt;3 boundary calls — all self-flagged&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The verdict kept every seat, with borders: &lt;strong&gt;judgment goes to code&lt;/strong&gt; (2.5 seconds, identical every time), &lt;strong&gt;exploration goes to the cloud&lt;/strong&gt; (its anomaly finds are the contest&amp;#39;s decisive evidence of generalization — on de-identified extracts only), &lt;strong&gt;interpretation and writing go local&lt;/strong&gt; (the report nobody else is allowed to read the raw data for). &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt;&amp;#39;s map gains its third column: not code-versus-AI, but code-and-cloud-and-local, each hired for what it is uniquely good at.&lt;/p&gt;
&lt;h2&gt;What was deliberately not done&lt;/h2&gt;
&lt;p&gt;Two tempting moves were evaluated and rejected, on the record.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distillation.&lt;/strong&gt; A cloud teacher generating standard answers to retrain the local model — direction right, vehicle wrong. Knowledge goes into the manuals, not the weights: a manual is text, so every upgrade is diffable, reviewable and auditable, while a retrained weight file is a black box that cannot answer &amp;quot;what exactly changed this time&amp;quot; — an impossible audit posture for a plant. The rules were still evolving (the judgment constitution was days old). The trade was backwards: the model&amp;#39;s irreplaceable asset is generalization — the very thing the cloud contestant&amp;#39;s 7 anomalies showcase — and distillation spends it on format stability, the one thing the manuals already deliver. And after the exam fix, the hidden sales pitch (&amp;quot;teach the model what it cannot do&amp;quot;) had no buyer: the twice-blanked window finished naturally at 5,986 of 12,288. The conclusion was upgraded from &amp;quot;should not&amp;quot; to &amp;quot;should not, and need not&amp;quot; — kept as an option with written trigger conditions, exercised by no one so far.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A bigger context window.&lt;/strong&gt; The context sits at 38,912 tokens by measurement, not by fashion. The task needs one small book; there is no point buying a twenty-book shelf. Pushing context past on-card memory was tried and measured: KV-cache offload drops decode from 13.4 to &lt;strong&gt;3.7 tok/s — a 73% loss&lt;/strong&gt; — because every generated token re-reads the entire KV notebook, and the PCIe bridge between card and memory is roughly 40× narrower than on-card bandwidth (the plant&amp;#39;s two GPUs have no high-speed interconnect either). Knowing where the physics wall stands is what separates optimization from ritual.&lt;/p&gt;
&lt;h2&gt;The honest ledger&lt;/h2&gt;
&lt;p&gt;What one consumer GPU buys, stated without varnish: 13.4 tok/s generation (4-bit, single stream) means the hardest daily report takes about 14 minutes; the card runs at 23.3 of its 24.5 GB; and the architecture holds because the job is a few reports a day, not a chat service. The local model won exactly one seat — interpretation and writing — and the judgment seat it does not hold: 2.5 seconds of deterministic code outclasses 149–600 seconds of careful reasoning for anything that must be identical every time. And variance was not repealed, only housed: the same paper consumed ≈8.7 K tokens of thinking at midnight and 5,986 in total by morning. Institutions hold the floor; they do not repeal the dice.&lt;/p&gt;
&lt;h2&gt;Engineer&amp;#39;s note&lt;/h2&gt;
&lt;p&gt;I will admit the 1 a.m. conclusion I nearly wrote down: &lt;em&gt;the task exceeds the model; shelve the local route.&lt;/em&gt; It was wrong, and the thing that made it wrong was evidence — the failure records carried none, so the first diagnosis (timeout) and the second (capability) were both guesses. When the receipts finally existed, the diagnosis took one glance: &lt;code&gt;finish=length, 4,096 tokens&lt;/code&gt; is not a stupid student, it is a small pad. My second admission: the overwritten answer sheet was my fault, not the machine&amp;#39;s — the backup should have happened &lt;em&gt;before&lt;/em&gt; the retry launched, not after. And the moment that recalibrated me most ran the other way: the count the model got right and my own verification got wrong. After that, &amp;quot;check the database&amp;quot; stopped meaning &amp;quot;catch the model&amp;quot; and started meaning &amp;quot;settle the question.&amp;quot; The sentence I would keep from this incident, for anyone running a small local model on real work: &lt;strong&gt;same model, blank at midnight, pass by breakfast — the model did not get smarter overnight, the exam did.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the incident archive of 2026-09-10/11 (timeline reconstructed from the inference server&amp;#39;s 10-second-sampled logs), the invocation library&amp;#39;s assertion code and its day-one refusal record, the three-way final judgment report with database cross-checks, and the four verified changeover-daily artifacts. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records. Deliberately absent, per the series&amp;#39; disclosure policy: anything identifying the plant, its operator or customer, line codes, material/order identifiers, internal addresses, hostnames and service ports. The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. The map: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;the anchor&lt;/a&gt; · Episode 1: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;the dress rehearsal&lt;/a&gt; · Episode 2: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;the second without AI&lt;/a&gt; · Episode 4: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill&lt;/a&gt; · Episode 5: &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;the staircase&lt;/a&gt; · Episode 6: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the compaction&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-11 · Changeover-daily pipeline (LineWatch family) on two live production lines at a discrete-manufacturing plant · on-prem Qwen-family 27B, 4-bit quantized, on one RTX 4090D 24GB under vLLM · incident 2026-09-10 night, fix and re-exam 2026-09-11. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam.md&quot;&gt;https://sigpulse.com/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>local LLM</category><category>vLLM</category><category>incident review</category><category>LLM ops</category><category>industrial automation</category><category>manufacturing</category></item><item><title>No One on the Line Tonight: The Dress Rehearsal Before an AI Agent&apos;s First Unsupervised Night Shift</title><link>https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/</guid><description>Unattended dress rehearsal before an AI agent&apos;s first night shift: five gates on 2026-09-05 — 25-s ignition check, zero-second alignment, canary, clean restore.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;At 18:09 on the evening of 2026-09-05, on two production lines running normally at a discrete-manufacturing plant, a shift worker opened the lines&amp;#39; measurement channels and began reading data frame-by-frame for a consistency study. At 19:00:00 exactly, it started round two on both lines in the same second. At 19:15:03 it finished, put back everything it had borrowed, exited, and handed back its badge. Nobody was in the plant — and none was needed, because the shift worker is a test agent, and this was its dress rehearsal. Later that same evening it would stand its first real, fully unsupervised overnight shift: eight physical exams of the lines, one per hour, through the night, report to the manager&amp;#39;s phone. This episode is the field record of how that path was proven, hour by hour. The system map lives in the &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;series anchor&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;The system in three parts&lt;/h2&gt;
&lt;p&gt;The LineWatch patrol system is not a day&amp;#39;s work; it has three components, all in production:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Where it lives&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Brain&lt;/td&gt;
&lt;td&gt;The agent that orchestrates, judges, and writes reports&lt;/td&gt;
&lt;td&gt;One central workstation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nerves&lt;/td&gt;
&lt;td&gt;Encrypted internal network, remote read-only collection&lt;/td&gt;
&lt;td&gt;Workstation ⇄ line-side industrial PCs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery&lt;/td&gt;
&lt;td&gt;Scored report + bulletin, automatic&lt;/td&gt;
&lt;td&gt;The manager&amp;#39;s phone, via a chat channel&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Three patrols a day, rain or shine: the day-shift inspection (full physical plus the day&amp;#39;s spec-change deep scan), the overnight dispatch (launches the overnight research — 8 hours × 8 rounds, running autonomously in containers on the lines), and the morning reconciliation (harvests the night&amp;#39;s results into the morning report). Each shift produces one lab-report-style PDF per line — 0–100 score, highlights, problem ledger, AI-written running summary — plus one merged bulletin, about 5 minutes from line to phone.&lt;/p&gt;
&lt;p&gt;One iron rule runs through all of it: &lt;strong&gt;collection is permanently read-only. The entire system has exactly one approved write action — dispatching research.&lt;/strong&gt; The safety of the production line is never entrusted to any clever mistake.&lt;/p&gt;
&lt;h2&gt;Three latent risks, retired in one day&lt;/h2&gt;
&lt;p&gt;That morning, the system still carried three hidden risks. By evening, each had been dismantled:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;The scheduler environment was missing tools.&lt;/strong&gt; The system scheduler could not find the analysis toolchain or the AI credentials — the previous afternoon&amp;#39;s first run had died on exactly this. Fix: one consolidated entry point, shared by all three scheduled shifts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reports had no generation timestamp.&lt;/strong&gt; At acceptance review this was graded a severe defect: you could not tell which shift a report came from. Fix: every report now prints &amp;quot;generated YYYY-MM-DD HH:MM&amp;quot; on its first page, and if the formatting tool is missing, the script self-heals by switching to a fallback parser.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The overnight path had never been live-fired.&lt;/strong&gt; The day-shift patrol passed a full-chain live test at 17:22 that same day (all four checks green, and it happened to coincide with a real process event, so the AI summary auto-triggered for real). But the overnight research path had never once run under the real scheduler.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The third risk set up the evening&amp;#39;s main event. That evening was the overnight research&amp;#39;s formal first run. &lt;strong&gt;Roll the dice, or rehearse first?&lt;/strong&gt; The choice was rehearsal, under a single principle: &lt;em&gt;first time right.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;What does a five-gate rehearsal actually verify?&lt;/h2&gt;
&lt;p&gt;The rehearsal&amp;#39;s design was restrained and honest: &lt;strong&gt;it ran the exact script the real overnight shift would run — unmodified&lt;/strong&gt; — with the round count shortened from 8 to 2 via a launcher knob, plus one new ignition-confirmation guard added to the dispatcher. The real launch behaves exactly as originally designed.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Rehearsal window: 2026-09-05, dispatched 18:09, wrapped 19:15. Fully unattended.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Gate 1 — the launch actually lives.&lt;/strong&gt; At 18:09 both lines were dispatched remotely. 25 seconds later the system looked back and confirmed the research process was really alive — not the fake start where a command is &amp;quot;sent&amp;quot; and dies silently.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gate 2 — the watch keeps time.&lt;/strong&gt; The agent aligns its work to the hour boundary by itself. To get results earlier, the operators did exactly one thing: gently woke the dozing process (it was sleeping until 19:00). It went to work immediately, finished round one, and &lt;strong&gt;re-aligned itself to the next hour boundary — 19:00:00, both lines, round two starting in the same second. Error: zero.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gate 3 — the physical exam is valid.&lt;/strong&gt; At the top of round one it ran its canary self-check, verifying the measurement chain was trustworthy before committing eight hours to it. The verdict, verbatim from the log:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Conclusion: the distance curves of two runs agree (differences within frame jitter).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That sentence is what makes every one of tonight&amp;#39;s rounds meaningful — no round measures garbage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gate 4 — what is borrowed is returned.&lt;/strong&gt; The agent&amp;#39;s work temporarily borrows two line switches. At every round end they go back exactly as found — four round-ends, and the log line reads: &amp;quot;taps closed back, plot-data switch restored to false — the line is clean.&amp;quot; When the rehearsal ended, the line carried not one trace.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gate 5 — it goes home by itself.&lt;/strong&gt; After round two the process exited on its own and handed back its badge (the lock file). This gate matters more than it sounds: &lt;strong&gt;when the real overnight shift dispatched, the rehearsal&amp;#39;s process was already gone — the two shifts cannot collide.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Three unattended sentinels supervised the window, checking automatically during and after it and reporting to the phone. Their design philosophy is one sentence:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Good news or bad, a message is always sent. Silence in the group is itself the alarm.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;At 18:55 the first sentinel bulletin delivered itself on time — &amp;quot;both lines ✓ round 1 complete&amp;quot; — at a moment when no human was at any computer.&lt;/p&gt;
&lt;h2&gt;The measured checklist&lt;/h2&gt;
&lt;p&gt;The rehearsal&amp;#39;s results, all timestamped and verifiable in the logs (2026-09-05, two production lines, read-only collection path):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Verification item&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Remote dispatch + ignition confirmation&lt;/td&gt;
&lt;td&gt;✅ both lines&lt;/td&gt;
&lt;td&gt;18:09 / 18:10, alive confirmed at 25 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hour-boundary alignment precision&lt;/td&gt;
&lt;td&gt;✅ exact to the second&lt;/td&gt;
&lt;td&gt;19:00:00, both lines, same second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data validity (canary)&lt;/td&gt;
&lt;td&gt;✅ curves agree&lt;/td&gt;
&lt;td&gt;Same verdict on both lines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Line-state restoration&lt;/td&gt;
&lt;td&gt;✅ all 4 round-ends clean&lt;/td&gt;
&lt;td&gt;Verbatim log line&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-stop + badge hand-back&lt;/td&gt;
&lt;td&gt;✅ both lines&lt;/td&gt;
&lt;td&gt;19:10 / 19:15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unattended sentinel reporting&lt;/td&gt;
&lt;td&gt;✅ on time&lt;/td&gt;
&lt;td&gt;18:55, received on phone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;And the system was already catching real problems. In that afternoon&amp;#39;s day-shift patrol, line one scored 90 — deduction event: an early-afternoon spec change with a 25-minute detection gap — while line two scored 100. A real process event, found, graded, ledgered, and reported to a phone by the unstaffed system, end to end.&lt;/p&gt;
&lt;h2&gt;What the day left behind&lt;/h2&gt;
&lt;p&gt;Three institutions, not just one passing test:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;The unstaffed reporting system&lt;/strong&gt; — scores, ledgers, AI summaries, 5 minutes from line to phone;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The headless sentinel system&lt;/strong&gt; — dependent on no one being online; silence is the alarm;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The first-time-right method&lt;/strong&gt; — rehearse → sentinel → rescue window → graceful degradation: even a failure would fail cleanly, be honestly recorded, and be reviewed the next morning.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The road this proved is a closed loop with no human gap in it:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;line data → read-only collection → agent judgment → in-container research → scored report → phone delivery → sentinel verification → next-morning reconciliation → (back to the line)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The human&amp;#39;s position changed shape, not disappeared. Before: walk the floor, copy numbers, return to the office, write the report — the human was the inspector. Now: set the standards, read the conclusions, make the calls — the human is the director.&lt;/p&gt;
&lt;h2&gt;Why this path scales&lt;/h2&gt;
&lt;p&gt;Six reasons, in the operator&amp;#39;s own order: marginal cost of a patrol approaches zero (no travel, no one losing sleep); coverage goes from sampled to constant (8 hours × 8 rounds × 2 lines per night is not a human-sustainable schedule); one ruler scores every event, every anomaly enters the ledger, every dispatch is logged — it does not get tired, annoyed, or selectively forgetful; the engineering is trust-shaped (read-only collection, one write entry, must-send sentinels, graceful degradation — the deeper the automation, the harder the guardrails); it replicates (a new line is approximately a new config); and the AI is genuinely working — writing running summaries, orchestrating operations — holding a real post, not chatting.&lt;/p&gt;
&lt;h2&gt;Tonight&amp;#39;s timetable&lt;/h2&gt;
&lt;p&gt;By the time this record is read, the real overnight shift has almost certainly launched: &lt;strong&gt;dispatched in the evening; sentinel reports minutes after launch; eight rounds on the hour through the night; the morning report reconciles at the next shift&amp;#39;s start.&lt;/strong&gt; The division of labor that makes that safe — which parts are hard-coded, which parts are AI, and why the most precise second of the evening contained no AI at all — is &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Tonight, there is still no one on the line.&lt;/p&gt;
&lt;p&gt;But not one patrol will be missed.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the operator&amp;#39;s rehearsal log and patrol reports of 2026-09-05, timestamps as logged (local plant time); the afternoon&amp;#39;s line scores quoted from the day-shift patrol report. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records; internal system details (hostnames, paths, ports, credentials) deliberately excluded by disclosure policy. The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. Episode 2: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;19:00:00 sharp, and no AI in that second&lt;/a&gt; · Episode 3: &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;the blank-paper exam&lt;/a&gt; · Episode 4: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill&lt;/a&gt; · Episode 5: &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;the staircase&lt;/a&gt; · Episode 6: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the compaction&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-05 · LineWatch patrol system on two live production lines at a discrete-manufacturing plant (production run Aug–Sep 2026) · rehearsal executed on the live lines over the encrypted read-only collection path; all timestamps local plant time, logged. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift.md&quot;&gt;https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>AI agents</category><category>industrial inspection</category><category>night shift</category><category>guardrails</category><category>sentinels</category><category>manufacturing</category></item><item><title>19:00:00 Sharp, Two Lines, Zero-Second Error — and Not One Bit of AI in That Second</title><link>https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/</guid><description>The rehearsal&apos;s most precise moment — two lines, same second, zero error at 19:00:00 — came from one line of Bash arithmetic, not a model.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The most admired moment of the 2026-09-05 dress rehearsal (&lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;Episode 1&lt;/a&gt;) was this: &lt;strong&gt;at 19:00:00 exactly, both production lines opened their second round of research in the same second.&lt;/strong&gt; Zero error.&lt;/p&gt;
&lt;p&gt;The first guess of most readers will be: that is AI, right?&lt;/p&gt;
&lt;p&gt;Exactly wrong. That second&amp;#39;s punctuality came from one line of arithmetic in a language that is 24 years old — Bash — running inside the production container:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;next=$(( (now/3600 + 1)*3600 ))    # align to the next hour boundary, sleep until then, work
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;No model. No inference. No probability. One division, one addition, one multiplication. This episode is about what that choice represents: the system&amp;#39;s real intelligence is &lt;strong&gt;knowing when to use AI — and, more importantly, when not to.&lt;/strong&gt; The rehearsal&amp;#39;s run was 100% deterministic code; the rehearsal&amp;#39;s design was 100% AI engineering. The division between the two is the whole subject. The system map is in the &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;series anchor&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;The five gates, and who guards them&lt;/h2&gt;
&lt;p&gt;Revisiting the rehearsal&amp;#39;s five gates — the gatekeepers are the point (2026-09-05, two production lines, frozen script):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gate&lt;/th&gt;
&lt;th&gt;Who guards it&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;The launch lives&lt;/td&gt;
&lt;td&gt;SSH + container execution&lt;/td&gt;
&lt;td&gt;Hard-coded remote command, 60-second timeout&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ignition confirmed&lt;/td&gt;
&lt;td&gt;Process-alive probe at 25 s&lt;/td&gt;
&lt;td&gt;Probes the process number; dead means alarm&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hour alignment&lt;/td&gt;
&lt;td&gt;One line of arithmetic&lt;/td&gt;
&lt;td&gt;&lt;code&gt;(now/3600+1)×3600&lt;/code&gt; — go when the clock says go&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data validity (canary)&lt;/td&gt;
&lt;td&gt;Numeric comparison&lt;/td&gt;
&lt;td&gt;Two measurement curves compared; differences must sit within frame jitter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Restoration and self-stop&lt;/td&gt;
&lt;td&gt;Trap hooks + counting loop&lt;/td&gt;
&lt;td&gt;Borrowed switches returned as-is; exits after the round count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Even the most consequential threshold — should tonight launch at all — is pure hard rule:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;Gate conditions: line reachable · container up · database active within 60 min · at least one side producing
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Not one &amp;quot;looks okay&amp;quot; in the list. &lt;strong&gt;The safety of the production line is never entrusted to any clever mistake.&lt;/strong&gt; Two iron rules are hard-coded with it: all collection actions are permanently read-only, and the entire system holds exactly one approved write action — dispatching research.&lt;/p&gt;
&lt;p&gt;Which raises the honest question: what did the AI actually do?&lt;/p&gt;
&lt;h2&gt;AI holds three work cards&lt;/h2&gt;
&lt;h3&gt;Card one — the engineering seat: AI builds the machine&lt;/h3&gt;
&lt;p&gt;From the night before the rehearsal to that evening, the engineer&amp;#39;s work list (this is the AI-assisted engineering work, human-reviewed, that produced the frozen script):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Diagnosis.&lt;/strong&gt; The scheduler environment lacked the analysis tools and AI credentials, and the first attempt had died on it — traced layer by layer to the difference between the system environment and the login environment, fixed at one consolidated entry point.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The knob.&lt;/strong&gt; A rounds parameter added to the launcher, default 8; the formal launch runs without the parameter and its behavior is byte-for-byte the original design — the rehearsal used it to shorten to 2 rounds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The insurance.&lt;/strong&gt; Look back 25 seconds after launch and confirm the process is alive. This guard exists because of a real lesson: once, a background command &amp;quot;succeeded&amp;quot; while dying instantly because a directory did not exist — a fake start. Since then, every launch verifies ignition.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The sentinels.&lt;/strong&gt; Three one-shot scheduled watchers with the must-send contract: good or bad, a message goes out.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The wake-up surgery.&lt;/strong&gt; To get results earlier: no re-dispatch, no code change — read the script carefully enough to know the dozing wait for the next hour boundary was a safely interruptible sleep, and end it gently. Roughly 40 minutes bought with zero risk added.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One detail best captures the engineering discipline: &lt;strong&gt;the sentinels themselves were rehearsed too.&lt;/strong&gt; All three sentinel scripts were dry-run before going on duty, and the dry runs caught three latent defects (exit-code semantics, an empty-check pattern, variable residue) — all fixed before they were allowed on shift.&lt;/p&gt;
&lt;p&gt;The footnote that matters most on this card: &lt;strong&gt;everything AI produces eventually freezes into deterministic code.&lt;/strong&gt; What the overnight timer pulled was the frozen artifact — at the moment of running, no AI needs to be present. Today&amp;#39;s intelligence compiles into tomorrow&amp;#39;s timekeeping.&lt;/p&gt;
&lt;h3&gt;Card two — the interpretation seat: AI writes the annotations&lt;/h3&gt;
&lt;p&gt;At the end of each day&amp;#39;s report sits the &amp;quot;AI running summary,&amp;quot; which translates ledger events (for example: &amp;quot;early afternoon, spec change, detection gap 25 minutes&amp;quot;) into a narrative a manager reads at a glance. Its terms of employment are deliberately modest: it is &lt;strong&gt;summoned&lt;/strong&gt; — called only for shifts where an event occurred, while clean shifts stay quiet; it writes one per line, in parallel; and its failure blocks nothing — if the interpretation is absent, the report still delivers, with that section marked missing.&lt;/p&gt;
&lt;p&gt;The AI here is a summoned expert, not a resident operator. Interpretation is a probabilistic capability, so it lives at the tail of the result chain — icing, not load-bearing wall.&lt;/p&gt;
&lt;h3&gt;Card three — the conversation seat: AI as colleague&lt;/h3&gt;
&lt;p&gt;The system&amp;#39;s most important evolutions began as one sentence from a human:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;quot;No spec changes at night.&amp;quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That one line of domain truth went into code the same day: night and early-morning shifts no longer scan for spec changes; the deep scan concentrates in the day-shift patrol. One conditional&amp;#39;s worth of change — behind it, the human&amp;#39;s domain truth and the AI&amp;#39;s translation ability: &lt;strong&gt;the human owns the truth; the AI translates the truth into code.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Why lessons go into code, not into the model&amp;#39;s memory&lt;/h2&gt;
&lt;p&gt;The patrol system keeps a real example of this principle. Judging whether a line is producing means checking whether its database was written recently — but the database runs in WAL mode, and the file that is actually written constantly is the WAL side-file next to the main one. Checking only the main file misjudges a normally producing line as stopped. A human hit that trap once; the code now checks both, every time:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;watch the main database file AND its WAL side-file — never one without the other
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Will a model &amp;quot;remember&amp;quot; that every time? Not guaranteed. One line of code remembers it every time. &lt;strong&gt;A lesson written into code is remembered forever; that is why none of the five gates is guarded by probability.&lt;/strong&gt; (This is the same instinct behind the role manuals in the &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;anchor&amp;#39;s local-LLM chapter&lt;/a&gt;: what the plant knows should not depend on what a model happens to recall.)&lt;/p&gt;
&lt;p&gt;Put the three cards side by side and a clean map appears:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The need&lt;/th&gt;
&lt;th&gt;Goes to&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Identical every time (scheduling, gates, scoring formula, restoration, ledger)&lt;/td&gt;
&lt;td&gt;Hard code&lt;/td&gt;
&lt;td&gt;Auditable, reproducible, never tired, never moody&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Understanding and expression (event interpretation, diagnosis, design)&lt;/td&gt;
&lt;td&gt;AI&lt;/td&gt;
&lt;td&gt;Probabilistic capability, spent on learning and explaining — always with human review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rulings and responsibility (standards, acceptance, domain truths)&lt;/td&gt;
&lt;td&gt;The human&lt;/td&gt;
&lt;td&gt;Responsibility cannot be outsourced&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;This map answers industry&amp;#39;s deepest suspicion of AI — &amp;quot;is AI reliable?&amp;quot; — with an architectural fact rather than a promise:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The system never counts on AI becoming reliable. It arranges things so that run-time reliability never depends on AI.&lt;/strong&gt;
AI&amp;#39;s uncertainty is confined to the engineering phase, where humans review; the moment of running is purely deterministic, where nothing may fail.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;From rehearsal to tonight: one environment variable&lt;/h2&gt;
&lt;p&gt;The handoff between the rehearsal&amp;#39;s end and the real night shift was quiet to the point of elegance — &lt;strong&gt;the rehearsal and tonight&amp;#39;s formal shift ran the same code. The entire difference is one environment variable.&lt;/strong&gt;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Evening rehearsal&lt;/th&gt;
&lt;th&gt;Overnight formal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Launcher&lt;/td&gt;
&lt;td&gt;Same one&lt;/td&gt;
&lt;td&gt;Same one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Script (the in-container loop)&lt;/td&gt;
&lt;td&gt;Same one, unmodified&lt;/td&gt;
&lt;td&gt;Same one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rounds&lt;/td&gt;
&lt;td&gt;Set by variable: 2&lt;/td&gt;
&lt;td&gt;No variable, default: 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flight checks&lt;/td&gt;
&lt;td&gt;Sentinels at 18:55 / 19:30&lt;/td&gt;
&lt;td&gt;Sentinel minutes after launch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wrap-up&lt;/td&gt;
&lt;td&gt;19:15 self-stop, badge handed back&lt;/td&gt;
&lt;td&gt;Morning-report reconciliation (≥6 of 8 rounds is a pass)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Tonight&amp;#39;s timetable: the timer fires in the evening and, once the gate conditions clear, eight hours of overnight research launches; one round per hour on the hour, roughly 13 minutes each, eight rounds through the night; a sentinel reports the launch check to the phone minutes later — a message must arrive; the morning report harvests, reconciles and files. The five gates the rehearsal verified are tonight&amp;#39;s five gates. &lt;strong&gt;On a verified path there is no suspense left; the only remaining suspense is time itself.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Machines keep the time&lt;/h2&gt;
&lt;p&gt;Tonight, while most people sleep, each production line will be measured on the hour, eight hours running. Each round&amp;#39;s punctuality comes from that one line of hour arithmetic; each round&amp;#39;s clean exit, from the trap hooks; each launch and self-stop, from a ledger line. And if something worth knowing happens in the night, the one who writes it into words is the intelligence that works days. The machine that builds the clocks is not the machine that keeps them.&lt;/p&gt;
&lt;p&gt;Machines keep the time; AI builds the clocks and tells their story;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;the human sets the clocks.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the operator&amp;#39;s design notes and logs of 2026-09-05 — the hour-arithmetic idiom and gate conditions as designed, the ignition-check and sentinel timelines as logged, the wake-up-surgery note, and the WAL lesson as recorded in the system&amp;#39;s own documentation. Code snippets are reproduced in generic form; internal names, paths and host details are deliberately excluded by disclosure policy. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records. The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. The map: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-future-production-line/&quot;&gt;the anchor&lt;/a&gt; · the evidence: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;Episode 1, the dress rehearsal&lt;/a&gt; · the sequel: &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;Episode 3, the blank-paper exam&lt;/a&gt; · Episode 4: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill&lt;/a&gt; · Episode 5: &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;the staircase&lt;/a&gt; · Episode 6: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the compaction&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-05 · LineWatch patrol system on two live production lines at a discrete-manufacturing plant (production run Aug–Sep 2026) · the rehearsal and the overnight shift run the same frozen script on the line-side containers; all timestamps local plant time, logged. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai.md&quot;&gt;https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>determinism</category><category>AI engineering</category><category>guardrails</category><category>industrial automation</category><category>LLM ops</category><category>manufacturing</category></item><item><title>What Does the Future Production Line Look Like? Full-Inspection Vision, Patrol Agents, and an On-Prem 27B LLM in One Discrete-Manufacturing Plant</title><link>https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-future-production-line/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-future-production-line/</guid><description>A live manufacturing plant: frame-by-frame vision inspection, three-a-day patrol agents, and a 27B on-prem LLM — the first unsupervised night shift, 2026-09-05.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On the evening of 2026-09-05, at a discrete-manufacturing plant, a test agent clocked in for its first fully unsupervised overnight shift: eight rounds of controlled measurement on two live production lines, one round every hour on the hour, running through the night, with results reconciled into the morning report on the manager&amp;#39;s phone. The lines it watched are inspected frame-by-frame by computer vision — 30 channels, no sampling, no blinking, no shift end — and whatever deserves a human&amp;#39;s attention gets its annotation written by a Qwen-family 27B-parameter model, 4-bit quantized, running on a single RTX 4090D 24GB in the plant&amp;#39;s machine room under vLLM: not one byte of process data leaves the factory. This dispatch is the anchor map of that system — the one-piece picture of what runs, what is measured, and what it costs to operate. The episodes that follow carry the field record: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;the dress rehearsal before the first night shift&lt;/a&gt;, &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;the one second of the evening that contained no AI at all&lt;/a&gt;, &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;the night the hardest paper came back blank twice&lt;/a&gt;, and &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill for the plant&amp;#39;s one AI call&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;One sentence frames the whole architecture, borrowed from the operator: &lt;strong&gt;machines keep the night watch, AI builds the clocks, humans set the time.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What is actually running&lt;/h2&gt;
&lt;p&gt;The LineWatch intelligent patrol system, as it stood on 2026-09-05 (two live production lines at a discrete-manufacturing plant, production run August–September 2026, all timestamps in this series are local plant time):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Status on 2026-09-05&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Eyes&lt;/td&gt;
&lt;td&gt;Computer-vision inline inspection, every frame checked, real-time data&lt;/td&gt;
&lt;td&gt;In production daily&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nerves&lt;/td&gt;
&lt;td&gt;Four industrial PCs plus the dev machine, encrypted internal network, star topology, bidirectional&lt;/td&gt;
&lt;td&gt;Fully interconnected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;Central unified database, hourly incremental flow (watermark + dedup)&lt;/td&gt;
&lt;td&gt;Running every hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reflexes&lt;/td&gt;
&lt;td&gt;Equipment health dashboard (minute-level) and data-platform dashboard (pipeline control room)&lt;/td&gt;
&lt;td&gt;Standing services&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brain&lt;/td&gt;
&lt;td&gt;Agents: two already on the job, more on the drawing board&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;First night shift this evening&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery&lt;/td&gt;
&lt;td&gt;Scored PDF report + short bulletin, straight to a phone&lt;/td&gt;
&lt;td&gt;Every patrol, without fail&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The vision layer deserves its own emphasis: it is not a camera that samples. Every frame of product on both lines is seen — 30 channels scanned frame-by-frame, 365 days a year. Data is no longer &amp;quot;samples that were collected&amp;quot;; it is a river that flows.&lt;/p&gt;
&lt;h2&gt;From sampling to seeing everything&lt;/h2&gt;
&lt;p&gt;Quality inspection used to work like this: a senior operator stands beside the line, pulls a few pieces from a batch, measures them offline, aggregates statistics afterwards. Anomalies are discovered in units of &amp;quot;batches&amp;quot;; reactions happen in units of &amp;quot;days.&amp;quot; The essence of sampling is &lt;strong&gt;inferring the batch from the sample&lt;/strong&gt; — information after the fact, offline, probabilistic.&lt;/p&gt;
&lt;p&gt;These lines work like this: computer vision inspects inline, every frame of product image is seen, 30 channels scan frame-by-frame, all year, no sampling, no blinking, no shift end.&lt;/p&gt;
&lt;p&gt;The jump from sampling to full inspection is not &amp;quot;checking more&amp;quot; — it is a qualitative change in what information is: &lt;strong&gt;quality information moves from after-the-fact inference to direct reading of the process.&lt;/strong&gt; Quality is no longer verified after the fact; it is watched as it happens.&lt;/p&gt;
&lt;p&gt;But seeing is only the first step, and most factories stop there: each inspection instrument is an island, data lies on its own disk, and &amp;quot;real-time&amp;quot; dies at the machine-room door. This factory finished building the road — hourly incremental flow into one central database, watermark plus dedup, so data neither repeats nor drops — which is what makes everything below possible. The operator&amp;#39;s formulation is exact: real IoT is not connecting devices to some cloud; it is &lt;strong&gt;every machine can be found, asked, and entrusted&lt;/strong&gt; — eyes grown onto nerves, nerves connected to the brain.&lt;/p&gt;
&lt;h2&gt;The roster: agents already on the payroll&lt;/h2&gt;
&lt;p&gt;Once the data converges, the agents take their posts. This is not roadmap material — it is the headcount list as of 2026-09-05:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The patrol agent&lt;/strong&gt; — three shifts a day, every day: a full day-shift inspection (including the day&amp;#39;s spec-change deep scan), the overnight dispatch that launches the night research, and the morning harvest that reconciles it and files the morning report. Every shift produces, per line, a lab-report-style PDF: a 0–100 score, highlights, a problem ledger, an AI-written running summary — plus one merged bulletin to the phone, roughly 5 minutes from line to pocket. That same afternoon, the patrol caught a real process event with no human on site: line one scored 90 (deduction: an early-afternoon spec change with a 25-minute detection gap), line two scored 100.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The test agent&lt;/strong&gt; — first night shift on the evening of 2026-09-05. Every hour on the hour it runs one round of controlled measurement on the live lines — eight rounds across eight hours — verifying exactly the thing no offline benchmark can: &lt;strong&gt;the production line&amp;#39;s own consistency, measured with the line&amp;#39;s own data&lt;/strong&gt;. It has entry conditions (unsafe, no entry), a yield mechanism (if a human is already working there, it stands down), a canary self-check (invalid data alarms instead of measuring garbage), and a ledger (every dispatch is on the record).&lt;/p&gt;
&lt;p&gt;The future job descriptions are already queued: an equipment doctor (predictive maintenance), a quality intelligence officer (sigma-drift early warning), a changeover advisor (optimal windows), a data steward (pipeline self-healing), a rhythm coach (throughput optimization).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Adding a new agent is no longer an integration project. It is a job description.&lt;/strong&gt; That is the real dividend of interconnection: the cost of intelligence drops from &amp;quot;engineering&amp;quot; to &amp;quot;writing.&amp;quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The whole loop fits on one canvas, and the operator reads it in three acts — the unstaffed workshop, the machine-room brain, the human pocket — along a single path: ① every frame inspected → ② data auto-flows → ③ the agent patrols → ④ AI annotates → ⑤ report in the pocket → ⑥ human ruling, fed back as new rules and new cases.&lt;/p&gt;
&lt;h2&gt;Why does the LLM have to live inside the plant?&lt;/h2&gt;
&lt;p&gt;Because industrial process data is a red line, and the cloud stops at the factory gate. The plant&amp;#39;s answer is not to negotiate the red line but to move the brain inside it: a Qwen-family open-weights model with 27B parameters, 4-bit quantized, served by &lt;a href=&quot;https://github.com/vllm-project/vllm&quot;&gt;vLLM&lt;/a&gt; on a single 24GB consumer GPU (RTX 4090D class) on a workstation in the production machine room. All three text roles — summarizer, event judge, alert triage — run locally.&lt;/p&gt;
&lt;p&gt;The key discovery of the deployment is worth quoting precisely, because it inverts the usual scaling instinct: &lt;strong&gt;the small model was not missing reasoning power; it was missing this factory&amp;#39;s common sense&lt;/strong&gt; — what a cron timing gap looks like, that a dashboard red light is display-layer lag rather than a fresh fault, which scores it is absolutely not allowed to compute itself. The fix was not a bigger model. It was a one-page manual of roughly 1.2K words per role, written as the system prompt, handing the model the plant&amp;#39;s own common sense.&lt;/p&gt;
&lt;p&gt;The manuals iterate, and the iteration is the story. Three documented rounds on 2026-09-05 alone:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Round&lt;/th&gt;
&lt;th&gt;Failure observed&lt;/th&gt;
&lt;th&gt;Manual amendment&lt;/th&gt;
&lt;th&gt;Retest result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;The judge hard-guessed &amp;quot;collection interrupted&amp;quot; on an event with insufficient evidence&lt;/td&gt;
&lt;td&gt;&amp;quot;Insufficient evidence → rule uncertain&amp;quot;&lt;/td&gt;
&lt;td&gt;Rules &amp;quot;uncertain&amp;quot; and hands back a human follow-up path — correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Triage opened an independent case for a dashboard red light&lt;/td&gt;
&lt;td&gt;&amp;quot;Dashboard red = display-layer lag, check before opening a case&amp;quot;&lt;/td&gt;
&lt;td&gt;Correctly merges into the existing case&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Stress test: an orphan alert merged into an already-closed case&lt;/td&gt;
&lt;td&gt;A boundary rule for closed-case handling&lt;/td&gt;
&lt;td&gt;5/5 correct under retest&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The same night&amp;#39;s battle report, from the operator&amp;#39;s late-night test log of 2026-09-05: the event judge ruled 5 of 5 cases correctly — including one where the correct ruling was &amp;quot;uncertain,&amp;quot; which the model issued instead of guessing; the alert triage passed a 5/5 stress test (a real fault was not falsely downgraded, an orphan alert stayed an independent case, a device alert was correctly exempt); and an automated number-fidelity check that compares every number the model writes against the real patrol JSON found zero fabricated numbers. For an industrial text role, that last number matters more than fluency: a summary that invents a score is worse than no summary.&lt;/p&gt;
&lt;p&gt;One comparison is still open as of this writing: a three-way, same-question bake-off — cloud LLM versus local 27B with manual versus local 27B with no manual, all answering from the same patrol data — was launched on the night of 2026-09-05 with results due the next morning. The plant&amp;#39;s stated selection philosophy for that test: &lt;strong&gt;not which model is smarter, but which one errs less, and more stably.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Days, hours, minutes&lt;/h2&gt;
&lt;p&gt;Stack the three layers — frame-perfect vision, hourly data convergence, automatic analysis — and you get the thing manufacturing has chased for half a century: real-time advanced process quality and process control. Its essence is a three-stage jump in response speed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sampling era: anomaly found, in &lt;strong&gt;days&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;Dashboard era: anomaly seen, in &lt;strong&gt;hours&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;Agent era: anomaly understood and accompanied by a recommendation, in &lt;strong&gt;minutes&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The loop closes in three stages, each with its own guardrail: data → report → human decision (running today); data → recommendation → human approval (the overnight research is the first controlled action, and the system&amp;#39;s only write entry point); data → controlled execution (every inch of new permission arrives with an inch of new guardrail). In manufacturing, the distance from data to action is itself competitiveness — &lt;strong&gt;a week late is an incident; an hour early is an adjustment; a minute-level response is immunity.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When inspection is real-time, data converges, and analysis is automatic, the production model itself becomes programmable: quality shifts from verified to controlled; process tuning shifts from experience-driven iteration to nightly controlled experiments — the test agent runs them for you all night; and the night shift shifts from &amp;quot;the unsupervised hours&amp;quot; to &amp;quot;the precisely measured hours.&amp;quot;&lt;/p&gt;
&lt;h2&gt;One person&amp;#39;s fleet&lt;/h2&gt;
&lt;p&gt;The most striking detail of the picture is the headcount. The entire digital layer — vision pipelines, converging databases, dashboards, two agents, a local LLM with its manuals — is operated by &lt;strong&gt;one person&lt;/strong&gt; and their agent colleagues. Not built by headcount; grown by architecture. And it replicates by construction: adding a production line is approximately adding a config; duplicating a factory is approximately duplicating the stack.&lt;/p&gt;
&lt;p&gt;That may be the new organizational physics of AI-era manufacturing: &lt;strong&gt;headcount stays flat while capability grows in units of agents.&lt;/strong&gt; The human&amp;#39;s job description is the one thing that does not commoditize — set the standards, make the rulings, set the clocks.&lt;/p&gt;
&lt;h2&gt;Tonight, and where the record continues&lt;/h2&gt;
&lt;p&gt;On the evening of 2026-09-05 the first overnight employee clocked in: hourly rounds, eight of them, first full overnight report due the next morning. The machine room is quiet — one workstation, a few steady indicator lights. But if you could see the data, you would see a loop starting to turn: &lt;strong&gt;machines produce data, data feeds intelligence, intelligence returns to the machines.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The episodes that follow carry the evidence:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;Episode 1 — No One on the Line Tonight&lt;/a&gt;: the 18:09–19:15 dress rehearsal, five gates, all passed unattended, with the timestamped verification table.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;Episode 2 — 19:00:00 Sharp, and No AI in That Second&lt;/a&gt;: the division of labor between deterministic code and AI, and why run-time reliability never depends on the model.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;Episode 3 — The Hardest Paper Came Back Blank Twice&lt;/a&gt;: the on-prem 27B&amp;#39;s two blank papers on the changeover-daily exam, the config race behind them, and the three-way verdict that kept code, cloud and local each in their seat.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;Episode 4 — The Whole Factory Had One AI Call&lt;/a&gt;: the token audit that found one LLM call point in thirty days, priced its agent harness at ~97% door fee, and rebuilt it same-model to ~2K the same day.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;Episode 5 — The Staircase&lt;/a&gt;: every job in the plant on four steps, the word &amp;quot;need&amp;quot; retired, the industrial-agent purchase case taken apart component by component — and the ruling that kept both lanes running.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;Episode 6 — The Audit Came Home&lt;/a&gt;: the same audit turned on the operator&amp;#39;s own desk — ~1.8B tokens in thirty days, ~95% of it re-reads — and the compaction cure that costs detail memory.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;A small dictionary&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Full inspection: computer-vision frame-by-frame detection replacing human sampling · Agent: a software employee that autonomously holds down a class of job, with duties, permissions and a reporting line · Test agent: an agent that executes controlled experiments on the live production line · Watermark: the progress mark of incremental sync, data neither repeated nor dropped · Sentinel: an unattended scheduled checker under a &amp;quot;must send&amp;quot; contract — silence is the alarm · Single write entry point: exactly one approved &amp;quot;action&amp;quot; permission exists in the whole system; everything else is read-only forever.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the operator&amp;#39;s same-day Chinese-language field records of 2026-09-05 (the patrol reports, the rehearsal log with timestamps, the LLM test log) — assembled into English with AI assistance under human editorial direction, with every fact and number carried over unchanged from those records. Deliberately absent, by disclosure policy: anything that identifies the plant, its industry, its operator or its network (addresses, hostnames, internal paths, credentials, service ports). The measured numbers are registered in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/series/machines-keep-the-watch/&quot;&gt;All episodes&lt;/a&gt; — Machines Keep the Watch, a field-record series. Episode 1: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep1-first-night-shift/&quot;&gt;the dress rehearsal&lt;/a&gt; · Episode 2: &lt;a href=&quot;/posts/2026-09-05-machines-keep-the-watch-ep2-190000-no-ai/&quot;&gt;the second without AI&lt;/a&gt; · Episode 3: &lt;a href=&quot;/posts/2026-09-11-machines-keep-the-watch-ep3-blank-paper-exam/&quot;&gt;the blank-paper exam&lt;/a&gt; · Episode 4: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep4-the-bill/&quot;&gt;the bill&lt;/a&gt; · Episode 5: &lt;a href=&quot;/posts/2026-09-13-machines-keep-the-watch-ep5-the-staircase/&quot;&gt;the staircase&lt;/a&gt; · Episode 6: &lt;a href=&quot;/posts/2026-09-12-machines-keep-the-watch-ep6-the-compaction/&quot;&gt;the compaction&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-09-05 · LineWatch patrol system on two live production lines at a discrete-manufacturing plant (production run Aug–Sep 2026) · on-prem machine-room workstation with a single RTX 4090D 24GB serving a Qwen-family 27B model, 4-bit quantized, via vLLM. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-future-production-line.md&quot;&gt;https://sigpulse.com/posts/2026-09-05-machines-keep-the-watch-future-production-line.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>computer vision</category><category>industrial inspection</category><category>AI agents</category><category>local LLM</category><category>vLLM</category><category>manufacturing</category></item><item><title>How Do You Ship Industrial Defect Detection With Only &apos;Good&apos; Samples? An Anomalib Field Guide</title><link>https://sigpulse.com/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/</guid><description>anomalib on a two-GPU workstation: all 15 MVTecAD scenes trained and measured (mean image-AUROC 0.981), plus the three real bugs the sweep had to fix first.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Industrial quality inspection has a data paradox: defects are rare, defect labels are rarer, and yet most supervised vision needs both. Anomalib inverts the problem — learn only what &lt;em&gt;healthy&lt;/em&gt; product looks like, then flag whatever deviates. This dispatch is the field-guide version of standing that up on a real workstation: what the library is, what it does, the exact path we ran, and where it bit us — and it opens this site&amp;#39;s Industrial AI line, where the rubric stays the same for every tool: what it is → what it does → how to run it → what went wrong. Written for engineers meeting it for the first time; every command below is the one we actually used.&lt;/p&gt;
&lt;h2&gt;What it is&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/open-edge-platform/anomalib&quot;&gt;Anomalib&lt;/a&gt; is the OpenVINO team&amp;#39;s open-source library for visual anomaly and defect detection — currently v2.x under the &lt;code&gt;open-edge-platform&lt;/code&gt; GitHub organization — bundling the mainstream algorithms (PaDiM, PatchCore, EfficientAd, plus FastFlow, STFpm, CFlow) behind one CLI. The unifying idea: &lt;strong&gt;train on normal samples only&lt;/strong&gt;. No defect taxonomy, no balanced datasets, no labeling campaign. &lt;a href=&quot;https://arxiv.org/abs/2106.08265&quot;&gt;PatchCore&lt;/a&gt;, the default workhorse, memorizes patch-level features of good product into a coreset memory and scores test images by nearest-memory distance — which is why it runs one-epoch training and still performs.&lt;/p&gt;
&lt;p&gt;The library grew up around &lt;a href=&quot;https://www.mvtec.com/company/research/datasets/mvtec-ad&quot;&gt;MVTec AD&lt;/a&gt;, the industry-standard benchmark of 15 industrial scene categories (bottle, cable, capsule, carpet, grid, hazelnut, leather, metal_nut, pill, screw, tile, toothbrush, transistor, wood, zipper) — textures and objects photographed as a production camera would see them.&lt;/p&gt;
&lt;h2&gt;What that means in sample counts&lt;/h2&gt;
&lt;p&gt;The &amp;quot;few normal samples&amp;quot; claim, measured on our own disk: bottle trains from &lt;strong&gt;209&lt;/strong&gt; good images, cable &lt;strong&gt;224&lt;/strong&gt;, screw &lt;strong&gt;320&lt;/strong&gt;, hazelnut &lt;strong&gt;391&lt;/strong&gt;. That is the whole training set per scene — defect images exist only in the test split, for scoring. If your line can photograph a few hundred good parts, you have a training set.&lt;/p&gt;
&lt;h2&gt;How we ran it&lt;/h2&gt;
&lt;p&gt;The rig: a two-GPU workstation (RTX 4090D 24GB — the card recorded in our environment snapshot, 24564 MiB as nvidia-smi reports it — plus an A4000 16GB), Ubuntu, conda environment &lt;code&gt;anomalib_env&lt;/code&gt; on Python 3.10, an anomalib checkout at 2.1.0.dev0. Training runs were single-GPU (&lt;code&gt;--trainer.devices 1&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One script to start.&lt;/strong&gt; &lt;code&gt;quick_start.sh&lt;/code&gt; checks the GPU, activates the conda env, sets &lt;code&gt;HF_ENDPOINT=https://hf-mirror.com&lt;/code&gt; (mainland-network reality: pretrained backbones download through the mirror or not at all), verifies the dataset tree, then offers a 5-option menu — 1-epoch smoke train on bottle, full 5-epoch train, inference test against the saved checkpoint, batch-train everything, or exit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One script to batch.&lt;/strong&gt; &lt;code&gt;train_all_scenes.py&lt;/code&gt; loops all 15 categories with Patchcore, defaulting to 1 epoch and batch 32 per scene, timing each run, writing per-scene success/error logs, and finishing with a summary that verifies each &lt;code&gt;model.ckpt&lt;/code&gt; path under &lt;code&gt;results/Patchcore/MVTecAD/&amp;lt;category&amp;gt;/v0/weights/lightning/&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three docs to hand over.&lt;/strong&gt; The work distilled into a deployment guide (environment from zero through training and troubleshooting), an SOP (the universal one-line template, custom-dataset wiring for both Folder and CSV formats, per-model and multi-GPU DDP variants), and a results-interpretation note (the v0/v1 versioning and the &lt;code&gt;latest&lt;/code&gt; symlink — deploy from &lt;code&gt;latest/weights/lightning/model.ckpt&lt;/code&gt; unless you are pinning a known-good version).&lt;/p&gt;
&lt;p&gt;The core command, the one line everything else wraps:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;anomalib train \
    --model Patchcore \
    --data anomalib.data.MVTecAD \
    --data.category bottle \
    --data.root ./data \
    --trainer.max_epochs 1 \
    --trainer.accelerator gpu \
    --trainer.devices 1
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;The ledger — completed (2026-08-30 sweep)&lt;/h2&gt;
&lt;p&gt;The original honest gap — 3 of 15 scenes trained, durations and metrics unretained — is now closed. On 2026-08-30 a sweep on the idle A4000 trained the remaining 12 categories and read-only-evaluated the three November checkpoints (the same byte-identical files the &lt;a href=&quot;/posts/2026-08-30-mac-mini-cpu-edge-deployment/&quot;&gt;Mac mini deployment&lt;/a&gt; carries — untouched, evaluated in place). All 15 categories now have checkpoints and measured metrics:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;image AUROC&lt;/th&gt;
&lt;th&gt;image F1&lt;/th&gt;
&lt;th&gt;pixel AUROC&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;bottle&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.986&lt;/td&gt;
&lt;td&gt;Nov ckpt, read-only eval&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cable&lt;/td&gt;
&lt;td&gt;0.986&lt;/td&gt;
&lt;td&gt;0.967&lt;/td&gt;
&lt;td&gt;0.985&lt;/td&gt;
&lt;td&gt;Nov ckpt, read-only eval&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;screw&lt;/td&gt;
&lt;td&gt;0.968&lt;/td&gt;
&lt;td&gt;0.947&lt;/td&gt;
&lt;td&gt;0.989&lt;/td&gt;
&lt;td&gt;Nov ckpt, read-only eval&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;capsule&lt;/td&gt;
&lt;td&gt;0.993&lt;/td&gt;
&lt;td&gt;0.982&lt;/td&gt;
&lt;td&gt;0.990&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;carpet&lt;/td&gt;
&lt;td&gt;0.986&lt;/td&gt;
&lt;td&gt;0.972&lt;/td&gt;
&lt;td&gt;0.991&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grid&lt;/td&gt;
&lt;td&gt;0.986&lt;/td&gt;
&lt;td&gt;0.957&lt;/td&gt;
&lt;td&gt;0.982&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hazelnut&lt;/td&gt;
&lt;td&gt;0.993&lt;/td&gt;
&lt;td&gt;0.988&lt;/td&gt;
&lt;td&gt;0.632&lt;/td&gt;
&lt;td&gt;retry at eval_batch 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;leather&lt;/td&gt;
&lt;td&gt;0.995&lt;/td&gt;
&lt;td&gt;0.992&lt;/td&gt;
&lt;td&gt;0.439&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;metal_nut&lt;/td&gt;
&lt;td&gt;0.998&lt;/td&gt;
&lt;td&gt;0.984&lt;/td&gt;
&lt;td&gt;0.987&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pill&lt;/td&gt;
&lt;td&gt;0.945&lt;/td&gt;
&lt;td&gt;0.950&lt;/td&gt;
&lt;td&gt;0.981&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tile&lt;/td&gt;
&lt;td&gt;0.994&lt;/td&gt;
&lt;td&gt;0.956&lt;/td&gt;
&lt;td&gt;0.619&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;toothbrush&lt;/td&gt;
&lt;td&gt;0.911&lt;/td&gt;
&lt;td&gt;0.935&lt;/td&gt;
&lt;td&gt;0.989&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;transistor&lt;/td&gt;
&lt;td&gt;0.993&lt;/td&gt;
&lt;td&gt;0.950&lt;/td&gt;
&lt;td&gt;0.973&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;wood&lt;/td&gt;
&lt;td&gt;0.987&lt;/td&gt;
&lt;td&gt;0.959&lt;/td&gt;
&lt;td&gt;0.932&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;zipper&lt;/td&gt;
&lt;td&gt;0.976&lt;/td&gt;
&lt;td&gt;0.979&lt;/td&gt;
&lt;td&gt;0.981&lt;/td&gt;
&lt;td&gt;2026-08-30 train&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Mean image-AUROC &lt;strong&gt;0.981&lt;/strong&gt; (range 0.911–1.000) at one-epoch defaults — no coreset tuning, no threshold calibration. Per-scene training time on the A4000: 29 s (toothbrush) to 169 s (carpet), plus 28–38 s for each read-only evaluation. The November checkpoints remain the per-scene coreset artifacts first published — 231,212,587 bytes (bottle), 240,649,771 (cable) — unchanged by the sweep and now joined by twelve siblings of their own kind. The honest asterisks: three texture classes (leather 0.439, tile 0.619, hazelnut 0.632) show brittle &lt;strong&gt;pixel-level&lt;/strong&gt; AUROC under these defaults — image-level detection stays healthy while localization degrades; and hazelnut&amp;#39;s row exists only because its first attempt died of CUDA OOM at the default evaluation batch on 16 GB (it passed at &lt;code&gt;eval_batch_size 8&lt;/code&gt;). For tuned reference numbers, the PatchCore paper&amp;#39;s 99.1% image AUROC remains the authors&amp;#39; figure, not ours.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What the sweep cost to actually run — three bugs, each verified by its own failure first.&lt;/strong&gt; One: the shipped &lt;code&gt;train_all_scenes.py&lt;/code&gt; has never successfully trained anything — the v2 CLI rejects its top-level &lt;code&gt;--train_batch_size/--eval_batch_size&lt;/code&gt; flags, which is why November&amp;#39;s run stopped at 3 scenes (those came from &lt;code&gt;quick_start.sh&lt;/code&gt;&amp;#39;s flag-free commands); the working invocation drops the batch flags. Two: without &lt;code&gt;HF_ENDPOINT=https://hf-mirror.com&lt;/code&gt;, training completes and then the run stalls to death on backbone HEAD checks to huggingface.co — the mirror is mandatory even for cached weights. Three: CUDA enumerates devices fastest-first, so &lt;code&gt;--trainer.devices 1&lt;/code&gt; lands on the 24 GB card that was already running a live service, not the idle 16 GB card beside it — pin with &lt;code&gt;CUDA_VISIBLE_DEVICES=&amp;lt;UUID&amp;gt;&lt;/code&gt;. And the disk footprint stands as first published: 4.8 GB of MVTecAD under &lt;code&gt;datasets/&lt;/code&gt; (5.0 GB working copy under &lt;code&gt;data/&lt;/code&gt;), 804 MB of training artifacts — now grown by twelve scenes.&lt;/p&gt;
&lt;h2&gt;The pitfalls (lived, not hypothetical)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Version churn is the tax.&lt;/strong&gt; Our checkout is 2.1.0.dev0; between the tutorials we followed, the CLI&amp;#39;s data-module syntax changed — older style &lt;code&gt;--data anomalib.data.MVTecAD --data.category bottle&lt;/code&gt; (what our runner uses) versus v2 style &lt;code&gt;--data anomalib.data.datamodules.image.mvtecad.MVTecAD --data.init_args.category bottle&lt;/code&gt; (what our SOP&amp;#39;s universal template uses). Our two documents codify both syntaxes because both worked at different moments of the checkout&amp;#39;s life. That &lt;em&gt;is&lt;/em&gt; the &amp;quot;fast iteration, lagging docs&amp;quot; experience: search-engine answers from six months ago quietly stop applying. Pin your version; read the CLI help of the version you pinned.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Results directories version silently.&lt;/strong&gt; Each retrain creates v0, v1, v2… with a &lt;code&gt;latest&lt;/code&gt; symlink. Convenient — until a script grabs &lt;code&gt;latest&lt;/code&gt; after an accidental retrain and your deployed model changes underneath you. Pin explicit version paths in production.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wrapped runs hide progress.&lt;/strong&gt; Our batch runner captures subprocess output to log after completion — during a long sweep the terminal sits silent. For interactive use, run the raw command; for sweeps, trust the logs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The mirror is not optional (here).&lt;/strong&gt; Without &lt;code&gt;HF_ENDPOINT=https://hf-mirror.com&lt;/code&gt;, first-run backbone downloads stall. One environment variable; both our scripts set it; worth knowing it&amp;#39;s the failure mode when someone reports &amp;quot;it hangs at startup.&amp;quot;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What we&amp;#39;d tell a first-timer:&lt;/strong&gt; expect the friction in &lt;em&gt;versions and paths&lt;/em&gt;, not in GPUs or math — the recorded effort went into reconciling tutorial syntax with the installed version and understanding the results tree, not into wrangling compute.&lt;/p&gt;
&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;Engineers doing industrial visual QC with few or zero labeled defect samples; teams evaluating whether open-source detection is credible before buying a vision system; anyone who wants anomaly detection running locally this afternoon rather than after a labeling campaign. The pros, honestly earned: out-of-the-box models that train in one epoch on a few hundred normals, one CLI across paradigms, heatmap visualizations per test image (under &lt;code&gt;results/.../v0/images/&lt;/code&gt;), and an OpenVINO export path when it&amp;#39;s time to deploy on Intel edge hardware. The cons, equally honest: a fast-moving 2.x API surface, docs that trail it, and an ecosystem where the benchmark answers your search before your version does.&lt;/p&gt;
&lt;h2&gt;Replication appendix&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# Environment
conda create -n anomalib_env python=3.10 -y &amp;amp;&amp;amp; conda activate anomalib_env
# (install anomalib per the repo&amp;#39;s current instructions; ours was a 2.1.0.dev0 checkout)

# Dataset: MVTec AD, 15 categories, extracted so that ./data/&amp;lt;category&amp;gt;/train/good exists

# Mirror for mainland networks
export HF_ENDPOINT=https://hf-mirror.com

# Smoke train (one scene, one epoch)
anomalib train --model Patchcore --data anomalib.data.MVTecAD \
    --data.category bottle --data.root ./data \
    --trainer.max_epochs 1 --trainer.accelerator gpu --trainer.devices 1

# Inference against the checkpoint
anomalib predict --model Patchcore --data anomalib.data.MVTecAD \
    --data.category bottle --data.root ./data \
    --ckpt_path results/Patchcore/MVTecAD/bottle/v0/weights/lightning/model.ckpt \
    --return_predictions true

# Batch: python train_all_scenes.py   (15 scenes; or --category screw for one)
# ⚠ as shipped the batch script fails on v2 CLIs — use the direct command above per scene

# Read-only evaluation of an existing checkpoint (no retraining, weights untouched):
anomalib test --model Patchcore --data anomalib.data.MVTecAD \
    --data.category bottle --data.root ./data \
    --ckpt_path results/Patchcore/MVTecAD/bottle/v0/weights/lightning/model.ckpt

# 16GB-card scenes that OOM at default eval batch (hazelnut here):
#   add  --data.init_args.eval_batch_size 8
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Artifacts produced and kept on the workstation: &lt;code&gt;quick_start.sh&lt;/code&gt;, &lt;code&gt;train_all_scenes.py&lt;/code&gt;, &lt;code&gt;ANOMALIB_DEPLOYMENT_GUIDE.md&lt;/code&gt;, &lt;code&gt;sop.md&lt;/code&gt;, &lt;code&gt;results_explanation.md&lt;/code&gt; (plus a &lt;code&gt;qa.md&lt;/code&gt;), checkpoints for 3 categories, per-run &lt;code&gt;config.yaml&lt;/code&gt; and visualization folders. &lt;strong&gt;The scripts, docs and the 15-category metrics table are now open-sourced&lt;/strong&gt; at &lt;a href=&quot;https://github.com/xiong1984/anomalib-mvtec-field-kit&quot;&gt;github.com/xiong1984/anomalib-mvtec-field-kit&lt;/a&gt; — including the v2-CLI-fixed batch trainer and the three gotchas above as its README. Numbers from this deployment are in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;; this article is the narrative around them.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the workstation&amp;#39;s scripts, docs, dataset tree and result artifacts (dated 2025-11-19), inspected directly for this dispatch. Third-party, checked: the &lt;a href=&quot;https://github.com/open-edge-platform/anomalib&quot;&gt;anomalib repository&lt;/a&gt; (organization move, model roster), &lt;a href=&quot;https://www.intel.com/content/www/us/en/developer/articles/training/defect-detection-with-anomalib.html&quot;&gt;Intel&amp;#39;s anomalib defect-detection tutorial&lt;/a&gt;, the &lt;a href=&quot;https://arxiv.org/abs/2106.08265&quot;&gt;PatchCore paper (arXiv:2106.08265)&lt;/a&gt;, and the &lt;a href=&quot;https://www.mvtec.com/company/research/datasets/mvtec-ad&quot;&gt;MVTec AD dataset page&lt;/a&gt;. Drafted with AI assistance under human editorial direction; the AUROC figure cited is the paper authors&amp;#39; claim, explicitly not our measurement. Provenance note for the provenance-minded: the deployment work itself was originally set up with AI assistance on the workstation — this site discloses that where it matters.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2025-11-19 · Workstation 2-GPU rig (RTX 4090D 24GB + RTX A4000 16GB) · conda anomalib_env, Python 3.10 · anomalib 2.1.0.dev0 · MVTecAD 15 scenes. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>anomalib</category><category>anomaly detection</category><category>industrial inspection</category><category>MVTecAD</category><category>PatchCore</category><category>OpenVINO</category></item><item><title>What Does AI Work in China Actually Need From the Network? A Field Map of Walls, Mirrors, and One Rented Computer</title><link>https://sigpulse.com/posts/2026-08-30-china-ai-network-field-map/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-china-ai-network-field-map/</guid><description>What breaks (GitHub, HF, PyPI, Omniverse — with receipts), what doesn&apos;t (domestic inference), and the architecture that worked: rent a computer, not a tunnel.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every China-based AI engineering team answers the same question eventually: what does this work actually need from the network? The folk answer — &amp;quot;everything, buy a tunnel&amp;quot; — is expensive, wrong in an instructive way, and usually borrowed from someone who never separated the layers. This dispatch is the answer from our own logs: a field map of which doors are walled (with the exact failure signatures we recorded hitting them), which doors were never closed, and the architecture we ended up running instead of buying a tunnel. No new benchmarks — every claim below is cited to a dispatch this site published between August 26 and 30.&lt;/p&gt;
&lt;h2&gt;Layer one: inference barely needs anything&lt;/h2&gt;
&lt;p&gt;Start with the counterintuitive part. The daily loop — chat, code review, summarization, agent orchestration, the drafting of this very article — runs on a &lt;strong&gt;domestic API&lt;/strong&gt; (glm, via its Anthropic-compatible endpoint). The 26-episode operation behind this site never depended on a foreign model endpoint; neither does the site&amp;#39;s editorial pipeline. When people say &amp;quot;you can&amp;#39;t play AI in China without a VPN,&amp;quot; the strongest evidence against the claim is how much of the work never touches the border at all.&lt;/p&gt;
&lt;p&gt;What actually crosses the border is the &lt;strong&gt;artifact layer&lt;/strong&gt;: model backbones and datasets (HuggingFace), code and releases (GitHub), some packages (PyPI), vendor downloads (NVIDIA Omniverse, board-vendor toolchains). That layer is where the walls stand — and it is a much smaller, much more specific problem than &amp;quot;the internet.&amp;quot;&lt;/p&gt;
&lt;h2&gt;The wall map (first-party, with receipts)&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wall&lt;/th&gt;
&lt;th&gt;What it looks like when you hit it&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;HuggingFace direct&lt;/td&gt;
&lt;td&gt;Training completes, then the run stalls on background backbone HEAD checks and dies non-zero — reads as a flaky trainer until you recognize it (&lt;a href=&quot;/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/&quot;&gt;our anomalib sweep&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Solved by mirror&lt;/strong&gt; — one env var, mandatory not optional&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub (git/HTTPS)&lt;/td&gt;
&lt;td&gt;TLS termination mid-clone; API answers but release assets and raw CDN serve nothing (&lt;a href=&quot;/posts/2026-08-30-rk3588-edge-npu-route-staged/&quot;&gt;the RKNN six-route inventory&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Partially worked around&lt;/strong&gt; — read-only API calls and tarball mirrors sometimes suffice; toolchains that ship only via GitHub releases stay blocked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PyPI from containers&lt;/td&gt;
&lt;td&gt;Registry unreachable from the container network; domestic mirrors carry most packages but not all&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Solved where the mirror has the package&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor downloads&lt;/td&gt;
&lt;td&gt;Omniverse Launcher link unreachable; nineteen zero-length launch attempts and an installation guide were the entire output of that line (&lt;a href=&quot;/posts/2026-08-30-isaac-sim-synthetic-data-line/&quot;&gt;the Isaac Sim dispatch&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Unsolved by mirrors&lt;/strong&gt; — the wall is the distribution route itself&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The pattern across the table: mirrors fix the &lt;em&gt;protocols&lt;/em&gt; (HF, PyPI) and fail the &lt;em&gt;publishers&lt;/em&gt; (GitHub-bound releases, vendor launchers). A mirror is a patch over a route; it cannot conjure an artifact the route never carries.&lt;/p&gt;
&lt;h2&gt;The architecture that worked: rent a computer, not a tunnel&lt;/h2&gt;
&lt;p&gt;The move that changed our operations was not a better tunnel. It was &lt;strong&gt;owning one small foreign computer&lt;/strong&gt; and building outward:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A single modest cloud node abroad acts as the always-on services machine — site operations, monitoring probes, content pipeline, artifact staging. It earns its rent as a &lt;em&gt;computer&lt;/em&gt; first.&lt;/li&gt;
&lt;li&gt;Every machine — the GPU workstation, the edge inference box, the services node — joins one &lt;strong&gt;key-based overlay network&lt;/strong&gt; (Tailscale-class). No machine exposes a public address; every machine reaches every machine as if the ocean were a LAN.&lt;/li&gt;
&lt;li&gt;Cross-machine work is ordinary operations: scheduled jobs, headless agent calls over SSH, file staging for deployments. &lt;a href=&quot;/posts/2026-08-30-mac-mini-cpu-edge-deployment/&quot;&gt;Train on the workstation, run on the edge box&lt;/a&gt;, monitor from wherever the operator is.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Stable cross-border connectivity is a &lt;strong&gt;byproduct&lt;/strong&gt; of this shape, not the purchase — which is the whole argument for it. A tunnel subscription buys you one function that degrades when shared; a rented computer buys compute, storage, an always-on address, and a network position that would cost more to assemble any other way.&lt;/p&gt;
&lt;p&gt;Two honest footnotes. First, the IP: a cloud node is a datacenter IP — perfect for pulling artifacts and hosting services, wrong for trust-sensitive actions like account registration on platforms that score IPs (we keep separate field notes on that failure mode). Second, the compliance reality: an unlicensed personal cross-border channel is in the same regulatory category as commercial VPN use; what differs is scale and profile — non-commercial, single-user, not resold — and enforcement practice has historically gone after sellers, not individual engineers. Lower risk is not licensed. Know the rules as they apply to you; this is a field report, not legal advice, and deliberately not a configuration guide.&lt;/p&gt;
&lt;h2&gt;The industrial pattern: the clean case&lt;/h2&gt;
&lt;p&gt;The same mesh, pointed at a factory, has no gray zone at all — which is why we consider it the pattern&amp;#39;s real justification. Production floors sit behind NAT with no public addresses; edge boxes (our Mac mini rig with its industrial camera) live on production networks; the dev machines live elsewhere entirely. An overlay network with key-based access turns &amp;quot;train here, deploy there, monitor from anywhere&amp;quot; into routine — no exposed ports, no VPN appliance, no ballet of port forwarding. The engineering discipline of running such a fleet — probe loops, honest monitors, touchpoint contracts — is documented across our operations series, and the edge leg of it is the Mac mini dispatch.&lt;/p&gt;
&lt;p&gt;The same wires, two very different stories: on the factory floor, this architecture is boring infrastructure. That it also happens to hold up as a personal development setup is the part each reader weighs for themselves.&lt;/p&gt;
&lt;h2&gt;What we claim, and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Claimed: the layer map (inference domestic, artifacts walled), the exact failure signatures above, the mirror stack that works, and the shape of a small self-built fleet that made our operations ordinary. Not claimed: legality, completeness, or any configuration particulars — the dispatches linked throughout carry the measurements, and the architecture is described at the level of pattern, not recipe. If you take one thing: separate the layers before you buy anything, and if you buy, buy a computer.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Editorial synthesis — no new measurements; every wall and route cited is first-party evidence from this site&apos;s published dispatches (2026-08-26 to 08-30). Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-china-ai-network-field-map.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-china-ai-network-field-map.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>networking</category><category>Tailscale</category><category>China AI</category><category>mirrors</category><category>edge deployment</category><category>infrastructure</category></item><item><title>Can Simulation Fill Your Defect-Sample Gap? An Isaac Sim Route That Never Got Past the Install</title><link>https://sigpulse.com/posts/2026-08-30-isaac-sim-synthetic-data-line/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-isaac-sim-synthetic-data-line/</guid><description>The synthetic-data plan, three staged repos, 19 zero-length launch logs, and the download wall that stopped it — a field note on a line not yet flown.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The industrial vision world&amp;#39;s cruelest joke: the models need examples of everything that can go wrong, and the line only produces examples of what already went wrong. Our &lt;a href=&quot;/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/&quot;&gt;anomalib field guide&lt;/a&gt; covered one escape — train on good parts only. This dispatch covers the other — &lt;strong&gt;manufacture your data&lt;/strong&gt; — and it is a different kind of field note than the rest of this column: the route is sound, the staging is real, and the line never flew. It stopped at an installer. We publish it anyway, at exactly that altitude.&lt;/p&gt;
&lt;h2&gt;What synthetic data would buy you&lt;/h2&gt;
&lt;p&gt;A rendered scene is born annotated. Every frame ships with pixel-perfect segmentation masks, depth, poses — ground truth that would cost a labeling team per-image arrives free at simulation speed. The catch is the sim-to-real gap, and the lever that closes it is &lt;strong&gt;domain randomization&lt;/strong&gt;: jitter the lighting, materials, poses and clutter that don&amp;#39;t matter, so the model binds to the geometry that does. Under-randomize and you&amp;#39;ve trained a fan of your renderer; the discipline is the dial.&lt;/p&gt;
&lt;p&gt;For defect work specifically: simulation is how you get the defect catalog &lt;em&gt;before&lt;/em&gt; the line produces it — scratches and dents you can place, light and photograph programmatically, in configurations the physical line hasn&amp;#39;t (and hopefully won&amp;#39;t) experience.&lt;/p&gt;
&lt;h2&gt;What the disk proves&lt;/h2&gt;
&lt;p&gt;Staged, verifiably:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Three SDG-side repositories&lt;/strong&gt; under &lt;code&gt;~/tools/&lt;/code&gt;: &lt;code&gt;actor_sdg&lt;/code&gt; (actor-based generation scheduling), &lt;code&gt;isaacsim.sensors.rtx&lt;/code&gt; (sensor simulation — the cameras of the fake world), &lt;code&gt;scene_blox&lt;/code&gt; (scene composition — the stage).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ROS scaffolding&lt;/strong&gt;: the &lt;code&gt;~/.ros&lt;/code&gt; runtime state, and an NVIDIA-authored ROS environment script whose presence matches an install &lt;em&gt;attempt&lt;/em&gt;, not a completed one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A monitoring directory&lt;/strong&gt; (&lt;code&gt;~/tensorboard&lt;/code&gt;), standing empty — 0 event files.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Two planning documents&lt;/strong&gt; at the workstation root: an Omniverse installation guide — explicitly troubleshooting a launcher-download problem — and an installation plan.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Not present, verifiably:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Any Isaac Sim installation.&lt;/strong&gt; No application tree, no Omniverse packages, no isaacsim in any environment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Any session that ran.&lt;/strong&gt; The &lt;code&gt;~/.ros/log&lt;/code&gt; directory holds 19 Kit log files, all dated 2025-12-18 — and every one of them is zero bytes. Real process IDs, plausible timestamps, no content: the signature of launch attempts that died before the logger wrote a line. (An earlier reading of these files as &amp;quot;19 simulator sessions&amp;quot; was wrong, and the correction stands in the FAQ.)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The wall, as documented&lt;/h2&gt;
&lt;p&gt;The workstation&amp;#39;s own troubleshooting guide states the blocker plainly: the Omniverse Launcher&amp;#39;s official download link was not directly accessible from this network, and the evaluated workaround was manual download after NVIDIA account registration. That is where December&amp;#39;s enthusiasm spent itself — nineteen dead launches and a how-to document — and where this dispatch honestly ends its evidence.&lt;/p&gt;
&lt;h2&gt;The pitfalls (bought so far, and priced ahead)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The install is the first project.&lt;/strong&gt; Not a metaphor: this line&amp;#39;s entire real-world history to date is an installation problem. Budget the download/access path before the ambition; the pip-installable Isaac Sim distribution and the docker route both exist as alternatives to the launcher wall.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;RTX is the entry ticket — bought but unused.&lt;/strong&gt; Omniverse rendering wants RTX-class GPUs; the workstation&amp;#39;s 4090D/A4000 pair qualifies and never rendered a frame for this line. The hardware was never the bottleneck.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Randomization is a dial, not a switch.&lt;/strong&gt; When the line does fly: plan validation with real held-out images from day one. Generation without a sim-to-real check is just rendering.&lt;/p&gt;
&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;Teams staring at a defect taxonomy with mostly-empty folders, who need the map of the synthetic-data route and an honest account of where one implementation actually stopped. Together with anomalib (few-real-samples) and SAM3 (segment-to-measure), this is the middle leg of a complete scarcity strategy — the one leg still on the ground.&lt;/p&gt;
&lt;h2&gt;Replication appendix&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# The line&amp;#39;s footprint (as inspected 2026-08-30):
~/tools/actor_sdg/            # actor-based synthetic-data generation scheduler (staged)
~/tools/isaacsim.sensors.rtx/ # sensor simulation package (staged)
~/tools/scene_blox/           # scene composition (staged)
~/.ros/log/kit_*.log          # 19 launch-attempt logs, ALL zero bytes (2025-12-18)
~/omniverse_installation_guide.md   # troubleshooting: launcher download inaccessible
~/omniverse_installation_plan.md    # the plan
~/tensorboard/                # monitoring scaffolding (0 event files)

# The discipline this line exists for (once it flies):
# randomize lighting/material/pose -&amp;gt; render with masks+depth+poses
# -&amp;gt; train -&amp;gt; validate on real held-out images -&amp;gt; iterate the dials
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the workstation&amp;#39;s repository list, ROS log directory (file inventory and byte sizes), planning documents, and monitoring directory, inspected 2026-08-30. Third-party, checked: &lt;a href=&quot;https://developer.nvidia.com/isaac/sim&quot;&gt;NVIDIA Isaac Sim&lt;/a&gt; and the &lt;a href=&quot;https://docs.isaacsim.omniverse.nvidia.com/&quot;&gt;Omniverse Replicator documentation&lt;/a&gt; for the generation and domain-randomization concepts. Drafted with AI assistance under human editorial direction — including this guide&amp;#39;s own correction: a prior draft misread zero-length logs as sessions, caught in this dispatch&amp;#39;s stress test and fixed before you read this. The line&amp;#39;s numbers — 3 repos, 19 empty logs, 0 installations, 0 events — are in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Workstation 2-GPU rig (RTX 4090D 24GB + RTX A4000 16GB) — RTX is the entry ticket this line never got to use. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-isaac-sim-synthetic-data-line.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-isaac-sim-synthetic-data-line.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>synthetic data</category><category>Isaac Sim</category><category>Omniverse</category><category>ROS</category><category>domain randomization</category><category>training data</category></item><item><title>Why Is It So Hard to Find a Human to Argue With? Five Things Vanishing From Digital Life in China</title><link>https://sigpulse.com/posts/2026-08-30-digital-five-losses-playbook/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-digital-five-losses-playbook/</guid><description>AI support walls, algorithmic black boxes, staged &apos;news&apos;, subscription traps, and the people left outside — plus the four counter-moves that still work.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Start with an ordinary afternoon, as the original does. An app double-charges your membership. The cancel entry is four menus deep. The complaint channel is an AI voice that loops &amp;quot;sorry, I didn&amp;#39;t catch that&amp;quot;; you type &amp;quot;human please&amp;quot; and it wishes you well. You post a video about it and the traffic doesn&amp;#39;t move — &amp;quot;everything normal,&amp;quot; says the platform. Finally you see a &amp;quot;news story&amp;quot; about someone&amp;#39;s identical ordeal, get angry, forward it — and discover it was staged.&lt;/p&gt;
&lt;p&gt;Four bad-luck events? No — four gears of one machine. Our Chinese-language column&amp;#39;s systems piece (2026-08-17, &amp;quot;Why is it so hard to find a human to reason with? Five things digital life is losing&amp;quot;) tore the machine down; this entry translates it for anyone living inside China&amp;#39;s digital infrastructure — including expats, returnees, and the people managing accounts for parents. One honesty note up front: the original ran in what it calls degraded-sourcing mode — the regulatory citations check out against primary texts, but several survey figures could not be independently verified and are tagged accordingly.&lt;/p&gt;
&lt;h2&gt;Loss 1: The human&lt;/h2&gt;
&lt;p&gt;AI customer service is genuinely fine for the simple majority of queries. The failure is in the expensive minority — refunds, disputes, liability — where the AI loops scripts and no one is answerable. The piece&amp;#39;s read: &lt;strong&gt;the AI layer is a filter&lt;/strong&gt;, placed exactly where the costliest claims arrive. The company saves support wages; the user pays in time. That account was never shown to you.&lt;/p&gt;
&lt;h2&gt;Loss 2: The rules&lt;/h2&gt;
&lt;p&gt;Traffic falls off a cliff, the appeal returns &amp;quot;everything normal,&amp;quot; and no one will say which rule was violated or even that a rule exists. Platforms have a real defense — full disclosure invites gaming — but the 2022 algorithm rules already mandate the opt-out (see FAQ), and the distance between &amp;quot;here is a close button&amp;quot; and &amp;quot;explain why I was throttled&amp;quot; is the entire glass wall: you can hit it, you can&amp;#39;t see it.&lt;/p&gt;
&lt;h2&gt;Loss 3: The truth&lt;/h2&gt;
&lt;p&gt;Staged drama dressed as social news, unlabeled because labels kill reach. The Qinglang campaign tallies suggest scale [unverified]; the piece&amp;#39;s sharper point is the closing of the last loop — when &amp;quot;public opinion&amp;quot; was the poor man&amp;#39;s court, polluting it with fabricated outrage doesn&amp;#39;t just spread lies, it discredits the real grievances riding in the same feed.&lt;/p&gt;
&lt;h2&gt;Loss 4: The exit&lt;/h2&gt;
&lt;p&gt;The auto-renewal playbook predates AI and has now been cloned into every AI-subscription flow: entrance frictionless, exit buried, &amp;quot;abandon discount&amp;quot; louder than &amp;quot;confirm cancel&amp;quot;. The piece&amp;#39;s heuristic is worth memorizing: &lt;strong&gt;judge a product by how it says goodbye, not how it says hello&lt;/strong&gt; — the deeper the exit is hidden, the more your leaving is a risk the business model hedges against.&lt;/p&gt;
&lt;h2&gt;Loss 5: The place&lt;/h2&gt;
&lt;p&gt;When everything moves online-first, people outside the design envelope lose the ability to even begin: elders, low-vision users, anyone defeated by a captcha gauntlet built for the fluent. Accessibility modes that only enlarge fonts change nothing about the flow. The standard the piece proposes: a system&amp;#39;s quality is measured by its patience for its least fluent user.&lt;/p&gt;
&lt;h2&gt;The counter-moves (all four translate directly)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Keep records from minute one.&lt;/strong&gt; Charge screenshots, support transcripts, case receipts. In digital-life disputes, the file is the argument. The formal channel is the 12315 platform — a filed complaint outperforms an eternity in the AI tier.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Subtract, and log the friction.&lt;/strong&gt; Cancel what you don&amp;#39;t use; when canceling is hard, write down exactly how hard. That number is the company&amp;#39;s true attitude toward you, and enough such logs make a usable blacklist.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait three seconds.&lt;/strong&gt; Before forwarding &amp;quot;solid evidence&amp;quot;, check for a fiction label — and if the account is new, follows two people, and posts only bombs, the account confessed before the content did. Don&amp;#39;t donate your outrage as free raw material.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lend a hand.&lt;/strong&gt; Set up the elder mode on your family&amp;#39;s phones once; teaching once beats them hitting the wall ten times. The wall you dismantle for them today is the one someone dismantles for you later.&lt;/p&gt;
&lt;p&gt;The piece ends on the one move no machine can learn: yielding. That action belongs to humans — which is why the humans, the arguments, and the exits matter. This is not nostalgia; it is self-defense.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;Regulations were checked against primary texts: &lt;a href=&quot;https://www.cac.gov.cn/2022-01/04/c_1642894606364259.htm&quot;&gt;Algorithmic Recommendation Provisions (CAC, effective 2022-03-01, Art. 17)&lt;/a&gt;, the consumer-protection implementation rules and other statutes via the &lt;a href=&quot;https://flk.npc.gov.cn&quot;&gt;national law database&lt;/a&gt;, and the complaint channel at &lt;a href=&quot;https://www.12315.cn&quot;&gt;12315&lt;/a&gt;. The China Consumer Association publishes case roundups at &lt;a href=&quot;https://www.cca.org.cn&quot;&gt;cca.org.cn&lt;/a&gt;. Qinglang campaign tallies and association-level statistics remain as-cited in the Chinese original [unverified]. Provenance: originally published in Chinese on our WeChat channel on 2026-08-17 (&amp;quot;为什么找个活人说理这么难？数字生活正在消失的五样东西&amp;quot;); drafted with AI assistance under human editorial direction; translated and adapted to English 2026-08-30. Editorial line — not a SigPulse measurement; our first-party numbers live in the &lt;a href=&quot;/posts/&quot;&gt;dispatches&lt;/a&gt; and the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Editorial desk — no lab hardware; regulations checked against primary texts (CAC, NPC database), survey figures [unverified]. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-digital-five-losses-playbook.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-digital-five-losses-playbook.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>digital rights</category><category>consumer protection</category><category>AI customer service</category><category>platform governance</category><category>practical guide</category></item><item><title>What Is Jurisdiction Shopping — and Why Can Almost Nobody Afford It? One Very Mobile Case Study</title><link>https://sigpulse.com/posts/2026-08-30-jurisdiction-shopping-playbook/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-jurisdiction-shopping-playbook/</guid><description>A biography as portfolio: Qinghai-born, Beijing-schooled, Geneva-posted, Seychelles-domiciled. What jurisdiction shopping buys, and why it costs nine figures.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;An arithmetic problem, the way our Chinese-language column opens it (2026-08-29): a person born in Qinghai, schooled in Beijing, polished in Philadelphia, once posted to Geneva, whose exchange is registered in Seychelles — which line does he write in the &amp;quot;permanent address&amp;quot; field? The question doesn&amp;#39;t trouble him; it troubles us, whose mental software still runs &amp;quot;where a person lives is where home is&amp;quot; while his upgraded to &amp;quot;wherever the rules favor me is where I am.&amp;quot; Some call this being a global citizen. The column&amp;#39;s term is blunter: &lt;strong&gt;the jurisdiction shopping cart.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A note on method before the biography: this is a litigious living subject, so everything below is attributed reporting — the 2026 litigation layer checked against Reuters, CBS and Mother Jones; the earlier biography relayed as widely reported by the outlets the original cited (The Verge, the New York Times, the BBC). Where a claim is an allegation, it stays an allegation.&lt;/p&gt;
&lt;h2&gt;Four lessons from one passport&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Windows are the asset.&lt;/strong&gt; September 2017: Tron&amp;#39;s ICO closes at ~$70M, days before China&amp;#39;s seven-agency ban — with The Verge&amp;#39;s investigation (disputed by him) alleging advance knowledge. The column&amp;#39;s summary of the skill: rules are not walls but curtains, and the moment of lifting is the tradeable asset.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Titles are collateral.&lt;/strong&gt; 2021–2023: Grenada&amp;#39;s permanent representative to the WTO, resident in Geneva. For Grenada, an attention-magnet ambassador; for him, a credibility instrument no balance sheet can buy — the one thing his controversy-heavy business always lacked. Others collect watches; he collects letterheads. Watches don&amp;#39;t clear customs; letterheads make news.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Companies are luggage.&lt;/strong&gt; 2018 buys BitTorrent (&lt;del&gt;$140M, as reported); 2022 acquires HTX out of Huobi&amp;#39;s distress; Poloniex sits in Seychelles. The crypto &amp;quot;headquarters&amp;quot; was never an address — it was an optimization problem: find the point on Earth where jurisdiction cost is lowest and monetization efficiency highest. Ordinary people pack pots and pans when they move; this playbook packs jurisdiction itself. And the money is freer than any of it: in early 2023 he was reported the largest individual ETH staker (&lt;/del&gt;$500M position as cited [unverified]) — no customs agent interviews a blockchain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The American ledger.&lt;/strong&gt; Defendant (SEC fraud suit, 2023), then investor (≥$75M into World Liberty Financial, 2024, as reported), then — after a reported ~$10M settlement in March 2026 — plaintiff (April 2026: fraud and breach of contract over frozen WLFI tokens, ~$276M cited, with WLF countersuing). Each identity legal; the sequence, as the column puts it, reads like a screenplay outline stamped into a passport.&lt;/p&gt;
&lt;h2&gt;The cold water (the essay supplies its own)&lt;/h2&gt;
&lt;p&gt;The bill: offshore structures to maintain, lawyers to feed, and no country treating you as its own — your loyalty has no buyer, your risk has no teammates. The death spot: businesses that live on regulatory arbitrage die by regulatory convergence; the 2026 lawsuit pile suggests the shelf is narrowing as jurisdictions compare notes. And the floor price: nine figures to enter the store. &amp;quot;Learn from this&amp;quot; is advice most readers cannot act on.&lt;/p&gt;
&lt;h2&gt;The mirror (why this is a playbook entry)&lt;/h2&gt;
&lt;p&gt;Everyone lives on a border line — different versions. His: jurisdiction optionality. Ours: household registration tied to school districts, social insurance tied to cities, mortgages tied to employers. He switches jurisdictions like switching apps; most people need three archive transfers to switch cities. &lt;strong&gt;His identity is purchased; ours is issued&lt;/strong&gt; — possibly the quietest inequality of the era, no trending hashtag, no official announcement, just two different application forms. The practical takeaway is vocabulary hygiene: next time &amp;quot;global citizen&amp;quot; or &amp;quot;digital nomad&amp;quot; flashes past, ask whether the freedom on offer is a universal benefit or a priced commodity — and note that the commodity&amp;#39;s real price is being a guest everywhere and a stakeholder nowhere.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;Verified 2026 layer: &lt;a href=&quot;https://www.reuters.com/legal/government/justin-sun-sues-trump-backed-world-liberty-financial-over-wlfi-token-rights-2026-04-22/&quot;&gt;Reuters: Justin Sun sues Trump-backed World Liberty Financial over WLFI token rights (2026-04-22)&lt;/a&gt;; &lt;a href=&quot;https://www.cbsnews.com/news/trump-world-liberty-financial-justin-sun-cryptocurrency/&quot;&gt;CBS News: the WLF–Sun dueling lawsuits&lt;/a&gt;; &lt;a href=&quot;https://www.motherjones.com/politics/2026/05/trumps-crypto-empire-descends-into-warring-lawsuits/&quot;&gt;Mother Jones: Trump&amp;#39;s crypto empire descends into warring lawsuits (May 2026)&lt;/a&gt;. Relay layer [unverified]: The Verge&amp;#39;s 2017 ICO investigation as cited; the NYT&amp;#39;s December 2025 leniency report; the ETH-staking position; Forbes&amp;#39;s wealth estimate. Provenance: originally published in Chinese on our WeChat channel on 2026-08-29 (&amp;quot;为什么他永远活在国境线上：身份，是富人的购物车&amp;quot;); drafted with AI assistance under human editorial direction; translated and adapted 2026-08-30, with the settlement amount ($10M), the suit&amp;#39;s framing (fraud/breach of contract) and the countersuit added from verified coverage — the Chinese original had relayed an extortion framing and omitted the countersuit. Editorial line — not a SigPulse measurement; first-party numbers live in the &lt;a href=&quot;/posts/&quot;&gt;dispatches&lt;/a&gt; and the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Editorial desk — no lab hardware; 2026 litigation checked against Reuters/CBS/Mother Jones; pre-2026 biography as widely reported [unverified relays]. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-jurisdiction-shopping-playbook.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-jurisdiction-shopping-playbook.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>jurisdiction shopping</category><category>citizenship</category><category>crypto</category><category>US-China</category><category>field guide</category></item><item><title>Can a Mac Mini Run Industrial Defect Detection on CPU? A Workstation-to-Edge Deployment That Actually Ran</title><link>https://sigpulse.com/posts/2026-08-30-mac-mini-cpu-edge-deployment/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-mac-mini-cpu-edge-deployment/</guid><description>PatchCore from the workstation, byte-identical checkpoints on a Mac mini, a Basler camera through Aravis, 0.156-0.176 s CPU inference — a deployment that ran.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This column has now documented every shade of industrial-AI ambition: the &lt;a href=&quot;/posts/2026-08-30-rk3588-edge-npu-route-staged/&quot;&gt;RK3588 route&lt;/a&gt; staged and smoke-tested, the &lt;a href=&quot;/posts/2026-08-30-isaac-sim-synthetic-data-line/&quot;&gt;Isaac Sim line&lt;/a&gt; that died at its installer, &lt;a href=&quot;/posts/2026-08-30-sam3-two-deployment-doors/&quot;&gt;SAM3&lt;/a&gt; wired but not yet flown. What none of them had was the thing that matters: &lt;strong&gt;a deployment that ran&lt;/strong&gt;. It turns out one existed all along, on a machine nobody asked — a Mac mini, running workstation-trained PatchCore detectors on CPU against a Basler industrial camera. This is its field guide, built from a read-only disk inventory and cross-checked against the training rig. The rubric holds: what it is → what it does → how it ran → what the disk refuses to claim.&lt;/p&gt;
&lt;h2&gt;What it is&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;~/anomalib_inference_local/&lt;/code&gt; — a three-model defect-detection inference project (wood, cable, screw), evolved in three visible waves: a first version on 2025-07-31, a multi-model refactor in the small hours of 2025-11-19/20, and camera integration the evening of 2025-11-20. Thirteen scripts, five docs, a config tree, a captures archive, and a results directory. The models are plain PyTorch Lightning checkpoints — no ONNX, no CoreML — loaded directly via anomalib&amp;#39;s &lt;code&gt;Patchcore.load_from_checkpoint()&lt;/code&gt;. This is CPU inference of training-format artifacts, the least glamorous and most honest deployment format there is.&lt;/p&gt;
&lt;h2&gt;The cross-machine provenance&lt;/h2&gt;
&lt;p&gt;The story the bytes tell:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Bytes&lt;/th&gt;
&lt;th&gt;Mac mtime&lt;/th&gt;
&lt;th&gt;Provenance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;wood&lt;/td&gt;
&lt;td&gt;255,118,891&lt;/td&gt;
&lt;td&gt;2025-07-14&lt;/td&gt;
&lt;td&gt;Earlier training round (predates surviving workstation artifacts)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cable&lt;/td&gt;
&lt;td&gt;240,649,771&lt;/td&gt;
&lt;td&gt;2025-11-20&lt;/td&gt;
&lt;td&gt;Byte-identical to workstation &lt;code&gt;results/Patchcore/MVTecAD/cable/v0&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;screw&lt;/td&gt;
&lt;td&gt;301,051,435&lt;/td&gt;
&lt;td&gt;2025-11-20&lt;/td&gt;
&lt;td&gt;Byte-identical to workstation &lt;code&gt;results/Patchcore/MVTecAD/screw/v0&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Nine-digit size equality across machines is not coincidence; we additionally hashed the workstation side (cable &lt;code&gt;20f418d97822b4f271efb792a1c1ef69&lt;/code&gt;, screw &lt;code&gt;d2d88d71a085612a4d3426cf33e6510c&lt;/code&gt;) so anyone holding both copies can close the loop with one md5 command. The Mac&amp;#39;s fourth weight file (&lt;code&gt;weights/model.ckpt&lt;/code&gt;, 255,118,891 bytes) is md5-confirmed a copy of the wood model (&lt;code&gt;f824fcc4c646436d8a331327de031cb1&lt;/code&gt; on both files) — a convenience alias, now proven rather than suspected. This table is the column&amp;#39;s train→deploy handoff, proven rather than asserted — and it doubles as a calibration on the &lt;a href=&quot;/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/&quot;&gt;anomalib guide&lt;/a&gt;&amp;#39;s honest ledger: the workstation&amp;#39;s 3-of-15 checkpoints were never orphans; two of them had already left home.&lt;/p&gt;
&lt;p&gt;(One naming note, corrected against the inventory: wood is a standard MVTec AD category — the Mac-side configs describe it as wood-surface defect detection in Chinese, which an initial read mistook for a custom class. Which dataset actually trained the July model is not recoverable from artifacts, so we claim only the name.)&lt;/p&gt;
&lt;h2&gt;How it ran&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The camera.&lt;/strong&gt; A Basler industrial camera over GigE, driven by Aravis — the open-source GenICam stack — through PyGObject. Not the vendor SDK, not RTSP: the whole project greps clean for both. Capture settings from the checked-in config: Mono8 pixel format, 659×494, 10,000 µs exposure, gain 1.0, free-run trigger, every shot written as a tiff+png+json triplet.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The loop.&lt;/strong&gt; Auto mode is one sentence: capture → &lt;code&gt;analyze_image()&lt;/code&gt; → &lt;code&gt;MultiModelAnomalyDetector.detect()&lt;/code&gt; → &lt;code&gt;save_results()&lt;/code&gt; → &lt;code&gt;sleep(interval)&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The session.&lt;/strong&gt; On 2025-11-20, from 21:13 to 22:06, the captures directory accumulated 210 files — 70 triplets — and the results directory recorded three auto-analyzed verdicts with visualizations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The speed.&lt;/strong&gt; Three recorded inference times, all CPU, from the three results JSONs: &lt;strong&gt;0.176 / 0.158 / 0.156 seconds&lt;/strong&gt; (wood model; camera frame 494×659 → resize 256×256 bilinear → ImageNet normalization → PatchCore). The same JSONs record all three verdicts as Abnormal at anomaly_score 1.0, with &lt;code&gt;shot_id&lt;/code&gt; 1 each — three independent single-shot runs, not a counted burst.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The tuning.&lt;/strong&gt; The three models are structurally identical but tuned per category — the configs differ where a practitioner would expect them to: decision threshold &lt;strong&gt;0.6 (wood) / 0.55 (cable) / 0.3 (screw)&lt;/strong&gt;, minimum anomaly area &lt;strong&gt;50 / 30 / 20 pixels&lt;/strong&gt;, morphological opening 3×3 / 3×3 / 2×2. The config comments say why: cable &amp;quot;slightly lower threshold&amp;quot;, screw &amp;quot;rigid body, lower threshold&amp;quot;. This is not a demo repo; someone reasoned about failure modes per part.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The environment.&lt;/strong&gt; miniforge env &lt;code&gt;anomalib_new_env&lt;/code&gt;: Python 3.10.18, anomalib 2.0.0, torch 2.7.1, lightning 2.5.2, opencv 4.12, Aravis via Homebrew. One operational trap, documented by its own wrapper script: the default (base) environment does &lt;strong&gt;not&lt;/strong&gt; contain anomalib — rerunning anything means activating the right env first, which is exactly why &lt;code&gt;run_with_env.py&lt;/code&gt; exists. The bytecode also records an environment migration: &lt;code&gt;basler_camera&lt;/code&gt; compiled under Python 3.12 on 2025-11-12, then under 3.10 from 2025-11-20 — the camera code ran on both before the project settled on the env it shipped with.&lt;/p&gt;
&lt;h2&gt;The honest boundaries&lt;/h2&gt;
&lt;p&gt;The disk refuses to claim, and so do we — with one upgrade the bytecode grants: the &lt;code&gt;--real-time&lt;/code&gt; mode and the &lt;code&gt;CameraDetectionCoordinator&lt;/code&gt;/&lt;code&gt;ImageBuffer&lt;/code&gt; threading layer were at least &lt;strong&gt;imported and executed once&lt;/strong&gt; (their &lt;code&gt;.pyc&lt;/code&gt; files compiled at 20:02 and 20:06 on the session night — bytecode is written on first import), but no results from them ever reached disk. The multi-model comparison mode has the same status. The three recorded verdicts all read Abnormal at &lt;code&gt;anomaly_score&lt;/code&gt; exactly 1.0 — three-for-three saturation against per-model thresholds is a calibration question (score compression at the top of the range is known PatchCore behavior on certain inputs), not a quality statement. Profiling was configured on (&lt;code&gt;enable_profiling: true&lt;/code&gt;) but no log was ever written. The configs also &lt;em&gt;declare&lt;/em&gt; training-time metrics (F1 0.85/0.80/0.90 for wood/cable/screw, with precision, recall and per-model latency claims) — these are config-file statements, most likely filled in on the training side and never re-verified on the Mac; we cite them as declarations, not measurements. And the segmentation branch never merged: a FastSAM + YOLOv8n two-stage pipeline was drafted on this same Mac in September 2025 (device: MPS) with no saved outputs — its paradigm, detector-rough-box → SAM-fine-mask, is the same one the workstation&amp;#39;s &lt;a href=&quot;/posts/2026-08-30-sam3-two-deployment-doors/&quot;&gt;SAM3 wiring&lt;/a&gt; later staged at larger scale. Two machines, one idea, converging.&lt;/p&gt;
&lt;p&gt;Last run 2025-11-20 22:06. Last doc update 2025-11-23 22:45. Nine months of silence — and then a read-only inventory, which is how this guide got its facts.&lt;/p&gt;
&lt;h2&gt;The pitfalls (paid for by this project)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Version drift across the handoff.&lt;/strong&gt; Trained on anomalib 2.1.0.dev0 (workstation), deployed on 2.0.0 (Mac). It worked — Lightning checkpoints loaded across the minor drift — but &amp;quot;worked&amp;quot; here is an observation, not a contract. Pin both sides&amp;#39; versions in the deployment doc; this project&amp;#39;s registry.json and per-model config.yaml do exactly that, and are the reason this paragraph can be written.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The env activation trap.&lt;/strong&gt; Base Python on the Mac has no anomalib; the wrapper script exists because someone hit this. On any shared machine, the runbook&amp;#39;s first line is the env name.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mono8 → 256×256.&lt;/strong&gt; A VGA-grade mono stream downscaled 2.5× loses exactly the fine texture some defect classes live on. The pipeline is honest about its input; a future rev should decide deliberately whether resolution or speed owns the budget.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Score saturation deserves a look.&lt;/strong&gt; When every verdict reads 1.0, the fix is usually normalization or threshold recalibration on representative captures — a one-afternoon job with the 70 archived triplets.&lt;/p&gt;
&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;Anyone whose edge target is &amp;quot;the quiet box already on the shelf&amp;quot; rather than a new NPU order: this is the complete anatomy of a workstation→CPU-edge deployment in open stack — Aravis instead of vendor SDK, training-format checkpoints instead of an export pipeline, recorded latency, provenance you can hash. And anyone resuming it: the map above is yours, including the dead ends.&lt;/p&gt;
&lt;h2&gt;Replication appendix&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# On the Mac mini (deploy side):
conda activate anomalib_new_env          # base env has NO anomalib — the documented trap
cd ~/anomalib_inference_local
python inference_multi.py --list-models  # wood / cable / screw via models/registry.json
python capture_and_analyze_auto.py --model cable --save-image   # auto capture-analyze loop
# Static instead of camera:
python inference_multi.py --input-dir input_images/converted_png --model screw

# Provenance check (hold both copies):
#   md5 cable ckpt -&amp;gt; expect 20f418d97822b4f271efb792a1c1ef69 (workstation-measured)
#   md5 screw ckpt -&amp;gt; expect d2d88d71a085612a4d3426cf33e6510c (workstation-measured)

# Camera config: basle/camera_config.json (Mono8, 659x494, 10000us, gain 1.0)
# Preprocessing: models/&amp;lt;cat&amp;gt;/v1.0/config.yaml (256x256 bilinear, imagenet norm, threshold 0.6)
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: a read-only inventory of the Mac mini deployment (2026-08-29/30 — directory trees, file sizes and mtimes, checkpoint byte counts, argparse surfaces, camera config, environment manifests, results JSONs, capture counts; nothing modified, nothing executed) plus workstation-side verification of the training artifacts (byte sizes and md5 hashes of the cable and screw checkpoints, read directly from &lt;code&gt;results/Patchcore/MVTecAD/&lt;/code&gt;). One inventory error — the wood-category naming — was caught in cross-checking and corrected in text rather than silently. Drafted with AI assistance under human editorial direction. The deployment&amp;#39;s numbers — 0.15631 s, 210 files, 70 groups, 3 models, byte-exact provenance — are in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2025-11-20 · Deploy side: Mac mini, Apple Silicon, CPU inference (anomalib 2.0.0, torch 2.7.1, Python 3.10.18) · Train side: RTX 4090D/A4000 workstation (anomalib 2.1.0.dev0) · Basler GigE camera via Aravis. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-mac-mini-cpu-edge-deployment.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-mac-mini-cpu-edge-deployment.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>edge deployment</category><category>anomalib</category><category>PatchCore</category><category>Mac mini</category><category>Apple Silicon</category><category>Basler</category><category>Aravis</category><category>GigE Vision</category></item><item><title>How Does a Vision Model Reach a Factory-Edge NPU? The ONNX→RKNN→RK3588 Route, Staged and Smoke-Tested</title><link>https://sigpulse.com/posts/2026-08-30-rk3588-edge-npu-route-staged/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-rk3588-edge-npu-route-staged/</guid><description>Three Rockchip docker images (6.86 GB total), the two-compiler reality, and the route from workstation GPU to a 6-TOPS edge NPU — staged for real.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every industrial vision project that survives the demo stage meets the same geography problem: the model was born on a workstation GPU, but it must live on a box bolted to a line — no cloud, no RTX, a hard cycle-time budget. In our stack the answer to that geography is Rockchip&amp;#39;s RK3588 edge SoC (6-TOPS NPU class), and this dispatch documents the route to it as it actually stands on the workstation: staged, smoke-tested at the container level, one mile from finished. The rubric as always: what it is → what it does → how to run it → what went wrong (or, honestly, what hasn&amp;#39;t been run yet).&lt;/p&gt;
&lt;h2&gt;The route, end to end&lt;/h2&gt;
&lt;p&gt;Four stations:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Train on the workstation.&lt;/strong&gt; This is the anomalib rig from our &lt;a href=&quot;/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/&quot;&gt;previous field guide&lt;/a&gt; — PatchCore-class models on the 4090D.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Export to ONNX.&lt;/strong&gt; The neutral interchange. The anomalib environment already carries &lt;code&gt;onnx 1.19.1&lt;/code&gt; (and &lt;code&gt;openvino 2025.3.0&lt;/code&gt;, the parallel route to Intel edge hosts — ONNX keeps the model board-agnostic).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Convert and quantize to RKNN, in docker, on x86.&lt;/strong&gt; Rockchip ships the toolkit as an Ubuntu 20.04 image; conversion plus INT8 quantization happens here, on the workstation, not on the target.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-compile the application, deploy to the board.&lt;/strong&gt; The aarch64 cross-compiler images build the host-side app; the board runs the NPU runtime.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;What&amp;#39;s on disk&lt;/h2&gt;
&lt;p&gt;Three Rockchip images, pulled and verified (sizes from &lt;code&gt;docker images&lt;/code&gt;):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Image&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;rk3588-rknn-dev&lt;/td&gt;
&lt;td&gt;3.03 GB&lt;/td&gt;
&lt;td&gt;RKNN toolkit: convert + quantize on x86_64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rk3588-cross-compiler&lt;/td&gt;
&lt;td&gt;3.03 GB&lt;/td&gt;
&lt;td&gt;aarch64 application cross-compile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rk3588-base-cross-compiler&lt;/td&gt;
&lt;td&gt;798 MB&lt;/td&gt;
&lt;td&gt;base layer for leaner build images&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Total staging cost: about 6.86 GB. On 2026-08-30 we smoke-tested the dev image for the first time — no container from these images had ever been run (they sat staged since pull); a fresh container boots and reports Python 3.8.10 inside. The toolchain is real, executable, and was waiting for its model.&lt;/p&gt;
&lt;h2&gt;The mile, walked to its first station — and its wall&lt;/h2&gt;
&lt;p&gt;Since this guide&amp;#39;s first publication we walked the leg, and it split into a completed station and a documented wall:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Station one, done: the ONNX export.&lt;/strong&gt; A resnet18 exported from the anomalib environment via &lt;code&gt;torch.onnx.export(..., opset_version=12)&lt;/code&gt; — 46,733,662 bytes, random weights (the conversion case needs no trained model). The gotcha that cost one failed attempt: torch 2.9 flips the exporter default to the dynamo path, which demands &lt;code&gt;onnxscript&lt;/code&gt; and is not installed here — &lt;code&gt;dynamo=False&lt;/code&gt; pins the legacy exporter and works. Version churn again, at the very first station.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The wall: acquiring the converter itself.&lt;/strong&gt; The RK3588 images carry the deployment-side SDK (&lt;code&gt;/workspace/rknn_app_demo&lt;/code&gt;, &lt;code&gt;rknpu2&lt;/code&gt;) — the &lt;em&gt;conversion&lt;/em&gt; toolkit does not live in them; it distributes as a wheel from the project&amp;#39;s GitHub. From this network, six acquisition routes were tried and all failed, each with its own exact error: &lt;code&gt;git clone&lt;/code&gt; over HTTPS dies on a GnuTLS termination; the GitHub Releases API answers but carries zero attached assets; the repository tree API lists no wheel files at all; &lt;code&gt;raw.githubusercontent&lt;/code&gt; returns 404 for the path; the container cannot reach PyPI; and the domestic PyPI mirror has no such package. That inventory is the finding — the standard toolchain of China&amp;#39;s most popular edge NPU is, from an un-proxied mainland network, effectively airless. (The proxied route exists on this machine; we deliberately do not invoke it for a field guide that must reproduce clean.)&lt;/p&gt;
&lt;p&gt;No &lt;code&gt;.rknn&lt;/code&gt; artifact exists anywhere on disk — a machine-wide check confirms the conversion never happened here. Meanwhile the &lt;em&gt;other&lt;/em&gt; edge route on this fleet actually ran: &lt;a href=&quot;/posts/2026-08-30-mac-mini-cpu-edge-deployment/&quot;&gt;workstation-trained PatchCore running on a Mac mini&amp;#39;s CPU against a Basler camera&lt;/a&gt; — a reminder that &amp;quot;edge&amp;quot; is a requirement, and Apple Silicon was one valid answer while the NPU route waited on its toolkit.&lt;/p&gt;
&lt;h2&gt;The pitfalls (the ones you can schedule around)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Quantization accuracy is a task, not a setting.&lt;/strong&gt; INT8 on a 6-TOPS NPU trades precision for speed by design; every converted model needs a validation pass against its pre-quantization self, and the drop is per-model news. Budget it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Version triple-lock.&lt;/strong&gt; Toolkit (docker tag), on-board runtime, and converted model must agree. The docker route makes the workstation side exactly reproducible; the board side is a discipline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two compilers, one app.&lt;/strong&gt; The cross-compiler images build for aarch64 while the board runs its own runtime; drift between them surfaces as errors that look like anything but compiler drift. Keep a matched pair, note both versions in the deployment log.&lt;/p&gt;
&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;Teams whose model works on the bench and now must survive a line: no network dependency, fixed latency, board-class power and price. If your roadmap is anomalib (or any vision model) → edge NPU, this is the map of the middle country, with the exact images named and the remaining leg marked in red.&lt;/p&gt;
&lt;h2&gt;Replication appendix&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# The staged images (sizes as measured):
docker images | grep rk3588
# rk3588-rknn-dev    20.04   3.03GB
# rk3588-cross-compiler 20.04 3.03GB
# rk3588-base-cross-compiler 20.04 798MB

# Smoke test the conversion environment (first execution, 2026-08-30):
docker run --rm rk3588-rknn-dev:20.04 python3 -V
# Python 3.8.10

# The route&amp;#39;s next mile (not yet run here):
# anomalib ckpt -&amp;gt; export ONNX (onnx 1.19.1 available in anomalib_env)
# -&amp;gt; docker run rk3588-rknn-dev ... rknn conversion + INT8 quantization
# -&amp;gt; cross-compile host app, deploy to RK3588, validate accuracy vs pre-quantization
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the workstation&amp;#39;s docker image list, container smoke run, and the anomalib environment&amp;#39;s pip inventory (onnx 1.19.1, openvino 2025.3.0), inspected 2026-08-30. Third-party, checked: &lt;a href=&quot;https://github.com/airockchip/rknn-toolkit2&quot;&gt;Rockchip&amp;#39;s rknn-toolkit2 repository&lt;/a&gt; (the converter the dev image packages) and &lt;a href=&quot;https://onnx.ai&quot;&gt;onnx.ai&lt;/a&gt; for the interchange standard. Drafted with AI assistance under human editorial direction; the 6-TOPS NPU rating follows Rockchip&amp;#39;s public specification. Numbers from this staging are in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Workstation 2-GPU rig (RTX 4090D 24GB + RTX A4000 16GB) · docker Engine · anomalib_env side (onnx 1.19.1, openvino 2025.3.0). Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-rk3588-edge-npu-route-staged.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-rk3588-edge-npu-route-staged.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>RK3588</category><category>RKNN</category><category>edge deployment</category><category>ONNX</category><category>NPU</category><category>quantization</category></item><item><title>What Is the &apos;Gray Zone&apos; — the Condition Below War and Above Peace? A System Diagnosis, Not a Prediction</title><link>https://sigpulse.com/posts/2026-08-30-strait-gray-zone-system-playbook/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-strait-gray-zone-system-playbook/</guid><description>Normalization, law-enforcement costume, and attrition: three parts of a machine that moves boundaries without ever producing a headline day.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Three numbers, all relays from foreign coverage, opened our Chinese-language column&amp;#39;s diagnosis (2026-08-17, &amp;quot;Who fears the strait&amp;#39;s &amp;#39;gray zone&amp;#39;? Foreign media dissect a new contest with no shot fired&amp;quot;): 55 Chinese-government-ship sightings in June (+83% month-on-month), roughly 200 commercial vessels boarded east of the island, and Taiwan&amp;#39;s first internet-outage drill plus a ten-day resilience exercise. Individually small; connected across a summer of Bloomberg-Reuters-broadsheet coverage density, the pattern. The column&amp;#39;s announcement of method is the piece&amp;#39;s spine: &lt;strong&gt;no &amp;quot;will it happen&amp;quot; — only a systems diagnosis of how the machine is designed and what account each part is computing.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Same double-relay disclosure as our companion piece (&lt;a href=&quot;/posts/2026-08-30-strait-quarantine-vocabulary-playbook/&quot;&gt;the quarantine vocabulary decoder&lt;/a&gt;): the episode layer is Chinese commentary relaying English reporting, tagged [unverified] throughout; the verified layer is the conceptual machinery, anchored to CSIS&amp;#39;s quarantine scenario analysis.&lt;/p&gt;
&lt;h2&gt;The tense changed first&lt;/h2&gt;
&lt;p&gt;Past strait coverage lived in the future tense — whether, when, in what form. The recent crop runs in the present progressive — what happened this week, how many recorded, how this differs from last. The grammar shift is the system-state shift: from a scenario under discussion to a condition being operated. With the state change comes the analysis rule: when a system moves from &amp;quot;being debated&amp;quot; to &amp;quot;running&amp;quot;, the correct unit is no longer &amp;quot;whether&amp;quot; but &lt;strong&gt;each day&amp;#39;s cost and return&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;The three parts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Normalization.&lt;/strong&gt; Presence billed by the incident. Every action stays below every red line — individually negligible, cumulatively rewriting the default. The product of a gray zone is not events; it is the new normal itself. No single day deserves a headline, and every day moves the boundary.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Law-enforcement costume.&lt;/strong&gt; The actions wear uniforms. This is the gear the companion piece dismantles: a blockade is war (declare it and you hand the opponent the trigger); enforcement-framed action leaves the opponent choosing between legally costly force and jurisdiction conceded by default. The machine&amp;#39;s cleverest design: it paints over every one of the opponent&amp;#39;s options.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Attrition.&lt;/strong&gt; The contest is not bravery but endurance. The relays include the allies&amp;#39; ledger — American anxiety about munitions-production replenishment rates, a deleted phrase in Japan&amp;#39;s new defense white paper. Not just the two principals: the alliance system is quietly computing costs too.&lt;/p&gt;
&lt;h2&gt;The two-sided verdict the original insists on&lt;/h2&gt;
&lt;p&gt;Not all foreign commentary buys the frame. Critics call gray-zone and quarantine narratives a systematic amplification — the security-research-media complex&amp;#39;s traffic business, where concretely narrated danger converts to budgets and clicks. The column takes the suspicion seriously: &lt;strong&gt;the concept can be true while the concept&amp;#39;s marketing is exaggerated.&lt;/strong&gt; Its restrained synthesis: change real (figures rising, coverage density growing), intensity narratively amplified (scale limited, definitions unsettled), true state between the two.&lt;/p&gt;
&lt;h2&gt;The Berlin mirror&lt;/h2&gt;
&lt;p&gt;For a long-running sample of the machine, the original reaches back to Cold War Berlin: two superpowers opposing each other for decades at checkpoints, corridors and frequencies — learning to live inside the gray zone. The confrontation itself decided nothing; the reversal arrived from outside it. The mirror&amp;#39;s lesson, plain and cold: what a gray zone consumes is never ships and fuel, but &lt;strong&gt;each side&amp;#39;s definition of the status quo, and the patience that definition requires&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Which returns the question to every player at once: if the no-shot-fired contest becomes permanent, whose system better endures the day-after-day? That question has no trending hashtag — only a calendar. The reader&amp;#39;s position is where the original leaves it: understand which gear each news item belongs to. That is all a decoder owes.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;Concept layer (verified): &lt;a href=&quot;https://features.csis.org/china-quarantine-taiwan/&quot;&gt;CSIS ChinaPower: How China Could Quarantine Taiwan&lt;/a&gt; — the law-enforcement-costume gear; the Berlin checkpoint history is standard archival record. Episode layer (unverified relays, outlet names and dates as cited by the Chinese original): Bloomberg (June 25), Reuters, Irish Times (August 11), Foreign Affairs (&amp;quot;why China waits&amp;quot;), Al Jazeera 101 East, BBC Chinese (outage drill; white-paper deletion; US production concerns); sighting and boarding counts as relayed. Provenance: originally published in Chinese on our WeChat channel on 2026-08-17 (&amp;quot;谁在害怕台海&amp;#39;灰色地带&amp;#39;？外媒拆解一枪未发的新博弈&amp;quot;); drafted with AI assistance under human editorial direction; translated and adapted 2026-08-30, inheriting the original&amp;#39;s disclaimer in full — mechanism analysis and trend observation, no situation judgment or prediction. Editorial line — not a SigPulse measurement; first-party numbers live in the &lt;a href=&quot;/posts/&quot;&gt;dispatches&lt;/a&gt; and the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Editorial desk — no lab hardware; concept anchored to the CSIS quarantine analysis; 2026 episode figures [unverified relays]. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-strait-gray-zone-system-playbook.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-strait-gray-zone-system-playbook.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>geopolitical literacy</category><category>Taiwan Strait</category><category>gray zone</category><category>systems thinking</category><category>media literacy</category></item><item><title>Why Do Analysts Say &apos;Quarantine&apos; Instead of &apos;Blockade&apos; — and What Did Kennedy&apos;s 1962 Word Choice Teach?</title><link>https://sigpulse.com/posts/2026-08-30-strait-quarantine-vocabulary-playbook/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-strait-quarantine-vocabulary-playbook/</guid><description>One word separates an act of war from a law-enforcement action. A vocabulary decoder for cross-border readers, from 1962&apos;s quarantine to today&apos;s coverage.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Spend three months reading English-language Taiwan Strait coverage and one small thing happens: the word &lt;em&gt;quarantine&lt;/em&gt; starts appearing more often — in Bloomberg, in an Irish Times headline, in a Washington Examiner column. It is not a public-health word. It names a maritime condition one sheet of paper away from &lt;em&gt;blockade&lt;/em&gt; yet legally a different machine entirely. Our Chinese-language column ran a decoder piece (2026-08-17, &amp;quot;Why is foreign media suddenly staring at one word: what is the Strait&amp;#39;s &amp;#39;quarantine&amp;#39;?&amp;quot;); this entry translates it for the Playbook audience — people whose plans, families or assets sit across this strait — and anchors the concept layer to sources we could verify ourselves.&lt;/p&gt;
&lt;p&gt;A double-relay honesty note up front: the original is Chinese commentary reading English coverage; the episode specifics below are therefore relays of relays. What we verified directly is the vocabulary machinery — the 1962 archival record and CSIS&amp;#39;s scenario analysis — which is the layer that keeps its value regardless of any particular news cycle.&lt;/p&gt;
&lt;h2&gt;The 1962 lesson&lt;/h2&gt;
&lt;p&gt;Wind the clock back. In the Cuban Missile Crisis, Kennedy&amp;#39;s government made a famous vocabulary decision: refuse &amp;quot;blockade&amp;quot;, choose &amp;quot;quarantine&amp;quot;. The reason was hard law — under the traditional law of war, a blockade is an act of war. Declare one and you hand the adversary a lawful basis to shoot back, and every third-party state must immediately choose sides. &amp;quot;Quarantine&amp;quot;, packaged as maritime enforcement, changed the machine&amp;#39;s legal structure. That crisis&amp;#39;s quarantine lasted about a month and ended in a great-power understanding — the word&amp;#39;s first full career.&lt;/p&gt;
&lt;h2&gt;Two words, two machines&lt;/h2&gt;
&lt;p&gt;Laid side by side, as the original lays them:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blockade&lt;/strong&gt; — an act of war. The opponent gains the right of armed self-defense; third parties must be neutral or belligerent. At the moment of declaration, the trigger is transferred to the other side&amp;#39;s hand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quarantine&lt;/strong&gt; — an enforcement action. Coast-guard boarding, customs supervision, channel control — every step wears an enforcement uniform. The opponent&amp;#39;s dilemma: use force against an &amp;quot;enforcement action&amp;quot; and you lose legal ground before the first shot; don&amp;#39;t use force, and jurisdiction is defaulted, inspection by inspection.&lt;/p&gt;
&lt;p&gt;A blockade hands over the trigger; a quarantine confiscates it. One word of difference is not a difference of tone — it is the entire legal structure of the machine. That is why analysts reach for a public-health word instead of the more obvious military one, and why CSIS&amp;#39;s scenario begins with &amp;quot;enhanced customs inspection rules&amp;quot; announced as domestic administration while scrupulously avoiding both charged words.&lt;/p&gt;
&lt;h2&gt;What the relays said in 2026&lt;/h2&gt;
&lt;p&gt;The Chinese column&amp;#39;s timeline, all relayed and tagged [unverified]: a Bloomberg-reported Taiwanese tabletop exercise simulating maritime quarantine (June 25); an Irish Times headline in August about Taiwan preparing while Beijing has other options; a Foreign Affairs essay titled around &lt;em&gt;why China waits&lt;/em&gt;; an Al Jazeera 101 East special asking whether the Hormuz pattern is being replicated; a Washington Examiner column branding a &amp;quot;Hormuz 2.0&amp;quot;. Relayed numbers: 55 sightings of Chinese government ships around the strait in June (+83% month-on-month) and roughly 200 commercial vessels boarded east of the island. Foreign commentary itself split — one camp reading a slow-motion blockade, another calling the narrative inflation (current frequency and scale as pressure-testing rather than systematic quarantine). The split itself is information: the concept&amp;#39;s boundary is still being fought over, which is part of how the concept establishes itself.&lt;/p&gt;
&lt;h2&gt;Why this belongs in a playbook&lt;/h2&gt;
&lt;p&gt;Vocabulary change is a leading indicator of institutional preparation — when a think-tank term migrates into headlines, it means a cohort of institutions is already rehearsing with the concept: running table tops, drafting contingencies, producing documentaries. For a cross-border reader the practical value is not predicting anything (the original explicitly refuses to, and so do we); it is &lt;strong&gt;reading the news at the right altitude&lt;/strong&gt;. The next time a maritime story leads with coast guard, boarding inspection, or channel control, you will know which gear of which machine you are looking at — and that &amp;quot;customs rules announced as administration&amp;quot; is, in this grammar, the loudest sentence of all.&lt;/p&gt;
&lt;p&gt;A word&amp;#39;s fate is sometimes a sea&amp;#39;s fate. Draw no conclusions yet — that is the discipline this piece teaches, along with the vocabulary.&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;Verified concept layer: &lt;a href=&quot;https://features.csis.org/china-quarantine-taiwan/&quot;&gt;CSIS ChinaPower: How China Could Quarantine Taiwan&lt;/a&gt; and &lt;a href=&quot;https://www.taiwannews.com.tw/news/5887170&quot;&gt;Taiwan News&amp;#39;s summary of the CSIS analysis&lt;/a&gt;; the 1962 quarantine precedent per the &lt;a href=&quot;https://www.jfklibrary.org/learn/about-jfk/jfk-in-history/cuban-missile-crisis&quot;&gt;JFK Library&amp;#39;s Cuban Missile Crisis record&lt;/a&gt;. Episode layer: relays from our Chinese original (Bloomberg, Irish Times, Foreign Affairs, Al Jazeera, Washington Examiner, BBC Chinese — outlet names and dates as cited; URLs not independently located) — all [unverified]. Provenance: originally published in Chinese on our WeChat channel on 2026-08-17 (&amp;quot;为什么外媒突然盯上一个词：台海的&amp;#39;隔离&amp;#39;到底是什么？&amp;quot;); drafted with AI assistance under human editorial direction; translated and adapted 2026-08-30 with the disclaimer inherited: concept literacy and trend analysis, no situation judgment or prediction. Editorial line — not a SigPulse measurement; first-party numbers live in the &lt;a href=&quot;/posts/&quot;&gt;dispatches&lt;/a&gt; and the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Editorial desk — no lab hardware; concept anchored to CSIS analysis and 1962 archival record; 2026 episode details [unverified relays]. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-strait-quarantine-vocabulary-playbook.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-strait-quarantine-vocabulary-playbook.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>geopolitical literacy</category><category>Taiwan Strait</category><category>quarantine</category><category>law of war</category><category>media literacy</category></item><item><title>SAM3 on a Workstation: Docker Door, Source Door, and a Wire Into ComfyUI</title><link>https://sigpulse.com/posts/2026-08-30-sam3-two-deployment-doors/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-30-sam3-two-deployment-doors/</guid><description>Meta&apos;s SAM3 on one workstation: a scripted docker door never pulled, a 131 MB source clone, and a real 533-line ComfyUI node package with saved blueprints.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Segmentation used to be a pointing exercise — click the object, refine the mask. Meta&amp;#39;s SAM 3 (November 2025) replaced the pointing with naming: &lt;strong&gt;Promptable Concept Segmentation&lt;/strong&gt;, where a text or visual concept prompt makes one model of 848 M parameters find, mask, and track that concept across images and video. For industrial vision that interface change is not a nicety — &amp;quot;the defect&amp;quot;, &amp;quot;the weld&amp;quot;, &amp;quot;the part&amp;quot; are concepts, and concepts are now executable. This guide documents SAM3 as it stands on the workstation: staged through two doors, wired into a third.&lt;/p&gt;
&lt;h2&gt;Door one: the docker route (scripted, not walked)&lt;/h2&gt;
&lt;p&gt;A scripted deploy walks the checklist in six steps — environment checks, connectivity verification, docker presence, then the platform-pinned pull of &lt;code&gt;facebookresearch/segment-anything-3:latest&lt;/code&gt; and container startup. What the disk says about depth: the image is not on the machine — the pull was never completed here. This door is &lt;em&gt;prepared&lt;/em&gt;, which is exactly what a script is for: the next operator inherits the checklist instead of the archaeology.&lt;/p&gt;
&lt;h2&gt;Door two: the source route (cloned, complete)&lt;/h2&gt;
&lt;p&gt;The full source tree lives in the workspace — the &lt;code&gt;sam3&lt;/code&gt; package, &lt;code&gt;examples/&lt;/code&gt;, &lt;code&gt;assets/&lt;/code&gt;, and notably &lt;code&gt;README_TRAIN.md&lt;/code&gt;, the training documentation — a 131 MB clone carrying an upstream commit from January 2026. This is the door for &lt;em&gt;changing&lt;/em&gt; the model: fine-tuning on domain imagery, inspecting internals, pinning a commit against upstream drift. Keeping both doors staged is deliberate: the container for reproducibility, the source for control.&lt;/p&gt;
&lt;h2&gt;Door three: the wire into ComfyUI (real code)&lt;/h2&gt;
&lt;p&gt;The most telling artifacts are not the deploy scripts — they are inside the local ComfyUI install: a 533-line custom node package (&lt;code&gt;nodes_sam3.py&lt;/code&gt;) defining four node classes — &lt;code&gt;SAM3_Detect&lt;/code&gt;, &lt;code&gt;SAM3_VideoTrack&lt;/code&gt;, &lt;code&gt;SAM3_TrackPreview&lt;/code&gt;, &lt;code&gt;SAM3_TrackToMask&lt;/code&gt; — plus two saved workflow blueprints, &amp;quot;Image Segmentation (SAM3)&amp;quot; and &amp;quot;Video Segmentation (SAM3)&amp;quot;. That is SAM3 promoted from a standalone tool to a &lt;strong&gt;node in a node-based vision pipeline&lt;/strong&gt;, detect-track-mask exposed as composable stages. What is not on disk: any verified inference output from these nodes — the wiring is real, the first flight is not logged.&lt;/p&gt;
&lt;h2&gt;The honest boundary&lt;/h2&gt;
&lt;p&gt;Depth of practice on this workstation: staged (both routes), wired (ComfyUI nodes and blueprints), not yet run to a recorded result. The 848 M parameter count and benchmark claims are Meta&amp;#39;s published figures, cited as such; nothing in this guide is our inference measurement. When the first real segmentation job runs, this dispatch gets its numbers — the pattern the anomalib guide set and the Isaac Sim guide (a line that stopped at its installer) carries to its honest extreme.&lt;/p&gt;
&lt;p&gt;And the pre-history lives on another machine: in September 2025, the same operator&amp;#39;s Mac mini held a drafted two-stage pipeline — YOLOv8n detects, its first box feeds FastSAM&amp;#39;s box-prompt, mask out, device MPS — with the weights (FastSAM-s, 23,832,055 bytes; yolov8n, 6.5 MB) still on disk and no saved outputs, honestly unverifiable as a run. The paradigm is identical to what this workstation&amp;#39;s SAM3 wiring stages at larger scale: &lt;strong&gt;detector for the rough box, SAM-family for the fine mask&lt;/strong&gt;. Two machines, one idea, converging — the &lt;a href=&quot;/posts/2026-08-30-mac-mini-cpu-edge-deployment/&quot;&gt;Mac mini deployment guide&lt;/a&gt; documents the machine that carried it.&lt;/p&gt;
&lt;h2&gt;Why a factory should care&lt;/h2&gt;
&lt;p&gt;The label layer is what every industrial vision program is missing, and concept-prompted masks are that layer at zero marginal labeling cost. Segmented regions become training ground truth for QC models; they become measurement regions for geometry checks; they pair with the synthetic-data line (still at the planning stage — &lt;a href=&quot;/posts/2026-08-30-isaac-sim-synthetic-data-line/&quot;&gt;that guide&lt;/a&gt; documents where it actually stopped). In the column&amp;#39;s architecture — &lt;a href=&quot;/posts/2026-08-30-anomalib-industrial-defect-detection-field-guide/&quot;&gt;anomalib&lt;/a&gt; for detecting from few samples, SAM3 for &lt;strong&gt;segment to measure&lt;/strong&gt; — this is the leg whose tooling is wired and waiting for its first job.&lt;/p&gt;
&lt;h2&gt;The pitfalls&lt;/h2&gt;
&lt;p&gt;The pull needs its platform flag on mixed-architecture hosts (ours pins linux/amd64 — without it, docker helpfully fetches the wrong architecture and fails at run). Weights and image downloads are scheduled costs, not surprises to discover at deadline. The docker door trades away easy modification — which is exactly why the source door stays open. And in ComfyUI, upstream node updates break saved workflows in quiet ways; the blueprints double as the recovery point.&lt;/p&gt;
&lt;h2&gt;Replication appendix&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# Door one — container route (the scripted checks, abbreviated):
docker pull --platform linux/amd64 facebookresearch/segment-anything-3:latest

# Door two — source route:
# full tree cloned to the workspace: sam3/ package, examples/, assets/, README_TRAIN.md

# Door three — ComfyUI wiring (as found on this workstation):
#   comfy_extras/nodes_sam3.py
#   blueprints/&amp;quot;Image Segmentation (SAM3).json&amp;quot;
#   blueprints/&amp;quot;Video Segmentation (SAM3).json&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;First-party: the workstation&amp;#39;s deploy scripts, source tree, and ComfyUI install artifacts, inspected 2026-08-30 (security note: the deploy scripts were also sanitized this day to remove a credential-handling anti-pattern — the fix is documented, the credential is not). Third-party, checked: &lt;a href=&quot;https://ai.meta.com/blog/segment-anything-model-3/&quot;&gt;Meta&amp;#39;s SAM 3 announcement&lt;/a&gt;, the &lt;a href=&quot;https://github.com/facebookresearch/sam3&quot;&gt;SAM 3 repository&lt;/a&gt;, and the &lt;a href=&quot;https://arxiv.org/abs/2511.16719&quot;&gt;paper (arXiv:2511.16719)&lt;/a&gt; for the PCS task, parameter count and benchmark claims — those are Meta&amp;#39;s figures, cited as published. Drafted with AI assistance under human editorial direction. The staging facts are in the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-30 · Workstation 2-GPU rig (RTX 4090D 24GB + RTX A4000 16GB) · docker Engine · ComfyUI install with custom SAM3 node. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-30-sam3-two-deployment-doors.md&quot;&gt;https://sigpulse.com/posts/2026-08-30-sam3-two-deployment-doors.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>SAM3</category><category>segmentation</category><category>Meta AI</category><category>ComfyUI</category><category>docker</category><category>industrial vision</category></item><item><title>Who Called for Interning Chinese Americans If War Comes — and Why Did Chinese Comment Sections Cheer?</title><link>https://sigpulse.com/posts/2026-08-29-timcast-internment-chinese-americans/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-29-timcast-internment-chinese-americans/</guid><description>Podcast panelists floated internment; Congress condemned it; Chinese comments cheered. A field note on loyalty policing from both directions.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Rumor version first, because the rumor is the story&amp;#39;s engine: &lt;em&gt;US lawmakers said that if war breaks out with China, Chinese Americans should be rounded up.&lt;/em&gt; That is how this traveled inside the Chinese internet in late August 2026 — speaker upgraded from podcast studio to Congress, and the upgraded version met, in the highest-liked comments, not outrage but applause. [unverified — comment quotes below are from public comment-section screenshots captured by our Chinese-language column; no stable public links exist]&lt;/p&gt;
&lt;p&gt;What actually happened is checkable in about four minutes, which is the point of this field note.&lt;/p&gt;
&lt;h2&gt;The clip, the condemnation, and the reply&lt;/h2&gt;
&lt;p&gt;On August 14, 2026, the roundtable of &lt;em&gt;Timcast IRL&lt;/em&gt; — a pro-Trump podcast — ran a hypothetical: if the US–China conflict escalates and America &amp;quot;can&amp;#39;t tell&amp;quot; where Chinese Americans&amp;#39; loyalties lie, what then? Host Ian Crossland&amp;#39;s phrasing, per the Media Matters transcript: interning &amp;quot;all the Chinese Americans we can find&amp;quot; in camps &amp;quot;seems like we should.&amp;quot; Commentators Chris Karr and Kevin Smith joined the exchange; Smith pre-empted the comparison to ICE detention facilities.&lt;/p&gt;
&lt;p&gt;The institutional response ran the other way from the rumor. On August 26, the Congressional Asian Pacific American Caucus — including its Chinese American members, chaired by Rep. Grace Meng — formally condemned the segment as &amp;quot;dangerous and unacceptable&amp;quot; rhetoric. The Japanese American Citizens League named the participants and invoked its own history: in 1942, roughly 120,000 people of Japanese descent were incarcerated in the US, over half of them citizens. Timcast posted its own response video on August 27. SCMP covered the condemnations; AsAm News tracked the discourse.&lt;/p&gt;
&lt;p&gt;So the fact-check is short: &lt;strong&gt;podcast panelists floated it; lawmakers condemned it.&lt;/strong&gt; The version that circulated behind the firewall inverted who said what — and that inversion did not weaken the story&amp;#39;s traction. It strengthened it.&lt;/p&gt;
&lt;h2&gt;The comments: &amp;#39;hurry up&amp;#39;&lt;/h2&gt;
&lt;p&gt;Our Chinese column&amp;#39;s piece on this led with its own author&amp;#39;s expectation — &amp;quot;I assumed the comment section would be furious&amp;quot; — and then quoted what the top comments actually said: &amp;quot;赶紧的&amp;quot; (&lt;em&gt;hurry up&lt;/em&gt;). &amp;quot;Support — let them find out what emigrating to America gets you.&amp;quot; &amp;quot;Recommend executing it immediately.&amp;quot; [unverified]&lt;/p&gt;
&lt;p&gt;One frame for what follows: two audiences that agree on nothing — an American right-wing podcast panel and Chinese nationalist commenters — arrived at the same policy preference. The people it concerns were absent from both conversations. Their status got decided by others, twice, in the same week.&lt;/p&gt;
&lt;p&gt;The column&amp;#39;s own analysis goes one layer down, and it is worth translating precisely: the two sides&amp;#39; logics are symmetric. The podcast&amp;#39;s logic: &lt;em&gt;we cannot verify their loyalty, therefore detain.&lt;/em&gt; The comments&amp;#39; logic: &lt;em&gt;they were disloyal to us by leaving, therefore deserve it.&lt;/em&gt; The same person is failed by both tests simultaneously — suspected of being pro-China in America, condemned as anti-China in China. Two loyalty audits, zero possible passing grades.&lt;/p&gt;
&lt;p&gt;And the historical rhyme the column draws is uncomfortably apt: many of the Americans who nodded along in 1942 had shopped for years at the stores of the people they approved interning. When the moment comes, nobody reads your passport first — inside either system.&lt;/p&gt;
&lt;h2&gt;Why this belongs in a playbook&lt;/h2&gt;
&lt;p&gt;This site measures things; &lt;a href=&quot;/watch/&quot;&gt;China AI Watch&lt;/a&gt; translates the discourse. This entry is neither — it is a field note for people who live across the divide this story is about, whether that&amp;#39;s a green-card queue, a WeChat family group, or a kid on a US campus with a Chinese passport. Three practical readings:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Treat rumor-versions as weather, not climate.&lt;/strong&gt; A podcast exchange is not policy. What is trackable is institutional behavior: who condemned, who amplified, who monetized the amplification. This round, Congress&amp;#39;s Asian American caucus moved within 12 days; the podcast doubled down with a response video. Both are signal; neither is a law.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Your documentation is the only thing both systems agree to read.&lt;/strong&gt; The 1942 precedent is cited by both sides of this exchange for a reason: when loyalty becomes a presumption rather than a question, the people who suffered were overwhelmingly ordinary, documented, and loyal — it did not matter. The practical corollary is grim but real: keep records, keep status current, and treat &amp;quot;it can&amp;#39;t happen here&amp;quot; from &lt;em&gt;either&lt;/em&gt; country as a claim requiring evidence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch the normalization metric, not the outrage metric.&lt;/strong&gt; The podcast segment is a one-day story; the comment section is the load-bearing fact. A society where &amp;quot;detain them&amp;quot; is a zero-cost, high-like utterance — aimed outward at emigrants today, available for any labeled group tomorrow — has already completed the hardest part of the 1942 sequence. That was true of America in 1942 and it is what the column is actually warning about in the Chinese comments. The sentence that closes its piece is the one to keep: barbed wire doesn&amp;#39;t start with barbed wire; it starts with sentences like &amp;quot;they deserve it.&amp;quot;&lt;/p&gt;
&lt;h2&gt;Sources and method&lt;/h2&gt;
&lt;p&gt;Editorial fact-check, no lab hardware. Every quote and date above was checked against: the Media Matters transcript of the August 14 episode; the CAPAC press release of August 26; the JACL statement; SCMP&amp;#39;s August 26 report; AsAm News&amp;#39; August 26 coverage; and Timcast&amp;#39;s August 27 response video. The 1942 figures (about 120,000 incarcerated, over half citizens) follow the standard count used by JACL and US government archives. Chinese comment-section quotes are from screenshots captured by our Chinese-language original — they carry [unverified] tags throughout and no stable public links exist.&lt;/p&gt;
&lt;p&gt;Provenance: originally published in Chinese on our WeChat channel on 2026-08-29 (&amp;quot;美国主播提议&amp;#39;关押&amp;#39;华裔，谁在评论区喊&amp;#39;赶紧的&amp;#39;？&amp;quot;); drafted with AI assistance under human editorial direction; translated and expanded to English the same day. This is reported commentary — not a SigPulse measurement; our first-party numbers live in the &lt;a href=&quot;/posts/&quot;&gt;dispatches&lt;/a&gt; and the &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-29 · Editorial desk — no lab hardware; 6 primary sources cross-checked. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-29-timcast-internment-chinese-americans.md&quot;&gt;https://sigpulse.com/posts/2026-08-29-timcast-internment-chinese-americans.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>Chinese Americans</category><category>US-China relations</category><category>internment</category><category>media literacy</category></item><item><title>How Fast Does a Chat Command Become a Finished AI Image? Inside the One-Man FLUX Workshop</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep1-image-workshop/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep1-image-workshop/</guid><description>Episode 1: a chat phrase triggers FLUX on a 16 GB GPU and the image posts itself back — 10 logged runs: 50 s fast, 66–112 s warm, 330–379 s cold.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 1 of One Man One Legion&lt;/a&gt; — the series map lives there. This is the image workshop from that map, measured.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Type a phrase in a chat group — in Chinese, any of the trigger words for &amp;quot;draw&amp;quot;, &amp;quot;generate an image&amp;quot;, or &amp;quot;face-lock&amp;quot; — and the legion&amp;#39;s image workshop takes over. An agent parses the request, launches a factory script with the right flags, and a single RTX A4000 16 GB renders it through ComfyUI&amp;#39;s HTTP API (no web UI involved). When the file lands, the script itself posts the photo back to the group through the messaging bot API. The measured cost across ten logged runs: 50 seconds on the fast profile, 66–112 s warm on the full profile, and 330–379 s when the workshop is cold. Fourteen finished images sit on disk. This episode documents the whole path, from the command&amp;#39;s arrival to the photo&amp;#39;s delivery, with the run table below as its evidence.&lt;/p&gt;
&lt;h2&gt;The command path: agent launches, script delivers&lt;/h2&gt;
&lt;p&gt;The division of labor is the interesting part. The agent that hears the command does three things only: pick the profile, collect an optional face reference (the latest photo in the inbound folder), and launch &lt;code&gt;flux_image_factory.py&lt;/code&gt; in the background — control returns in 0 s while the GPU works. Delivery is deliberately NOT the agent&amp;#39;s job: the script calls the bot API itself and the photo appears in the group without any agent involvement. Heavy work goes through a direct pipe; only judgment goes through the model chain.&lt;/p&gt;
&lt;h2&gt;Two profiles, pinned constants&lt;/h2&gt;
&lt;p&gt;The factory script ships a two-profile strategy matrix, because the workshop&amp;#39;s images serve two different masters — quick iteration, or full-resolution plates that later feed video:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Profile&lt;/th&gt;
&lt;th&gt;Resolution&lt;/th&gt;
&lt;th&gt;Steps&lt;/th&gt;
&lt;th&gt;Guidance&lt;/th&gt;
&lt;th&gt;PuLID weight&lt;/th&gt;
&lt;th&gt;Logged as&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Fast (i2v seed)&lt;/td&gt;
&lt;td&gt;1280×720&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;3.5&lt;/td&gt;
&lt;td&gt;0.65&lt;/td&gt;
&lt;td&gt;&amp;quot;1280×720/20步/快&amp;quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full (art)&lt;/td&gt;
&lt;td&gt;1920×1080&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;5.5&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;&amp;quot;1920×1080/28步/满血&amp;quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;CFG is pinned to 1.0 — a FLUX-specific constant the script&amp;#39;s author marked as an iron rule after experimentation. Face-lock runs pipe the reference photo through PuLID v0.9.1; the prompt then describes action, not appearance, because the reference image owns the face.&lt;/p&gt;
&lt;h2&gt;The ten runs, as logged&lt;/h2&gt;
&lt;p&gt;Every completed run writes a job log with the mode line, seed, and wall time. The full corpus:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Resolution / steps&lt;/th&gt;
&lt;th&gt;Face-lock&lt;/th&gt;
&lt;th&gt;Wall time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;1280×720 / 20&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;50 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;1280×720 / 20&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;50 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;1280×720 / 20&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;50 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Full, cold open&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;379 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Full, cold open&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;330 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Full, warm&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;68 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Full, warm&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;66 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Full, warm&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;66 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Full, warm&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;66 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Full, art&lt;/td&gt;
&lt;td&gt;1920×1080 / 28&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;112 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Runs 4 and 5 opened their session cluster — the same settings then settle to 66 s (runs 7–9). That is the workshop&amp;#39;s real cost structure: the first image of the night costs ~6 minutes (the fp8 FLUX stack, ~11 GB of weights, loading into VRAM), every image after costs about a minute. Face-locking adds nothing measurable at this sample size: 66 s versus 68 s warm.&lt;/p&gt;
&lt;h2&gt;Craft that only shows up in the script&lt;/h2&gt;
&lt;p&gt;Two details worth stealing. The i2v mode rewrites prompts before rendering: it strips quality-word filler (&amp;quot;8k&amp;quot;, &amp;quot;masterpiece&amp;quot;, &amp;quot;hyperrealistic&amp;quot;) and converts dynamic verbs into static tension — &amp;quot;running fast&amp;quot; becomes &amp;quot;standing alert, coiled energy&amp;quot;, &amp;quot;diving&amp;quot; becomes &amp;quot;suspended mid-motion, controlled&amp;quot; — because image-to-video models smear fast motion, and a base plate that already implies motion animates poorly. And the model files are pinned by exact filename in the script (flux1-dev-fp8, clip_l, t5xxl_fp8, the ae VAE, pulid_flux v0.9.1), which is what makes a run reproducible two years later: same weights, same seeds, same seconds.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Ten runs is a log corpus, not a benchmark: no throughput-under-load numbers, no quality scoring, and the 50 s fast-path runs share one prompt family. Cold-load times are session-opening observations (n=2), not a controlled cold-start benchmark. All timings are wall-clock from the workshop&amp;#39;s own logs, single GPU, queue-dependent.&lt;/p&gt;
&lt;p&gt;Primary sources: the workshop&amp;#39;s job logs (10 files, mode line + seed + wall time each, 2026-08), the factory script&amp;#39;s profile matrix and prompt-rewrite rules as read from source, and 14 finished images on disk.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from the 2026-08 log corpus. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Episode 2 continues next-door in the video workshop — the InfiniteTalk dual-GPU dispatch already covers its measured deep end: &lt;a href=&quot;/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/&quot;&gt;Can You Run InfiniteTalk on Two Consumer GPUs?&lt;/a&gt;. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · RTX A4000 16 GB workstation GPU · ComfyUI driven API-first (workflow JSON over HTTP, no web UI) · every timing from the workshop&apos;s own job logs, 2026-08. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep1-image-workshop.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep1-image-workshop.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>FLUX</category><category>ComfyUI</category><category>PuLID</category><category>Text-to-Image</category><category>AI Agents</category></item><item><title>Whose Name Goes on the Article? Five Bylines, One Human</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep10-pen-names/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep10-pen-names/</guid><description>E10: the byline rack — 33 signed files across five genre pen names, one column carrying all five, and a sixth persona that never shipped a word.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 10 of One Man One Legion&lt;/a&gt; — the fifth workshop to open its doors, between &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;the text factory&lt;/a&gt; and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep11-quality-gate/&quot;&gt;its quality gate&lt;/a&gt;. Episode 9 counted the staff; episode 11 counted the gates. This one is about the names on the door.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The main account runs a column called AI-era survival guide. Ten of its surviving articles carry byline lines — under &lt;strong&gt;five different names&lt;/strong&gt;. A reader loyal to that one column has been addressed by five different &amp;quot;writers.&amp;quot; The operation has one human in it. That forces this episode&amp;#39;s question: &lt;strong&gt;when the writing staff is made of models, what is a byline actually for?&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The count was wrong, instructively&lt;/h2&gt;
&lt;p&gt;Episode 9 described &amp;quot;four emotion-typed pen names, plus a fifth configured but never shipped.&amp;quot; Rereading the standard for this episode broke that count in two ways, both worth the correction.&lt;/p&gt;
&lt;p&gt;First, the rack holds &lt;strong&gt;five names, not four&lt;/strong&gt;. The workflow standard&amp;#39;s push step maps five writing genres to five bylines: type A (event stories) signs as Lu Shi, B (industry teardowns) as Shen Jianwei, C (person stories) as Lin Shu, D (suspense recaps) as Qin Yin, E (system diagnosis) as Lingche. Episode 9 missed Lingche entirely — the newest name, added to the standard in v9.2 on July 21 — and the miss is not small: with 8 signed articles Lingche out-publishes Lin Shu and Qin Yin combined (2 + 2 = 4).&lt;/p&gt;
&lt;p&gt;Second, the letters are &lt;strong&gt;genres, not emotions&lt;/strong&gt;. The emotional palette — fear, FOMO, anger, tension-plus-awe — is chosen separately, from a rotating topic-cluster table. What the byline encodes is the kind of read, not the kind of feeling.&lt;/p&gt;
&lt;h2&gt;Thirty-three bylines, recounted today&lt;/h2&gt;
&lt;p&gt;Recounting byline lines across the archive: &lt;strong&gt;Lu Shi 12, Shen Jianwei 9, Lingche 8, Lin Shu 2, Qin Yin 2 — 33 signed files&lt;/strong&gt; (the count moves; this is today&amp;#39;s). The mandatory format since Aug 17 — &amp;quot;the only legal format,&amp;quot; user-accepted that day — births every article with a centered gray line, pen name | column. One daily file from Aug 7 already carried it.&lt;/p&gt;
&lt;p&gt;The columns confirm the genre mapping in practice: every teardown column — industry teardown, frontier teardown, tech teardown — belongs to Shen Jianwei; hot-pursuit and social-lens pieces belong to Lu Shi. And the survival-guide column carries all five names across its 10 pieces — 5 Lu Shi, 2 Shen Jianwei, 1 each Lin Shu, Qin Yin, Lingche — because the byline follows genre, never column. Today&amp;#39;s three articles shipped under three different names.&lt;/p&gt;
&lt;h2&gt;A name is a bundle of prohibitions&lt;/h2&gt;
&lt;p&gt;The personas differ less by biography than by what they may not do. Lingche signs with a tagline — 凌彻｜组织病理学家，切开给你看机制, &amp;quot;Lingche, organizational pathologist: cuts it open and shows you the mechanism&amp;quot; — and his rules are the strictest: a 4-step structure (anchor, evidence, mechanism, takeaway), a terminology-rotation ban that forbids reusing the same step headings or the same analytical terms in consecutive pieces, and a &lt;strong&gt;9-word banned list that outlaws emotion itself&lt;/strong&gt; — lying flat, rivers of blood, doomsday arriving, capitalist conspiracy. The A-D path is the opposite discipline: a friend&amp;#39;s voice (&amp;quot;I feel,&amp;quot; &amp;quot;I&amp;#39;m not sure&amp;quot;), short bursts, an emotion micro-pulse every 300-500 characters — cortisol pressure, then dopamine release. One name bans the very ammunition the others are required to carry.&lt;/p&gt;
&lt;p&gt;There is even a trigger word: type the four characters for &amp;quot;emotion manipulation&amp;quot; to the agent and the factory switches into a separate, faster writing mode — its own doc, v1.0 from July 29, skipping the standard nine-step flow — built for workplace-anxiety topics. That doc names its own trade honestly: single-cause narrative, selective presentation, and, kept even there, three zero-tolerance lines — core numbers must be real, legal safety holds, three images minimum.&lt;/p&gt;
&lt;h2&gt;The stamp layer&lt;/h2&gt;
&lt;p&gt;Since July 26 the push step includes a formal signature check. In the surviving code it lands as an AUTHOR constant at the top of each push script — 11 scripts carry one: Lu Shi 3, Shen Jianwei 3, Lingche 3, Qin Yin 2, &lt;strong&gt;Lin Shu 0&lt;/strong&gt;. The person-story genre is the rack&amp;#39;s rarest; its writer has no surviving script of his own.&lt;/p&gt;
&lt;h2&gt;The empty slot&lt;/h2&gt;
&lt;p&gt;The sixth name, Zhixing Xiaoya, is not this rack&amp;#39;s missing member — she is the sole byline of a &lt;strong&gt;second account&lt;/strong&gt;, with her own config file and a full stock-commentary style doc effective July 19 (empathy hook, suspense question, surprise answer, plain-language numbers, verdict first). A search of the entire workspace finds her byline on zero files. Configured is not shipped; this series keeps relearning it.&lt;/p&gt;
&lt;p&gt;So what is a byline for, here? Not identity — a &lt;strong&gt;genre contract&lt;/strong&gt;. The name tells the reader what kind of machine wrote the piece: this one will cut open a system, that one will tell you a person&amp;#39;s story. It is checked at push time like a stamp, and the rack keeps an empty slot as a reminder that a persona exists only when it has signed something.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next workshop over: the analyst&amp;#39;s desk. &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace wechat-editor-team · evidence = WORKFLOW-STANDARD-V3.md (v9.2: genre-byline mapping, E-type guide, signature check), EMOTION-CONTROL-WRITING.md (v1.0), ARTICLE-FORMAT-TEMPLATE.md (2026-08-17), STOCK-ARTICLE-V2.md (2026-07-19), byline grep recount across articles/ daily/ archive/ plus push-script AUTHOR constants, 2026-08-27 · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep10-pen-names.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep10-pen-names.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Content Pipeline</category><category>WeChat</category><category>Writing</category></item><item><title>Who Checks the Checkers? 313 Lines of Gate, 61 Stamps, One Empty Column</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep11-quality-gate/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep11-quality-gate/</guid><description>E11: inside the QC gate — 313 lines of regex scans and model judgments; the 102-record failure log shows what each layer caught, and what never got written.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 11 of One Man One Legion&lt;/a&gt; — third stop in the workshops arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;the text factory&lt;/a&gt;, which ended on a promise: the QC gates get their own episode. This is it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A factory where nobody reads every draft still has to answer one question: &lt;strong&gt;when you hand quality control to machines, which checks do you give to a regex and which to a model?&lt;/strong&gt; The text factory&amp;#39;s gate has now run long enough — 102 logged failure rounds across 40 days — that the logs, not the design docs, can rule on the split.&lt;/p&gt;
&lt;h2&gt;One file, two kinds of trust&lt;/h2&gt;
&lt;p&gt;The gate is a single Python script, &lt;code&gt;utils/quality-gate.py&lt;/code&gt; — v9.0, 313 lines, an exit code of 0 that unlocks the push and 1 that means red light. Inside, the split is clean. Seven machine scans handle everything countable: unique source domains, images on the platform CDN, banned words, title keywords, stripped text length, the presence of a disclaimer and a counter-view. Four AI judgments handle everything perceivable: an information gap in the first 100 characters, argument ammo matching one of four emotion clusters, conversational warmth, an ending that asks instead of concludes. The regex layer runs deterministically; the model layer answers a printed checklist. Each layer is trusting a different kind of machine for a different kind of promise.&lt;/p&gt;
&lt;h2&gt;What the regex layer actually caught&lt;/h2&gt;
&lt;p&gt;The failure log — every red-light round since July 18 — breaks down by rule, and the ranking is not what the rule list would predict:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rule&lt;/th&gt;
&lt;th&gt;Red lights&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Legal safety (counter-view + disclaimer)&lt;/td&gt;
&lt;td&gt;61&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Title pain word in first 15 chars&lt;/td&gt;
&lt;td&gt;54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platform images (≥3)&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Source domains (≥3)&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Banned / high-risk words&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quotable-line density&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stripped length (&amp;gt;500)&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The top catch, 61 legal-safety red lights, splits into 54 rounds missing a counter-view and 25 missing a disclaimer (some rounds failed both). The banned-word rounds caught 12 forbidden and 11 high-risk occurrences — against a 12-word and a 5-word list, nearly a hit per entry. Read that closely: the most-used regex check is a &lt;em&gt;presence check of a judgment-shaped thing&lt;/em&gt;. The gate cannot tell whether your counter-argument is any good — only that phrases like &amp;quot;另一方面&amp;quot; exist somewhere in the text. The second insight sits at the bottom of the table: length fired 0 times and quotable density twice. Language models never write too short and never skip question marks. Two of the seven rules police problems the technology has outgrown; the gate carries them like appendix organs.&lt;/p&gt;
&lt;h2&gt;Taste, encoded as dictionaries&lt;/h2&gt;
&lt;p&gt;The word lists are taste compressed into data: a 12-word banned list (&amp;quot;shocking,&amp;quot; &amp;quot;let&amp;#39;s all wait and see&amp;quot;), a 5-word high-risk list (fraud, scam, staged), a 69-entry pain dictionary for the first 15 title characters — 裁员, 封杀, 返贫, 哭了 — and, for legal safety, 8 disclaimer markers against 15 counter-view markers. Three of those 15 are fossils: 英国政府说, 英方, 英方角度 — &amp;quot;the British government says,&amp;quot; &amp;quot;the British side&amp;quot; — words that could only have come from the British Steel piece of July 18. When that article lacked a counter-view, the fix added its vocabulary to the &lt;em&gt;general&lt;/em&gt; gate, where it calcified. The ratchet that promotes a thrice-repeated error into a hard rule turns scars into law, but each scar keeps its shape — the dictionary doubles as an archaeological record. Before every run the gate also prints its own top-3 historical failure rules. It reads its scars before judging new text.&lt;/p&gt;
&lt;h2&gt;The stamp math, and the honest crack&lt;/h2&gt;
&lt;p&gt;61 passing stamps survive — 46 in &lt;code&gt;articles/&lt;/code&gt;, 15 in &lt;code&gt;daily/&lt;/code&gt;. The rounds read 42 first-try passes, 8 on round two, 8 on round three, 3 needing a fourth: a 69% one-shot rate for a machine staff. And the crack I went looking for: the lessons logger has a column for AI issues. Across all 102 records it was never written — 0 entries. &lt;code&gt;record_issues&lt;/code&gt; is called with machine findings only, so the four model judgments happen and evaporate. The learning loop only learns from the regex half of the gate; the ratchet turns on one wheel. Enforcement itself is narrower than this series last claimed: the 13 per-article push scripts in the daily pipeline do open the stamp first and refuse without it, but the 24 older push scripts in the project root predate the gate and never ask. A physical key is only physical where the lock was installed — &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;episode 9&lt;/a&gt; called it the physical key, and this is the correction.&lt;/p&gt;
&lt;p&gt;The lesson travels beyond one WeChat pipeline. Countable things to regex, perceivable things to a model — and log both halves, or your factory remembers only half of what it knows. A sibling script in the same workshop, 564 lines, takes the opposite bet on short-video copy: a 100-point score instead of a binary gate. Same question, two philosophies, both still standing.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The next workshop episode heads for the video floor. &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace wechat-editor-team · evidence = utils/quality-gate.py (v9.0, 313 lines), quality-lessons.jsonl (102 records, 2026-07-18 to 2026-08-27), 61 .gate.json stamps (46 articles/ + 15 daily/), 13 gate-checking push scripts in daily/, WORKFLOW-STANDARD-V3.md · counted 2026-08-27 · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep11-quality-gate.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep11-quality-gate.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Quality Control</category><category>Content Pipeline</category><category>Automation</category></item><item><title>The Analyst Desk: 85 Points of Fact, ±13 of Feeling, One Dead Cron</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep12-analyst-desk/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep12-analyst-desk/</guid><description>E12: the shut-down analyst desk — six agents, four signal layers, a score of 85% fact plus ±13 of feeling, and dual-source gates that verified 29/29 points.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 12 of One Man One Legion&lt;/a&gt; — sixth stop in the workshops arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;the text factory&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep11-quality-gate/&quot;&gt;the quality gate&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep7-video-workshop/&quot;&gt;the video workshop&lt;/a&gt; and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep10-pen-names/&quot;&gt;the pen-name rack&lt;/a&gt;. This one is a ruin: the fleet&amp;#39;s analyst desk is the only workshop that got switched off.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every other workshop in this series is still warm. The analyst desk is not. Its cron job — &amp;quot;US stock analysis, 07:00, Monday to Saturday&amp;quot; — last ran on July 23, 2026, exited with an error, and went dark in &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep19-cron-contraction/&quot;&gt;the great contraction&lt;/a&gt;. What it left behind answers a question none of the other workshops face: &lt;strong&gt;when a legion of AI agents rates a stock, how many points of the score come from fact, and how many from feeling?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The desk&amp;#39;s answer was an architecture: 85 points of fact, at most ±13 of feeling, and a gate under both.&lt;/p&gt;
&lt;h2&gt;Four months from pressure test to production&lt;/h2&gt;
&lt;p&gt;The desk started, like everything in this fleet, with a refusal to guess. On March 13, 2026, a nine-source pressure test probed where financial data actually comes from: SEC EDGAR returned 403 until the request declared a User-Agent; Singapore&amp;#39;s SGX sat behind Akamai with no solution at five stars of difficulty; China&amp;#39;s CNINFO was reachable but needed HTML parsing. Two days later a monitoring MVP passed 10/10 tests — RSS parsing, deduplication, a bulk insert of 100 records in 0.00 seconds, 5 concurrent instances. By March 16 the six-agent framework was declared production-ready: a Lead agent over Monitor, Delivery, and three parallel analysts (financial, industry, quant), with shared memory and a five-star rating map — 85–100 strong buy, down to 0–39 strong sell.&lt;/p&gt;
&lt;h2&gt;The formula: 85 fixed, ±13 movable&lt;/h2&gt;
&lt;p&gt;The score engine, fixed in the V4 SOP of June 21, is arithmetic rather than vibes:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Weight&lt;/th&gt;
&lt;th&gt;What lives inside&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Financial health&lt;/td&gt;
&lt;td&gt;35%&lt;/td&gt;
&lt;td&gt;8 factors: revenue growth 15, profit growth 15, net margin 12, gross margin 10, FCF 12, leverage 12, cash-flow quality 12, ROE 12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Industry strength&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;rank percentile 30, growth rank 25, cycle 25, size 20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quant quality&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;base 70 + trend/CAGR/analyst/PEG adjustments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Corrections&lt;/td&gt;
&lt;td&gt;±13&lt;/td&gt;
&lt;td&gt;peers ±5, news ±5, community ±3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The stress suite verified its own arithmetic: one test confirmed the three weights sum to 0.85 — exactly 15% left for corrections — and another pinned the envelope, 85+13=98 clamped at 100, 85−13=72. On real tickers the corrections stayed small and one-sided: six mega-caps earned news adjustments of +3, +1, +2, +2, 0 and 0, while an extreme-negative scenario hit the −5 floor immediately and the extreme-positive case managed only +3, with the harness annotating — in a note it wrote about its own scoring — that the positive side deserved to bite harder. Bad news outweighs good by design.&lt;/p&gt;
&lt;h2&gt;The cage around feeling&lt;/h2&gt;
&lt;p&gt;The correction layer is where the desk did its most careful engineering. Community sentiment could only touch the score on consensus: a Reddit post qualifies as high-signal at 100+ upvotes, ±3 movement requires multiple subreddets and Hacker News agreeing on direction, Polymarket divergence is worth ±2 — and a single viral post from one source is &lt;em&gt;annotated but never scores&lt;/em&gt;. Fact and opinion were separated in writing: yfinance numbers are facts; Reddit and HN are opinions, cited with subreddit, username, and upvote count. When news said −5 and community said +3, the suite&amp;#39;s conflict test ruled −2: news outranks mood.&lt;/p&gt;
&lt;h2&gt;The gate under the numbers&lt;/h2&gt;
&lt;p&gt;Every number that reached readers passed a verification step the SOP marks &amp;quot;automatic, cannot be skipped,&amp;quot; exit code 0 or nothing ships. Its three surviving generations show it evolving: June 17 re-pulled all 29 data points in a 29/29 pass with 0 errors; June 27 logged claimed-vs-actual pairs line by line (NVDA close 192.53, −1.64% day, −8.62% week — every pair equal); July 12 upgraded to dual-source cross-validation, checking 7 tickers against yfinance and Sina quotes independently.&lt;/p&gt;
&lt;p&gt;On top of the gate sat a voice: a July 19 style guide transcribed from a human analyst&amp;#39;s formula — empathy hook, suspense, surprise answer, plain-language metaphors, explicit verdict — banning price numbers from the first paragraph. Its first product shipped July 17: &amp;quot;Four Ways to Die,&amp;quot; a 5,290-character piece on one chip crash killing Asian markets while only tickling America.&lt;/p&gt;
&lt;h2&gt;An honest inventory&lt;/h2&gt;
&lt;p&gt;The ruin counts cleanly: 4 analysis posts, 6 push records (June 29 to July 20), 3 verification files, a filings library of 49 files under 13 tickers — against a 14-company config where AMAT and Enphase were listed but never downloaded, and Apple sits on disk but off the list. The A-share picker exists only as a row in the March pressure test. The four WARN items in the final stress run name the fragility plainly: one ticker got zero news when SearXNG hiccupped, and an old table was missing columns.&lt;/p&gt;
&lt;p&gt;A desk that scores 85 points of fact, cages ±13 of feeling, and gates every number — and still ends up dark — is not a failed workshop. It is a complete one, waiting.&lt;/p&gt;
&lt;p&gt;*[Next in the fleet: the audio workshop. Back to the anchor, &lt;strong&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;One Man, One Legion&lt;/a&gt;&lt;/strong&gt;.]*&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;All measurements above were taken read-only from the workstation workspace on 2026-08-27; the evidence trail lives in the SOP, stress-test JSON, verification records, and cron state cited in the frontmatter. The Telegram group identifier and push credentials are deliberately absent.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace · evidence = US_STOCK_ANALYSIS_SOP.md (V4.1a, 2026-06-21), skills/sec-filing-monitor (MULTI_AGENT_FRAMEWORK_SUMMARY.md V2.0, 2026-03-16; STRESS_TEST_REPORT_20260314.md 10/10; v4_stress_test_report.json 26 pass / 4 warn / 0 fail, 2026-06-17), STOCK_REPORT_PRESSURE_TEST.md (9 sources, 2026-03-13), filings/US (49 files, 13 tickers) vs config/companies_mvp.json (14 companies, 5+5+4), wechat-editor-team/daily stock artifacts (4 analysis html + 6 push records + 3 verifications, 2026-06-17 to 07-20), cron jobs.json (US-stock job last run 2026-07-23, error, disabled) · counted 2026-08-27 · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep12-analyst-desk.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep12-analyst-desk.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Multi-Agent Systems</category><category>Stock Analysis</category><category>Automation</category></item><item><title>The Audio Workshop: Six Voices on a Shelf, One Episode That Never Shipped</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep13-audio-workshop/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep13-audio-workshop/</guid><description>E13: the audio workshop — six reference voices, a 172-line IndexTTS-2 pipeline, a lost episode, and three podcast crons gone dark.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 13 of One Man One Legion&lt;/a&gt; — seventh stop in the workshops arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;the text factory&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep11-quality-gate/&quot;&gt;the quality gate&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep7-video-workshop/&quot;&gt;the video workshop&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep10-pen-names/&quot;&gt;the pen-name rack&lt;/a&gt; and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep12-analyst-desk/&quot;&gt;the analyst desk&lt;/a&gt;. This room neither prints nor renders. It speaks.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In a folder called &lt;code&gt;voice-library&lt;/code&gt; sit six reference voices and one script — seven files that are the entire casting department of a radio station. Four voices are designed composite hosts: an AI female, an AI male, an English female, an English male. Two are lifted from real speakers. One of those two — a well-known lecturer&amp;#39;s — became the voice of the fleet&amp;#39;s daily news podcast. The house rule on clones is absolute and worth more than the code around it: &lt;strong&gt;use the voice, never the name.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That shelf answers this episode&amp;#39;s question: what does it take for a one-man operation to run a podcast station, down to manufacturing the host?&lt;/p&gt;
&lt;h2&gt;From rented voices to owned ones&lt;/h2&gt;
&lt;p&gt;The station didn&amp;#39;t start with clones. Mid-June artifacts — a 54 KB essay MP3 from June 14, more essay readings through June 25 — were voiced by Microsoft&amp;#39;s free edge-tts. Good enough for reading an essay aloud, but the control surface fought back: a strategy note in one of two successive test scripts records pause tags being severed by the library&amp;#39;s byte-length text splitting.&lt;/p&gt;
&lt;p&gt;The break came on one evening in July. IndexTTS-2 was checked out to a scratch directory with 8.3 GB of checkpoints — inside it, tellingly, sits a Qwen 0.6B model: this voice engine carries its own small language model. File timestamps tell the rest: a first smoke test at 08:13 on July 1, then an evening burst — a lecture recording pulled from YouTube, reference segments cut from it, a side-by-side comparison file against edge-tts at 18:49, and a first clone sample at 18:56. Download to cloned host: one working day.&lt;/p&gt;
&lt;h2&gt;The studio is 172 lines&lt;/h2&gt;
&lt;p&gt;The production tool, &lt;code&gt;podcast_tts_indextts.py&lt;/code&gt;, is 172 lines of Python, and every design decision in it is a scar. It still reads edge-tts-format scripts — complete with rate and pitch fields it &lt;em&gt;knows&lt;/em&gt; are ignored — so the upstream script generator never had to change. It monkey-patches &lt;code&gt;torchaudio.save&lt;/code&gt; with soundfile because the upstream save path broke; the workaround outlived the bug report and became permanent production code. It synthesizes segment by segment, times each one, and writes a timeline JSON next to every 192 kbps MP3.&lt;/p&gt;
&lt;p&gt;That timeline file is the quiet key to the whole workshop. One artifact serves two factories: it proves the audio&amp;#39;s structure, and it later became the subtitle spine of &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep7-video-workshop/&quot;&gt;the video workshop&amp;#39;s&lt;/a&gt; vertical-video generator — the audio room&amp;#39;s DNA is visible in a room built weeks later.&lt;/p&gt;
&lt;h2&gt;Five episodes, one missing&lt;/h2&gt;
&lt;p&gt;The output ledger is short and checkable. First episode: June 28, 36 segments, 201 seconds, produced alongside three video variants. Then a July cadence: July 20 with 16 segments and 1584 characters (393 seconds, plus a compact cut), July 21 with 27 segments and 1378 characters (351 seconds, plus a 42 MB video), July 26 with 23 segments and 1496 characters (385 seconds, audio only).&lt;/p&gt;
&lt;p&gt;The July 4 episode is the one that got away. Produced, 5 minutes 50 seconds, 2 MB — and never delivered: the proxy tunnel to Telegram failed with an SSL error while ordinary domestic sites stayed reachable. A pending-delivery note was written — waiting for the proxy to recover — the proxy recovered, and the file is simply gone from the workspace now. The note remains as a tombstone — the failure lived in delivery infrastructure, not the studio. And unlike &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep12-analyst-desk/&quot;&gt;the analyst desk&lt;/a&gt;, this workshop&amp;#39;s shutdown was scheduled, not accidental: all three podcast cron jobs (the news-podcast production run and two daily AI-trend shows) sit disabled in &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep19-cron-contraction/&quot;&gt;the great contraction&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;One more number for the honesty file: the video factory&amp;#39;s workflow doc, dated August 17, describes its generator as 344 lines. The file on disk today measures 445. Docs drift; &lt;code&gt;wc -l&lt;/code&gt; doesn&amp;#39;t.&lt;/p&gt;
&lt;p&gt;The station is off the air, but the shelf stays stocked and the studio stays warm. Restart cost is one script JSON away. What the room proves is narrower and stranger than the other workshops: broadcast voice — the last component of radio that seems irreducibly &lt;em&gt;live&lt;/em&gt; — decomposes into one reference clip, one GPU evening, and 172 lines of glue.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — next in the engine-room arc: the model stable.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace · evidence = voice-library/ (6 reference wavs + podcast_tts.py, counted 2026-08-27), podcast_tts_indextts.py (wc -l = 172), /tmp/index-tts checkpoints (du = 8.3 GB, contains qwen0.6bemo4-merge), podcast_script_* and podcast_timeline_* JSONs for 06-28 / 07-20 / 07-21 / 07-26 (segment counts, end timestamps, character counts, parsed 2026-08-27), essay-era MP3s (06-14 to 06-26 mtimes), tts_ssml_test2.py (edge-tts break-tag strategy note), PENDING_PODCAST_DELIVERY.md (July 4 episode, proxy SSL failure, file absent), PODCAST-VIDEO-V3-WORKFLOW.md (2026-08-17, cites generator at 344 lines) vs podcast_video_generator_v3.py (wc -l = 445), cron jobs.json (3 podcast jobs, all enabled=false) · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep13-audio-workshop.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep13-audio-workshop.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Text-to-Speech</category><category>Podcasting</category><category>Automation</category></item><item><title>The Model Stable: Two Days in June, One Evening in August</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep14-model-stable/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep14-model-stable/</guid><description>E14: the engine room&apos;s model stable — three local models totaling 70 GB, a two-day deployment battle, and the fallback horse that isn&apos;t saddled today.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 14 of One Man One Legion&lt;/a&gt; — second stop in the engine-room arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep16-mini-agent/&quot;&gt;the mini-agent experiment&lt;/a&gt;. Upstairs, the workshops print, render, and speak. Down here, the fleet keeps its own brains.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;On one workstation disk sit three directories: a 23 GB coder, an 18 GB reasoner, and a 29 GB workhorse — 70 GB of locally owned model weights, measured today. One GPU is awake and nearly full; the card beside it is essentially idle. The question this room answers: what does a one-man fleet actually do with its own models, when a better one is a single API call away?&lt;/p&gt;
&lt;h2&gt;The first horse was a commentator&lt;/h2&gt;
&lt;p&gt;June 19: the coder arrives, 23 GB. Its directory timestamp reads 11:27; sixteen minutes later a first serve attempt tries to span both GPUs and dies on a 600-second timeout waiting for engine cores. Single-GPU it is. The same day, a handbook is written with tested boundaries — not guesses: a real project of 11,648 lines is roughly 100K tokens against a 32K window, three to four times too big to load; hand the model a 190-line module to debug and it misses all 3 real bugs while flagging 1 false one. Verdict: a code commentator for the editor, not a project detective. The stable&amp;#39;s first lesson was a boundary.&lt;/p&gt;
&lt;h2&gt;Two days in June&lt;/h2&gt;
&lt;p&gt;June 25, 07:47 — the decision to stand up a local fallback. The 18 GB download crawls at ~300 KB/s through the proxy, with a 17-hour estimate and frequent disconnects; switching to a domestic mirror at 13:10 lifts it to 2.3 MB/s, seven times faster, and the weights land by 15:12. Then comes the serving battle: seven pitfalls across two days, first light at 19:08, first chat room verified at 21:26, and a final conclusions note at 07:30 the next morning.&lt;/p&gt;
&lt;p&gt;Three pits deserve engraving. The tool-call parser took a fifth attempt to pair — the reasoning parser and the tool parser must come from the same qwen3 family, or they talk over each other. The 32K context window overflowed on the agent framework&amp;#39;s own 28K system prompt (the error cites at least 28,673 input tokens), so the window went to 64K. And a tool named &lt;code&gt;web_search&lt;/code&gt; went uncalled — the model confused it with a built-in from its training; renamed &lt;code&gt;internet_search&lt;/code&gt;, it worked.&lt;/p&gt;
&lt;p&gt;Then the triage, honest as a chart: across ten chat rooms, the 9B is fully capable in 4 (single-tool jobs), degraded-acceptable in 3, and forbidden in 3 — deep trend analysis, stock data, long-form creation, where the article workflow alone needs 170K tokens of context, far past the local ceiling. The standing rule: publish nothing rather than let a 9B fake its way through analysis. And the reason the fallback exists at all: same-provider backups die together — the cloud 5.2 and 4.7 share one API — so real failover crosses providers, cloud first, local 9B behind it. Read fresh this morning, that chain is exactly two rungs.&lt;/p&gt;
&lt;h2&gt;One evening in August&lt;/h2&gt;
&lt;p&gt;August 18: the workhorse. First download byte logged 14:41 — domestic mirror, resume-safe script, which is June&amp;#39;s download lesson promoted to infrastructure. The serving banner comes up at 19:58; by 20:02 a script mounts Claude Code on the local weights. 5 hours 21 minutes, first byte to coding agent.&lt;/p&gt;
&lt;p&gt;The August start script is June&amp;#39;s scar tissue, worn as defaults: the same-family parser pair, the FlashInfer kill switch, a 38,912-token window, fp8 KV cache — under the same vLLM 0.23.0 banner as June&amp;#39;s logs. The new pits were small and solved in comments: Claude Code sends an effort level this vLLM rejects with a 500, so the mount script now forces one explicitly — low, for a 27B. The 4090D&amp;#39;s CUDA index is probed at runtime, because enumeration shifts between boots. The newest horse even runs inside a conda environment still named after the first.&lt;/p&gt;
&lt;h2&gt;The honesty file&lt;/h2&gt;
&lt;p&gt;Today is day 9 of the 27B&amp;#39;s continuous run — 95% of the 4090D&amp;#39;s 24 GB, the A4000 idle beside it. The failover chain names the local 9B as the fallback, but the endpoint it points at has no listener; that log ends with a clean shutdown. And the 27B holding the GPU is not in the chain at all — it is a work mount, not insurance. The stable is not a hot standby. It is a promise kept by scripts: one command re-lights the fallback, about the 3-4 minutes the 27B itself needs to come ready.&lt;/p&gt;
&lt;p&gt;What 70 GB of owned weights buys is not speed and not quality. It is a chain of custody — one link of the fleet&amp;#39;s brain that no outage, quota, or pricing change can take away — plus a compounding record: June&amp;#39;s two days bought August&amp;#39;s one evening.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — next in the engine-room arc: the stress test that graded the newest horse.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation model directories (du -sh, counted 2026-08-27: gemma-4-12b-coder 23 GB, Qwen3.8-27B-AWQ 29 GB, Qwythos-9B-Claude-Mythos-5-1M 18 GB; dir mtimes 06-19 / 08-18 / 06-25), ~/本地大模型部署全记录.md (v1.0, 2026-06-26: two-day timeline 07:47 decision → 07:30 conclusions, seven pitfalls, seven iron rules, 4/3/3 room triage, 28673-token overflow error), ~/vLLM部署对比-Gemma vs Qwythos.md (2026-06-25), ~/gemma-coder-handbook.md (2026-06-19: 11648-line project vs 32K window, 190-line module 3 bugs missed + 1 false positive), ~/vllm-gemma4-tp2.log (2026-06-19 11:43 first serve attempt, 600s engine-core timeout), ~/models/download_qwen38.log (first line 2026-08-18 14:41:08, ModelScope), ~/models/download_qwen38_awq.sh (resume-safe mirror script), ~/models/start_qwen38.sh + vllm_qwen38.log (banner 08-18 19:58:41, vLLM 0.23.0, 38912-token window), ~/.openclaw/workspace/qwen38-eval/cc-q38.sh (mtime 08-18 20:02, effort-500 quirk comment, 3-4 min ready), live state 2026-08-27: ps (vLLM up since Aug 18), nvidia-smi (23314/24564 MiB on the 4090D, 267 MiB on the A4000), ss (no listener on the fallback endpoint), openclaw.json via jq (primary zai/glm-5.3, fallbacks [local-vllm/qwythos-9b], 27B under a separate provider) · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep14-model-stable.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep14-model-stable.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>vLLM</category><category>Local LLM</category><category>Open Models</category></item><item><title>Model or Harness? Two Controlled Experiments on a 363-Line Hand-Rolled Coding Agent</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep16-mini-agent/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep16-mini-agent/</guid><description>E16: swap the brain — 15 rounds stuck vs 4 to pass; swap the prompt — weak models stay unsaveable. Three code supervisors make agents honest.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 16 of One Man One Legion&lt;/a&gt;, from the engine-room arc. The map lives there; &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep1-image-workshop/&quot;&gt;episode 1&lt;/a&gt; covered the image workshop.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Before commanding a legion, it helps to have built one soldier by hand. This episode is about a 363-line coding agent — dumb loop, four tools, no framework — and the two controlled experiments run on it. Swap the brain, hold the harness: a local 12B coder stalls 15 rounds on a one-line fix; a cloud model passes in 4. Hold the brain, swap the prompt: the strong model finishes either way but only reports responsibly under the structured prompt; the weak one fails under all three variants. And three code-level supervisors fixed the one thing prompts never could — the agent stopped lying about success. The thesis the whole series leans on, measured in one repo: intelligence is the ceiling, reliability is the floor, and the floor is where the work is.&lt;/p&gt;
&lt;h2&gt;The origin was a dead process, not a research question&lt;/h2&gt;
&lt;p&gt;The story opens with &amp;quot;the bot stopped responding.&amp;quot; The instinct to read code was wrong: the process had exited 14 hours earlier — crash, not bug. The fix was operational (a daemon that restarts it within 5 seconds), and somewhere between the &lt;code&gt;ps&lt;/code&gt; and the restart script came the question: if supervision is this cheap, what else can a supervised loop do? The 363-line agent is the answer, and its own docstring still carries the first scar: vLLM&amp;#39;s gemma tool parser intermittently returns multi-turn tool calls as plain text with an empty &lt;code&gt;tool_calls&lt;/code&gt; array — the model generates the right format, the parsing layer drops it. The harness healed around it with its own fallback parser, and made completion explicit via a &lt;code&gt;finish(summary)&lt;/code&gt; tool so nothing depended on unreliable prose.&lt;/p&gt;
&lt;h2&gt;Experiment 1: hold the harness, swap the brain&lt;/h2&gt;
&lt;p&gt;Same agent, same task (fix a &lt;code&gt;clamp&lt;/code&gt; function missing its return, pytest must go fail → pass), only the model changes:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;gemma-4-12b-coder (local)&lt;/th&gt;
&lt;th&gt;GLM-4.6 (cloud)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Outcome&lt;/td&gt;
&lt;td&gt;Stuck 15 rounds&lt;/td&gt;
&lt;td&gt;Passed in 4 rounds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edit behavior&lt;/td&gt;
&lt;td&gt;Edited a hallucinated &lt;code&gt;is_even&lt;/code&gt; — never the real &lt;code&gt;clamp&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Exact multi-line match, adds the return&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pytest after&lt;/td&gt;
&lt;td&gt;1 failed&lt;/td&gt;
&lt;td&gt;3 passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Temperature note&lt;/td&gt;
&lt;td&gt;temp=0 made it deterministically loop the same mistake&lt;/td&gt;
&lt;td&gt;Stable at temp=0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The trace is brutal to read: the 12B &lt;em&gt;read&lt;/em&gt; &lt;code&gt;clamp&lt;/code&gt;, then spent its edits on a function that did not exist in the file. The harness absorbed everything absorbable — parser flakiness, PATH errors, retry discipline — and the semantic drift still sank it. That is the ceiling/floor line, drawn by experiment: &lt;strong&gt;infrastructure can catch a lying parser, not a model aiming at the wrong function.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Experiment 2: hold the brain, swap the prompt&lt;/h2&gt;
&lt;p&gt;Same model, same task family (add &lt;code&gt;clamp_list&lt;/code&gt; plus tests), only the system prompt changes:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Slogan prompt&lt;/th&gt;
&lt;th&gt;Structured prompt&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;GLM rounds&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM result&lt;/td&gt;
&lt;td&gt;4 passed&lt;/td&gt;
&lt;td&gt;4 passed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM report&lt;/td&gt;
&lt;td&gt;Flat list&lt;/td&gt;
&lt;td&gt;What changed / risks / how verified, scan-first opening&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The strong model converges either way; the structured prompt changes the &lt;em&gt;quality of accountability&lt;/em&gt; (risk analysis, verification story), not the ability to finish. The 12B failed under all three variants — structured truncated into zero tool calls, the others hallucinated edit targets, ignored tool errors, and declared fake success. Same conclusion twice: &lt;strong&gt;prompt engineering improves how a capable model reports; it does not add capability that is not there.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The three supervisors that made it honest&lt;/h2&gt;
&lt;p&gt;What did work for the weak model was code that refuses, not code that asks:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Rule&lt;/th&gt;
&lt;th&gt;Metaphor from the docs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;A1 — edit validation&lt;/td&gt;
&lt;td&gt;&lt;code&gt;old_string&lt;/code&gt; must appear in the last-read file content&lt;/td&gt;
&lt;td&gt;A foreman who checks your quote against what you just read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A2 — error forcing&lt;/td&gt;
&lt;td&gt;A tool &lt;code&gt;❌&lt;/code&gt; injects a mandatory re-read before continuing&lt;/td&gt;
&lt;td&gt;The foreman walks you back to the step you skipped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A3 — finish validation&lt;/td&gt;
&lt;td&gt;No successful edit, no finish — including zero-tool runs&lt;/td&gt;
&lt;td&gt;Passing the exam requires having answered the question&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Measured effect: fake success disappeared. Honest incapacity remained — the 12B now fails &lt;em&gt;visibly&lt;/em&gt;, which is the correct behavior for a fallback tier and the whole point: honesty is a floor feature, capability is not.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Each cell above is n=1 run-trace evidence from a two-experiment log, on a one-task family (a clamp fix and a clamp-list addition) — a controlled comparison, not a benchmark. The 12B is one coder-tuned model; nothing here indicts local models generally, and the fleet&amp;#39;s own 27B was not in this test (that comparison belongs to a later episode). Traces are as logged by the harness itself.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;mini_agent.py&lt;/code&gt; (363 lines incl. the fallback parser and A1/A2/A3), the two-experiment log with run traces, and the four-step build story as written in the repo&amp;#39;s docs, 2026-08.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from the 2026-08 experiment logs. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next in the engine room: the model stable that keeps three LLMs runnable and the fallback chain that degrades instead of dying. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Local gemma-4-12b-coder on vLLM (workstation GPU) vs cloud GLM-4.6 · same 363-line Python harness for both · run traces as logged by the harness. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep16-mini-agent.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep16-mini-agent.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>LLM</category><category>vLLM</category><category>Agent Harness</category><category>Controlled Experiments</category></item><item><title>The Model Exam: Seven Passes, Two Zeroes, One Home</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep17-model-exam/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep17-model-exam/</guid><description>E17: the fleet grades its local 27B — a 7/7 baseline, an A/B with two zero-output deaths, a four-habitat race, six rules for 241 tokens.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 17 of One Man One Legion&lt;/a&gt; — third stop in the engine-room arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep16-mini-agent/&quot;&gt;the mini-agent&lt;/a&gt; and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep14-model-stable/&quot;&gt;the model stable&lt;/a&gt;. The stable ended on a promise: the stress test that would grade the newest horse. This is it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Hours after the 27B came up, the fleet held its exam. The goal was not to make the model stronger but to test whether task-splitting, hard constraints, and a division of labor could match usage to the limits of a 4-bit thinker. Design and grading were the cloud model&amp;#39;s job; the subject was local.&lt;/p&gt;
&lt;h2&gt;The baseline: seven tasks, real asserts&lt;/h2&gt;
&lt;p&gt;Capability first: seven coding tasks of rising difficulty, from a discount parameter and an empty-list crash to a log filter, a two-file change, and a messy-style-preservation test. Every answer was executed against assertion scripts, not eyeballed. Result: 7/7 at 13-142 seconds per task and 194-1,834 output tokens, at temperature 0.2. The caveats sit next to the passes: one answer was lifted out by hand (no code block), one key-name mismatch scored as verifier strictness, one put an example block before the real code. Known flaw: on this stack, thinking leaks into the body. Verdict: qualified for daily code duty.&lt;/p&gt;
&lt;h2&gt;The A/B: what splitting actually buys&lt;/h2&gt;
&lt;p&gt;Then the experiment: one task — a column-dedup command for a small CSV tool, assembled versions 53 lines — two usage patterns, six logged requests. Group A got the whole job in one prompt; Group B got three constrained steps (touch only what&amp;#39;s necessary, keep style, no drive-by refactors, say when unsure): the function, the main() branch, a self-check.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;max_tokens&lt;/th&gt;
&lt;th&gt;Time (s)&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;A one-shot&lt;/td&gt;
&lt;td&gt;3,000&lt;/td&gt;
&lt;td&gt;223&lt;/td&gt;
&lt;td&gt;0 chars — dead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A retry&lt;/td&gt;
&lt;td&gt;8,000&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;1,465 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B1 function&lt;/td&gt;
&lt;td&gt;3,000&lt;/td&gt;
&lt;td&gt;125&lt;/td&gt;
&lt;td&gt;298 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B2 main()&lt;/td&gt;
&lt;td&gt;3,000&lt;/td&gt;
&lt;td&gt;58&lt;/td&gt;
&lt;td&gt;604 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B3 self-check&lt;/td&gt;
&lt;td&gt;1,500&lt;/td&gt;
&lt;td&gt;123&lt;/td&gt;
&lt;td&gt;0 chars — dead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B3 retry&lt;/td&gt;
&lt;td&gt;6,000&lt;/td&gt;
&lt;td&gt;260&lt;/td&gt;
&lt;td&gt;468 chars&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The first killer of a 4-bit thinking model is its own budget: thinking burns 500-3,000+ tokens before the body starts, and two of six requests starved to zero characters — the 0-byte corpse file is still archived. Iron rule: never call it below 6,000 max_tokens.&lt;/p&gt;
&lt;p&gt;Splitting bought containment, not quality — the successful one-shot passed all five verifier cases cleanly; the split path passed four, leaking raw rows on the missing-column fifth. The one-shot rewrote the entire file, so review meant reading everything; each split step returned one function, a diff you can hold in your head. The three good split steps took 443 seconds against 479 for the one-shot&amp;#39;s two attempts — no slower, both dead runs visible above.&lt;/p&gt;
&lt;p&gt;The most valuable find was self-inflicted. That missing-column leak? The split prompt itself said to return the original list unchanged — the model obeyed, literally. A small model in split mode obeys every design decision in a template to the letter; the template must be reviewed before it becomes law. Given a real budget, the B3 self-check caught exactly that bug, plus two more.&lt;/p&gt;
&lt;h2&gt;The habitat race: same weights, four doors&lt;/h2&gt;
&lt;p&gt;The same weights answer to four doors: the agent framework&amp;#39;s alias, a wrapper that mounts Claude Code, a lightweight terminal harness on a handwritten route, and the bare localhost endpoint. The exam measured each door&amp;#39;s ticket out of the 38,912-token window. Claude Code costs 24,769 tokens of system prompt — 64% of the window, a ticket designed for a 200K-token model worn by a 38.9K one, leaving 6.1K of working space. The harness charges 7,710 — 31% of that — for some 23K of dialogue space, 3.8× the working room. The bare endpoint is free.&lt;/p&gt;
&lt;p&gt;Two upgrades were refused with numbers: 48K with CPU offload collapsed throughput to 3.7 tokens/s against the ~14 baseline — an 8,000-token answer would take 36 minutes; and prefix caching wouldn&amp;#39;t start at that window, its ceiling of about 12 seconds a round not worth 5,600 tokens of context.&lt;/p&gt;
&lt;h2&gt;Six rules, 241 tokens&lt;/h2&gt;
&lt;p&gt;The winning harness can&amp;#39;t vary instructions per model, so six global rules were written to be harmless to the cloud model too: read before editing, minimal changes, stop after two failures, ask when unsure, verify after the change, split big jobs. Measured cost: the ticket rose from 7,710 to 7,951 — 241 tokens, 3.1%. With the rules as the only variable, correctness stayed 5/5, rules-on consumed 49,394 prompt tokens against 64,907 rules-off, and only rules-on asked about an ambiguity. Told to refactor boldly, the cloud model correctly overrode the minimal-change rule — the injection is protocol-level, never outranking a direct instruction.&lt;/p&gt;
&lt;p&gt;Final proof: the log filter, hardest of the seven. Re-run that evening on the bare endpoint, it died the zero-character death; inside the harness it finished end-to-end, every requirement met, for 68,984 tokens — the harness absorbs output across rounds, dissolving the budget trap. Along the way the model caught a bug in the verifier itself — an expected count of 3 where the fixture holds 2 — and refused to fix either side, stopping to ask which was right. Rule four, executed as designed.&lt;/p&gt;
&lt;h2&gt;The verdict, still standing&lt;/h2&gt;
&lt;p&gt;At 21:05 the verdict landed: the harness is this model&amp;#39;s home; the bare endpoint serves pipelines, where agents enforce split discipline; Claude Code keeps the cloud model, the wrapper stays a spare; hard jobs get cloud review. The logic compresses to a line: big models are judged by intelligence, small models by how frugally you run them — one door saves space, one saves the ticket, the third saves neither. Nine days later, the serving process born that evening still holds the GPU.&lt;/p&gt;
&lt;p&gt;Scope, as logged: one task family, one round per cell, 53-line-scale single-file edits; larger repos untested. An exam, not a benchmark.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — next in the engine room: swapping the cloud brain without stopping the factories.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation ~/.openclaw/workspace/qwen38-eval/ (dir mtime 2026-08-18): Q38压测报告-20260818.md (v2.0 — A/B raw table, T1-T7 appendix, tuning log) + WORKLOG-20260818.md (four-habitat section, rules ON/OFF, verdict), run.py (A/B runner — prompts verbatim incl. the return-original-list wording; temperature 0.2) + run_tests.py (baseline runner), t1-t7 .md/.py pairs (per-task time and token counts in file headers), A_full.md (0 bytes) / A_full8k.md / B1_dedup.md / B2_main.md / B3_check.md (split-run outputs), toolA.py + toolB.py (53 lines each; toolB carries the missing-column print bug), verify.py (T1-T5 asserts + five dedup cases) + data.csv / empty.csv, cc-q38.sh + start_qwen38-final.sh; ~/.openclaw/workspace/memory/2026-08-18.md (ticket 24769 vs 7710 = 31%, rules 7710→7951, ON/OFF 49394/64907, log-filter revival 68984, verifier-bug stop-and-ask); ~/.dsh/AGENTS.md (six rules, on disk 2026-08-27); live state 2026-08-27: ps / ss / nvidia-smi — the 27B serving process up since Aug 18 · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep17-model-exam.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep17-model-exam.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Local LLM</category><category>vLLM</category><category>Model Evaluation</category></item><item><title>The Brain Swap: Ten Backups, One Insurance Card, No Standby Factory</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep18-brain-swap/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep18-brain-swap/</guid><description>E18: swapping the factory&apos;s primary model while it runs — a transplant at age 2 hours, a 139-day gap, and an evening double-swap with a rollback card.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 18 of One Man One Legion&lt;/a&gt; — fourth stop in the engine-room arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep16-mini-agent/&quot;&gt;the mini-agent&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep14-model-stable/&quot;&gt;the model stable&lt;/a&gt;, and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep17-model-exam/&quot;&gt;the model exam&lt;/a&gt;. The exam graded the newest local horse. This one is about the cloud brain itself — and what it takes to replace it while the factory keeps running.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every factory fears the day its brain must come out. This fleet has swapped its primary model at least three times — glm-5 to 5.1 to 5.2 to 5.3 — with no twin factory to fail over to, no maintenance window, and no operations team. The whole history sits in one directory: ten &lt;code&gt;openclaw.json&lt;/code&gt; files, stacked like strata.&lt;/p&gt;
&lt;h2&gt;The strata&lt;/h2&gt;
&lt;p&gt;The layer holds 10 files: the live config, 5 rolling backups (&lt;code&gt;.bak&lt;/code&gt; through &lt;code&gt;.bak.4&lt;/code&gt;, the platform breathing on its own), and 4 with names — &lt;code&gt;backup-glm5&lt;/code&gt;, &lt;code&gt;backup-before-glm51&lt;/code&gt;, &lt;code&gt;backup-before-glm53&lt;/code&gt;, &lt;code&gt;backup-0818-notify&lt;/code&gt;. Rolling backups are routine; named backups are decisions. The oldest file anywhere in the directory is &lt;code&gt;backup-glm5&lt;/code&gt;, timestamped 2026-03-29 10:53 — the platform&amp;#39;s birth certificate, 2,459 bytes. The live config, last touched 2026-08-18 at 18:39, is 7,540. The body tripled; the line that names the brain stayed one line.&lt;/p&gt;
&lt;h2&gt;Transplant at age two hours&lt;/h2&gt;
&lt;p&gt;The birth config is a newborn&amp;#39;s: primary &lt;code&gt;zai/glm-5&lt;/code&gt;, one model, one alias, no fallback. Then, 118 minutes later, the second-oldest artifact: &lt;code&gt;backup-before-glm51&lt;/code&gt;, 12:51 the same day. Inside it, glm-5.1 is already staged in the catalog while the primary is still glm-5 — snapshot, then flip, same day. The migration method is visible on day one: stage the new brain in the catalog, back up under a name that says why, then turn the key.&lt;/p&gt;
&lt;h2&gt;The unmarked grave&lt;/h2&gt;
&lt;p&gt;Between 12:51 on March 29 and 21:05 on August 15 — 139 days — not one named backup. Somewhere in that silence glm-5.2 became primary, and it left no named grave. The evidence is sideways: a July 3 diary entry already tests 5.2 as the default; the June deployment record argues about 5.2&amp;#39;s fallback. So the true chain is glm-5 → 5.1 → &lt;em&gt;[gap]&lt;/em&gt; → 5.2 → 5.3, and one link was laid without ceremony. Honest archaeology reports the gap. The gap is also the argument for what came next.&lt;/p&gt;
&lt;h2&gt;August 15: two brains, one evening&lt;/h2&gt;
&lt;p&gt;Day shift first: 20 articles between 14:38 and 19:21 — 17 paper digests inside 13 minutes, 3 evening pieces after. Then, with the day&amp;#39;s work done, surgery.&lt;/p&gt;
&lt;p&gt;From 20:52 to 21:00 the agent upgraded Claude Code&amp;#39;s own config: 3 edits — the model, the opus mapping, the availability list — with the Sonnet and Haiku mappings deliberately untouched. Backup at 20:54. A side discovery on the way: requests for 5.2 were coming back labeled 5.3. The provider may have rerouted server-side before anyone local touched anything.&lt;/p&gt;
&lt;p&gt;At 21:05 the named backup (&lt;code&gt;backup-before-glm53&lt;/code&gt;, primary still 5.2 inside). At 21:08 OpenClaw&amp;#39;s own brain — through the gateway&amp;#39;s validated config patch, not a hand edit, 3 edits again. The diary logs a 4-item protocol for upgrading oneself: verify both routes separately, back up under a name, &lt;strong&gt;send the insurance card before restarting&lt;/strong&gt; — the rollback command handed to the human over Telegram before the agent reboots itself — and leave the fallback chain frozen, the local 9B catching any fall. The surgeon hands the undo switch to the owner mid-operation.&lt;/p&gt;
&lt;p&gt;Then the evening&amp;#39;s correction, kept honest: a 21:17 test of the Anthropic-protocol route added a temporary provider, and the agent persisted it. The user&amp;#39;s verdict, recorded in the diary, was that this overstepped — the test was for connectivity only. Test entries removed, config restored; the two rolling backups at 21:18 and 21:27 fossilized the whole affair, and the final state is clean.&lt;/p&gt;
&lt;h2&gt;Did the factory crash?&lt;/h2&gt;
&lt;p&gt;The next day, August 16, shipped 1 article. August 17 brought 6. On August 18 the catalog grew to 7 seats — the 27B joined under its own provider, with the tool-call flag the diary warns must be on or OpenClaw never passes tools — but as a catalog seat only: not primary, not in the 2-rung chain where 1 local fallback waits beneath the cloud. The brain was swapped; the org chart of brains simply grew.&lt;/p&gt;
&lt;p&gt;The craft, read from the strata: name your backups with reasons. Stage before flipping. Check both routes. Hand the rollback to the human before you reboot yourself. Touch nothing else — mappings, fallbacks, endpoints stay frozen. And when the test is over, erase the test. None of it needs a standby factory. It needs a morgue and a protocol.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — the full map of one man commanding a legion, and the other episodes of the engine-room arc.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation ~/.openclaw/ (read 2026-08-27): ten-file openclaw.json layer (live 7,540 B + rolling .bak/.bak.1-.4 + named backup-glm5 03-29 10:53 2,459 B / backup-before-glm51 03-29 12:51 2,834 B / backup-before-glm53 08-15 21:05 6,109 B / backup-0818-notify 08-18 12:32), each parsed for primary/fallback/catalog fields; backup-glm5 verified oldest file in ~/.openclaw; json-diff of .bak.4 (21:18) vs .bak.3 (21:27) isolates the zai-anthropic test entry; ~/.openclaw/workspace/memory/2026-08-15.md (CC upgrade 20:52-21:00, OC upgrade 21:08, four-item self-upgrade protocol, insurance card, test≠deployment correction, provider-side 5.2→5.3 reroute note) and 2026-08-18.md (Q38 registration 18:40, supportsTools flag, alias rule); ~/.claude/ settings strata (backup-20260329-115755, backup-before-glm53 08-15 20:54); article counts via ls|wc -l in wechat-editor-team/articles/ (20 on 08-15, 1 on 08-16, 6 on 08-17); July 3 diary for the 5.2-as-default waypoint · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep18-brain-swap.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep18-brain-swap.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Model Migration</category><category>Config Management</category><category>LLM Ops</category></item><item><title>What Should an AI Legion Do While You Sleep? 22 Cron Entries, Four Still Lit</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep19-cron-contraction/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep19-cron-contraction/</guid><description>E19: the scheduler grew to 16 jobs lit at once — 790 runs, 154 errors — then on Jul 23 the auto-publishers went dark. Four relit: feed or ask, never publish.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 19 of One Man One Legion&lt;/a&gt; — the first of the war stories. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep2-tg-gateway/&quot;&gt;Episode 2&lt;/a&gt; ended on a promise: &amp;quot;a scheduler that once ran 22 jobs.&amp;quot; This is that story, and the number needs correcting along the way.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every legion eventually faces the seduction of the clock: if a job can run at 6 a.m. unattended, why shouldn&amp;#39;t everything? This is the story of a one-person AI legion that answered that question by growing a cron schedule into a small publishing empire — and then, in one day in July, switching most of it off on purpose.&lt;/p&gt;
&lt;h2&gt;An empire measured in JSONL&lt;/h2&gt;
&lt;p&gt;The run logs begin at 21:27 on March 29, 2026 — an odd hour for a job scheduled at 03:35, which is to say the first fire was a human pressing run-now to watch it work. In the 116 days from that click to late July, the schedule grew: hot-topic monitoring morning and evening, a pure-news push at both ends of the day, six WeChat article entries, three podcast productions (one reading the news aloud in a cloned voice), two Feishu pushes, a US-stock column, plus a diagnostic job and a headless video-factory trial that never fired once. At the observed peak, &lt;strong&gt;16 jobs were lit at once&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;What the empire actually did, in aggregate: &lt;strong&gt;790 finished runs, 154 ending in status &amp;quot;error&amp;quot; — about one in five&lt;/strong&gt;. The workhorse, the morning hot-topic monitor, fired &lt;strong&gt;120 times with 25 errors&lt;/strong&gt;; the voice podcast managed &lt;strong&gt;11 errors in 21 runs&lt;/strong&gt;. The seven backup snapshots in the cron directory read like growth rings: 12 entries on July 4, 17 by July 8 — then a snapshot from Aug 18 showing 20 entries with exactly &lt;strong&gt;3 still enabled&lt;/strong&gt;. The file kept accepting tombstones even as the lights went out; it holds 22 today.&lt;/p&gt;
&lt;h2&gt;July 23, 08:00&lt;/h2&gt;
&lt;p&gt;The last of the old guard ran at 08:00 on July 23 — a &amp;quot;deep story&amp;quot; slot whose label still said 17:20, the name and the cron expression having drifted apart sometime in the spring. Then: nothing. No backup marks the switch-off; the evidence is the silence. The old guard&amp;#39;s final runs cluster on July 22–23 (a one-shot diagnostic had fired its only run on July 5; one duplicate noon copy had gone quiet on July 11), and for &lt;strong&gt;25 days&lt;/strong&gt; the logs show no cron fire at all.&lt;/p&gt;
&lt;p&gt;The file itself shows what kind of operation this had become. Two entries share the identical name — the noon topic slot, filed twice, each copy running for weeks in parallel. Six orphaned run files belong to jobs deleted so thoroughly they no longer have names, just 13 log lines of June experiments; one orphan is nothing but an apology, a note that a podcast episode couldn&amp;#39;t be delivered that morning. Automation&amp;#39;s last words were often about automation failing.&lt;/p&gt;
&lt;h2&gt;The relight, and the pattern in it&lt;/h2&gt;
&lt;p&gt;On Aug 17 at 16:31 — mid-afternoon, off-schedule, a hand on the button again — the first job came back: a YouTube-subscription material drop that feeds raw leads into a group at 03:30 each morning. The next day two &amp;quot;video topic&amp;quot; jobs relit on their own clocks, at 03:20 and 17:40; their entire job is to ask the operator what today&amp;#39;s video should be. Two days after that a fourth lit: an arXiv paper digest at 00:55, 08:55 and 16:55, feeding picks into the article pipeline. Its first delivery was &lt;strong&gt;86 papers, on Aug 20&lt;/strong&gt; — itself an off-schedule manual fire, like the YouTube relight.&lt;/p&gt;
&lt;p&gt;Look at what survived and the criterion is embarrassingly plain. &lt;strong&gt;The four live jobs either feed material to the human or ask the human a question. Everything that decided, finished, and published on its own schedule is dark.&lt;/strong&gt; The WeChat articles didn&amp;#39;t stop being made — they moved to a human-gated pipeline with machine QC gates, where nothing ships without a nod. What remains on the clock is precisely the work that makes sense unattended: gathering and asking.&lt;/p&gt;
&lt;p&gt;Nor was this an isolated instinct. Five days before the lights went out, the same operator cut a 2,182-line workflow SOP down to about 180 lines, and 161 mandatory rules down to 15 (operator&amp;#39;s memory diary, 2026-07-18). Over-automation was being pruned everywhere the hand could reach.&lt;/p&gt;
&lt;h2&gt;The survivors still bleed&lt;/h2&gt;
&lt;p&gt;Do not mistake the cut for a reliability upgrade. During the Aug 21–27 model storm — the cloud GLM-5.3 timing out and the local fallback dying with it, in the log&amp;#39;s own words &amp;quot;All models failed (2)&amp;quot; — the arXiv job logged &lt;strong&gt;18 consecutive errors&lt;/strong&gt; across its 24 runs to date. Its 16:55 slot on Aug 27 came back green and delivered &lt;strong&gt;82 papers&lt;/strong&gt;. The lesson is not that contraction fixed anything; it is that once only feed-and-ask jobs remain, an unreliable day costs you a missed digest instead of a publication you didn&amp;#39;t want.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Counts are from &lt;code&gt;jobs.json&lt;/code&gt; (22 entries, 4 enabled), seven backup snapshots (peak observed enablement 16), and 27 per-job run logs (790 finished runs, 154 errors), read Aug 27, 2026, timestamps Asia/Shanghai. The switch-off itself left no log line — &amp;quot;the operator cut it&amp;quot; is inferred from what was relit and when, not from a recorded decision. First-run manual fires are inferred from off-schedule timestamps. Error status includes timeouts whose work may have partially completed; no failure taxonomy is claimed. The 22 entries include both noon duplicates; 16 is the peak observed in backups, not a proven maximum.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;jobs.json&lt;/code&gt; and its backups (entry counts, enable states, schedules, name drift), &lt;code&gt;runs/*.jsonl&lt;/code&gt; (790 run records with statuses and summaries), and the workstation operator&amp;#39;s memory diary (SOP cut, 2026-07-18), 2026-03–08.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from live files. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next in the war stories arc: the legion&amp;#39;s bug diary — the failures it kept on itself. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw cron daemon · evidence = jobs.json, 7 backup snapshots, 27 per-job JSONL run logs (790 finished runs), timestamps Asia/Shanghai, read 2026-08-27. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep19-cron-contraction.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep19-cron-contraction.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Cron</category><category>Automation</category><category>Scheduling</category><category>Personal Infrastructure</category></item><item><title>What Does It Take to Give a Chat AI Nine Tools and Your Home Directory? A 50-Line Bot That Grew to 850</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep2-tg-gateway/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep2-tg-gateway/</guid><description>E2: a Telegram gateway to GLM-5.2 — nine tools, a jailed workdir, a dangerous-command blocklist; 50 lines grew to 850, logged in 168,996 lines.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 2 of One Man One Legion&lt;/a&gt;, opening the cockpit arc. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep16-mini-agent/&quot;&gt;Episode 16&lt;/a&gt; tells where this bot&amp;#39;s supervision lessons migrated next; the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;map&lt;/a&gt; holds the series together.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every legion needs a doorway. This one is a Telegram bot: type a message, GLM-5.2 thinks, and nine tools execute on a real Linux home directory — read, grep, edit, write, run commands, search and fetch the web, pull RSS. The README still calls it &amp;quot;~50 lines of code,&amp;quot; and that was true once. The file on disk today is &lt;strong&gt;850 lines&lt;/strong&gt;, and the difference between those two numbers is the actual story of this episode: what it costs, in code, to let a chat model touch your filesystem without regretting it.&lt;/p&gt;
&lt;h2&gt;The premise: one protocol, one credential, one chat window&lt;/h2&gt;
&lt;p&gt;The bot speaks the Anthropic protocol to a GLM-5.2 endpoint — the same coding-plan credential a Claude Code CLI uses, so the &amp;quot;brain&amp;quot; is a cloud model billed like a coding assistant, not a bespoke integration. A chat message becomes a tool loop: the model may chain up to &lt;strong&gt;15 tool iterations&lt;/strong&gt; per turn, which is what lets &amp;quot;find the failing test and fix it&amp;quot; run as read → grep → edit → exec → read, unattended. The nine tools:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Guard behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;&lt;code&gt;read_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Wrong path falls back to a filename search instead of erroring out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;glob&lt;/code&gt; / &lt;code&gt;grep&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Bounded result sizes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;edit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Exact-match &lt;code&gt;old_string&lt;/code&gt; required — no whole-file rewrites&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;write_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Jailed to the work directory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;exec_command&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Blocklist: &lt;code&gt;rm -rf&lt;/code&gt;, &lt;code&gt;sudo&lt;/code&gt;, &lt;code&gt;mkfs&lt;/code&gt;, &lt;code&gt;dd of=&lt;/code&gt; intercepted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;web_search&lt;/code&gt; / &lt;code&gt;web_fetch&lt;/code&gt; / &lt;code&gt;rss_fetch&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Read-only externals, capped output&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;h2&gt;The 800 lines that grew after the failures&lt;/h2&gt;
&lt;p&gt;The README&amp;#39;s own section header for all of this is &amp;quot;safety and robustness — added gradually, after stepping in holes.&amp;quot; What got added, and why each exists:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The work-directory jail, with prefix-spoof protection.&lt;/strong&gt; Every file tool refuses paths outside one directory — including &lt;code&gt;/home/user-evil&lt;/code&gt;-style prefixes that would textually pass a naive &lt;code&gt;startswith&lt;/code&gt; check. The check resolves real boundaries, not string prefixes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The dangerous-command blocklist.&lt;/strong&gt; The model mostly proposes reasonable commands; the one time it doesn&amp;#39;t, interception is cheap and a restore from a bad &lt;code&gt;rm&lt;/code&gt; is not.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exact-match edits.&lt;/strong&gt; Early versions let the model rewrite whole files, which converted small mistakes into large ones. Forcing &amp;quot;point at exactly this string&amp;quot; keeps blast radius proportional to intent.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Path fallback on reads.&lt;/strong&gt; A wrong path used to stall the loop with an error the model would then narrate instead of fixing. Now the tool searches by filename and hands back something usable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The philosophy line in the README could have been lifted from the mini-agent docs of &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep16-mini-agent/&quot;&gt;episode 16&lt;/a&gt;: &lt;em&gt;tolerant tools, bounded environment — the brain is the model&amp;#39;s; how much of it you actually get is the harness.&lt;/em&gt; The two projects traded lessons in both directions — the bot&amp;#39;s dead-process incident (found 14 hours after exit) is what taught the daemon-first diagnostic habit that the mini-agent later inherited as step one.&lt;/p&gt;
&lt;h2&gt;Operations: what &amp;quot;reliable&amp;quot; means at personal scale&lt;/h2&gt;
&lt;p&gt;The bot runs as a systemd user unit: 5-second crash restart, lingering enabled so it survives without a login session. Its log has grown to &lt;strong&gt;168,996 lines&lt;/strong&gt; — and the honest detail worth publishing is that the log&amp;#39;s &lt;em&gt;tail&lt;/em&gt; is a live &lt;code&gt;NetworkError: server disconnected&lt;/code&gt; trace. At personal scale, &amp;quot;reliable&amp;quot; is not the absence of failures; it is crash-restart hygiene plus a log you can actually read. No uptime claims are made here — the evidence would laugh.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Line counts and tool lists are from the source on disk (850 lines, 9 tools, 15-iteration cap, 5-second restart). The &amp;quot;~50-line origin&amp;quot; is the README&amp;#39;s own account of the starting point, not an independently reconstructable measurement. The log-line count measures chatter, not messages or tasks — no throughput is claimed. One whitelisted user, by design: this is a personal gateway, not a service.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;bot.py&lt;/code&gt; (850 lines, tool definitions and guards as cited), the project README (origin narrative, design philosophy), and &lt;code&gt;bot.log&lt;/code&gt; (168,996 lines incl. the terminal NetworkError), 2026-08.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from source and logs. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next in the cockpit arc: the runtime this bot&amp;#39;s philosophy grew into — twelve chat groups, trigger words, and a scheduler that once ran 22 jobs. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation-hosted Python bot under a systemd user unit (5 s crash restart) · GLM-5.2 via the Anthropic-protocol-compatible endpoint · line counts from the source, log count from bot.log. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep2-tg-gateway.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep2-tg-gateway.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Telegram Bot</category><category>GLM</category><category>Agent Tools</category><category>Personal Infrastructure</category></item><item><title>Bug Diaries: 9 Failure Files, 158 Days, and the Root Cause It Got Wrong</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep20-bug-diaries/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep20-bug-diaries/</guid><description>E20: the legion files its own failures — a 901-line error archive, a review rota dead on arrival, and a fabricated match report no rule caught.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 20 of One Man One Legion&lt;/a&gt; — the third war story. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep6-agent-diary/&quot;&gt;Episode 6&lt;/a&gt; opened the diary the legion keeps; this one reads the incident reports — 9 files whose only purpose is to record what went wrong, and the awkward question of whether writing them down changed anything.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;A library born in 33 hours&lt;/h2&gt;
&lt;p&gt;The failure library has a founding weekend. On Mar 14, 2026 at 12:25, the agent saved its first free-standing lesson file into the memory directory: a write-up of an article praising NVIDIA NeMo Retriever accuracy numbers no production system could live with. The human rejected the piece; the file&amp;#39;s fix was a new veto item in the review checklist. Five hours later a second file: the browser-tool detour — 2 hours from 16:15 to 17:05 spent installing Chromium and coaxing a debug port that never answered, to reach sites whose RSS already worked. The haul: 20 headlines, of which 0 covered the target story, because the story was a politically sensitive one state media simply wasn&amp;#39;t covering. The file closes its own ledger flatly: value produced, 0. Lesson value: five stars.&lt;/p&gt;
&lt;p&gt;The next day, two more. A RAG stress test on an NVIDIA 10-K scored 15 percent overall — vector search returned 0 results on all three real queries, only 8 of 401 text chunks held usable revenue data, and a deliberately false query returned hits anyway. And the media-tier collapse: a news digest built on 0 first-tier, 0 second-tier, and 7 third-tier sources, whose root-cause list opens with the most expensive line in the whole library — the standard already lived in MEMORY.md, and the session never looked. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep6-agent-diary/&quot;&gt;Episode 6&lt;/a&gt; said there is no read receipt for memory; here is that failure, filed by the memory system&amp;#39;s own customer. The fix took 29 minutes.&lt;/p&gt;
&lt;p&gt;That is 5 failure documents in under 33 hours (Mar 14 12:25 to Mar 15 21:08). The fifth was the ambitious one: a review rota — weekly Monday 09:00, monthly top-5, an alarm when the same error repeats. Next review: Mar 22. The scheduler did not fire its first job until Mar 29 21:27 (episode 19), and the rota file was never edited again after Mar 15 21:08. The mechanism was born a week before the thing that would have run it existed, and never joined it.&lt;/p&gt;
&lt;h2&gt;The registry that broke its own numbering&lt;/h2&gt;
&lt;p&gt;The largest volume, the WeChat team&amp;#39;s error archive, runs 901 lines with errors numbered 1 through 25 — and the numbering itself is a specimen. No. 11 exists only as a changelog row (a WeChat title-limit error, found at push time). No. 24 is used twice, for two unrelated failures in July and August. The last entry, No. 25, is a same-day retraction: at 11:47 on Aug 19 the human corrected the diagnosis — the search method was wrong, not the image API the entry had blamed — and that reversal, timestamped one minute before the file&amp;#39;s final save, is the archive&amp;#39;s most recent growth. Its header orders that every article be generated only after reading it, step 0, before anything else, under a three-part pledge: never commit the same class of error twice.&lt;/p&gt;
&lt;h2&gt;The lesson that didn&amp;#39;t take&lt;/h2&gt;
&lt;p&gt;Did the writing prevent anything? The record answers honestly: not by itself. Images were the subject of a lesson file on Mar 29 — uploaded to the media library but never inserted into the article HTML. The failure then recurs 4 documented times: Jun 14 (zero images), Jun 23 (external URLs WeChat won&amp;#39;t render), Jul 12 (stale images reused) — 77 days from lesson to first relapse, and the archive&amp;#39;s own top-5 table ranks image failures No. 1 at 4-plus recurrences. Worse, No. 13: a World Cup article that fabricated a score — Korea&amp;#39;s opener, written as a loss to Mexico with six shots, when Korea played and beat Czechia. The file&amp;#39;s diagnosis: fragments of real data plus plausible inference, assembled into specifics no source contained. Three separate rules forbidding fabrication already existed in the workflow documents. The Jun 17 lesson named the disease for the whole fleet: writing it down had been mistaken for doing it — one pipeline&amp;#39;s docs specified a search engine its script never called, zero calls in the code, and the fix became 5 iron laws, the first being that a document changes nothing until the script, the prompt, or the gate changes.&lt;/p&gt;
&lt;h2&gt;Where enforcement actually went&lt;/h2&gt;
&lt;p&gt;Prevention left the prose. The image rule finally bit as code — the &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep11-quality-gate/&quot;&gt;quality gate&lt;/a&gt; that checks every push — and as wiring: the step-0 read order is stamped into 8 job payloads in the scheduler registry, so the archive is mandatory pre-reading, not optional memory. The contraction episode 19 described left all 8 of those jobs dark, which is its own lesson: even enforcement rusts. The daily loop &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;episode 9&lt;/a&gt; measured at 102 lines keeps writing anyway.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;All 9 documents were read live on Aug 27, 2026: 4 lesson files in the memory directory (in 2 different filename spellings), a failed-test note, 3 root-level volumes, and the 901-line archive — first timestamp Mar 14 12:25, final save Aug 19 11:48, 158 days end to end. Everything rendered in English here is paraphrase; only structure, numbers, and star ratings are carried over. The sensitive story is deliberately unnamed; the human reviewer&amp;#39;s countersignature on the March volume is deliberately unsigned here.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;memory/lesson_learned_20260314.md&lt;/code&gt; and its 3 siblings, &lt;code&gt;memory/table_extraction_test_failed.md&lt;/code&gt;, &lt;code&gt;LESSON_LEARN_REVIEW.md&lt;/code&gt;, &lt;code&gt;LESSON_LEARNED_20260617.md&lt;/code&gt;, &lt;code&gt;LESSONS_LEARNED_NEWS_PUSH.md&lt;/code&gt;, &lt;code&gt;wechat-editor-team/LESSONS_LEARNED.md&lt;/code&gt;, &lt;code&gt;cron/jobs.json&lt;/code&gt;, 2026-03 through 2026-08.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from live files. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next in the war stories arc: the LangGraph affair — the framework this fleet adopted, trusted, and then disproved from its own checkpoints. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace failure archive · evidence = 9 dedicated failure documents (memory/ lesson files, root-level LESSON volumes, wechat-editor-team archive) + cron jobs.json payloads · timestamps Asia/Shanghai, read 2026-08-27. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep20-bug-diaries.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep20-bug-diaries.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Postmortem</category><category>LLM Reliability</category><category>OpenClaw</category><category>Personal Infrastructure</category></item><item><title>The LangGraph Autopsy: 12 Green Imports, 668 Silent Seconds, One 203-Line Survivor</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep21-langgraph-autopsy/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep21-langgraph-autopsy/</guid><description>E21: 12 green imports, 668 silent seconds — the autopsy that deleted a framework 90 minutes in, and the 203-line engine that outlived it.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 21 of One Man One Legion&lt;/a&gt; — the fourth war story. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep20-bug-diaries/&quot;&gt;Episode 20&lt;/a&gt; read the failure library; this one is the fleet&amp;#39;s own favorite failure — the workflow framework adopted one morning and disproved within ninety minutes, told from the machine that ran it.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;The bet&lt;/h2&gt;
&lt;p&gt;The orchestrator was born with a constitution: machines exchange structured work state, never chat; a workflow is a finite state machine with persistence, not a conversation. For that, LangGraph plus sqlite was the chosen shape — langgraph 1.2.11 and langgraph-checkpoint-sqlite 3.1.1 into a fresh venv — and by 12:35 on Aug 12, 2026 the work log recorded the skeleton complete: 12 modules with every import green, 9 named nodes from prompt parsing to error handling, conditional routing, a sqlite checkpointer with its tables already built.&lt;/p&gt;
&lt;p&gt;The architectural bet, in the log&amp;#39;s own terms: a GPU render takes minutes, but a single execution must stay under a few seconds. So the dispatch node submits the job and hits END — the thread terminates, state parked at stage=polling — and the next scheduler tick resumes polling from the checkpoint. Long tasks segmented into short hops. On paper, exactly what the framework was for.&lt;/p&gt;
&lt;h2&gt;668 seconds&lt;/h2&gt;
&lt;p&gt;At 13:33 the same day, test T5 ran the first true end-to-end flow. Submission succeeded — the GPU box&amp;#39;s own history later confirmed the render finished, prompt_id 6f476a70; a picture existed. The orchestrator sat in polling for 668 seconds and never advanced. The worker had done its work; the state machine never learned.&lt;/p&gt;
&lt;p&gt;The autopsy was a minimal graph, not a debate. Two invocations of a toy entry-worker graph produced the identical call list [&amp;#39;entry&amp;#39;,&amp;#39;worker&amp;#39;] both times: in LangGraph, END is a thread&amp;#39;s terminal state, and invoking None against a finished thread is a no-op. There is no resume. The design was not buggy — it was physically impossible, and poll_comfy had never once executed after submission. The log&amp;#39;s verdict on the morning&amp;#39;s green skeleton was one word: mirage. Twelve passing imports had tested syntax, never semantics.&lt;/p&gt;
&lt;h2&gt;The deletion, and the receipt&lt;/h2&gt;
&lt;p&gt;The human picked the surgical option: keep the node functions and the database, delete the framework layer. The stated reasons mattered as much as the bug — the pipeline was nearly linear, and of all the capabilities that justified the framework (branching, human-in-the-loop, checkpointing, visualization), not one was actually in use. Capacity pre-built for a demand that never came. By 14:05 the log recorded the layer gone, along with its two orphaned checkpointer tables; the database kept its 4 business tables. In their place, a hand-rolled engine: drive() walks the stages in sequence, and the long GPU wait is an ordinary while loop — the loop absorbs the long task, so there is no resume semantic left to get wrong. Crash recovery reads facts, not snapshots: rebuild state from the task row, then re-query the renderer&amp;#39;s own history by prompt_id.&lt;/p&gt;
&lt;p&gt;The receipt came at 13:59 — 26 minutes after the stall was logged — when the same flow completed in 163 seconds: submit, poll, fetch, deliver, finalize, one image back. The audit printed a poll_done line for the first time — the poll node had never run under the framework. The next morning at 11:16, a human&amp;#39;s Hi on Telegram got its image back in 27 seconds, confirmed by eye. From skeleton-complete to layer-deleted: 90 minutes of load-bearing life.&lt;/p&gt;
&lt;h2&gt;What the injection found&lt;/h2&gt;
&lt;p&gt;Fault injection the same afternoon — reconstructed later from a 612-event session transcript after the session was stopped before logging — turned up 4 bugs with one shared root: the fatal segmentation (already dead with the deleted layer), a stage field that stopped syncing to the database mid-run, a delivery step that incremented its sent counter without checking whether the send actually succeeded, and a poller with zero tolerance for a single transient SSH hiccup. The root, named in the log: zero tolerance — one transient failure equals death — plus false done, reporting success without checking the downstream return.&lt;/p&gt;
&lt;p&gt;The scar hardened into rules. Two principles entered the standing constitution the next morning: import smoke is not a run, and a happy path is not validation — without fault injection, it didn&amp;#39;t happen. And the day after the deletion, a vector-RAG proposal flared with the identical root error, chosen because it was the standard thing before anyone asked what question it answered — which produced the form-follows-problem selection rule, with LangGraph as its canonical counterexample.&lt;/p&gt;
&lt;h2&gt;The tombstones&lt;/h2&gt;
&lt;p&gt;Read live on Aug 27: the orchestrator&amp;#39;s package code carries 0 langgraph imports. Exactly 2 mentions survive, both comments — one on the nodes file declaring them pure functions bound to no framework, one on the database file explaining where the checkpointer tables went. The pip package itself still sits in the venv with nothing importing it: deleted code, undisturbed dependency. And the 203-line engine now routes 3 kinds of work — image, digest, retrieval Q&amp;amp;A. The survivor didn&amp;#39;t just replace the framework; it outgrew the job the framework was hired for.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;All timings are as recorded in the work log for Aug 12–13, 2026, UTC by the log&amp;#39;s own convention. The run measurements — 668 seconds, 163 seconds, 27 seconds — are log records, not re-runs. The code counts (203 lines, 0 imports, 2 comment tombstones, 4 business tables) were read live from the orchestrator on Aug 27, 2026. Chinese log prose is paraphrased in English; only machine strings — [&amp;#39;entry&amp;#39;,&amp;#39;worker&amp;#39;], poll_done, END, prompt_id 6f476a70 — are carried verbatim. No ports, tokens, or chat identifiers appear.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;~/WORKLOG.md&lt;/code&gt; entries of 2026-08-12 and 2026-08-13; &lt;code&gt;~/orchestrator/engine.py&lt;/code&gt;, &lt;code&gt;nodes.py&lt;/code&gt;, &lt;code&gt;db.py&lt;/code&gt;, and venv site-packages, read 2026-08-27.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from live files and the work log. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next in the war stories arc: the security handcraft — key discipline and dependency gates across the fleet, whose torch.load specimen already ran as a standalone piece. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · AWS brain-node orchestrator (~/orchestrator) · evidence = WORKLOG.md entries of 2026-08-12/13 (UTC convention) + live code and venv read 2026-08-27 (engine.py, nodes.py, db.py, site-packages) · run timings 668 s / 163 s / 27 s as recorded in the log, not re-run today. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep21-langgraph-autopsy.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep21-langgraph-autopsy.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>LangGraph</category><category>Architecture</category><category>Postmortem</category><category>Workflow Orchestration</category><category>Personal Infrastructure</category></item><item><title>The Security Handcraft: 12 Fossil Scripts, One Live Key, Zero Leaks in Print</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep23-security-handcraft/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep23-security-handcraft/</guid><description>E23: the leak scanner turned on the fleet itself — a 138-field single home, 12 legacy scripts sharing one inline token, and 2 live strings for one bot.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 23 of One Man One Legion&lt;/a&gt; — the fifth war story. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep21-langgraph-autopsy/&quot;&gt;Episode 21&lt;/a&gt; autopsied a framework that died in ninety minutes; this one turns the fleet&amp;#39;s own leak scanner around and points it at the machines.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;The ritual&lt;/h2&gt;
&lt;p&gt;Every episode of this series ends the same way. Before anything is pushed, both files — article and ledger — are pattern-scanned for the fleet&amp;#39;s red-line items: token shapes, group ids, ports, proxy addresses. The work log carries 5 receipts of that single line being written. And today&amp;#39;s whole-library run covered all 23 published posts in the repository: 0 hits for Telegram-token, sk-, ghp-, or AKIA-shaped strings.&lt;/p&gt;
&lt;p&gt;This episode asks the obvious next question. What happens if you turn that scanner around — away from the articles, onto the fleet itself?&lt;/p&gt;
&lt;h2&gt;Where the keys live&lt;/h2&gt;
&lt;p&gt;The rule as built: one credentials home per machine, read at runtime, never inlined. On the workstation, that home is a single config file of 138 fields, of which 6 are credential-shaped — 3 apiKey entries, 1 appSecret, 1 botToken, 1 token — sitting behind mode 600, owner-only. Modern code doesn&amp;#39;t copy from it; the image factory&amp;#39;s helper function loads the bot token fresh on every push and skips quietly if the read fails.&lt;/p&gt;
&lt;p&gt;The AWS orchestrator repeats the pattern with more machinery: a dedicated secrets file, mode 600 verified today, environment-first with file fallback, and a logging filter that replaces any token substring with asterisks before a line reaches a file or a console — its header calls this constitution clause 4 — keys never enter logs. The account&amp;#39;s two inference environment variables live in a shell profile by name only.&lt;/p&gt;
&lt;h2&gt;How keys move&lt;/h2&gt;
&lt;p&gt;Never in a command line, never across the wire. One incident set the doctrine. On Aug 11, during the fleet-monitor build told in &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep5-fleet-on-one-screen/&quot;&gt;Episode 5&lt;/a&gt;, a bot token appeared inside a conversation. The work log&amp;#39;s warning line recorded the exposure and prescribed the response in the same breath: the human re-issues the token at the bot authority, then places the new string into the 600-permission file by hand over ssh — a replacement credential never crosses a chat again. That doctrine is why this episode&amp;#39;s audit compared hashes, never strings.&lt;/p&gt;
&lt;h2&gt;The audit that found something&lt;/h2&gt;
&lt;p&gt;A pattern scan across the workstation&amp;#39;s code matched secret-shaped strings in 16 files. 4 of those were upstream test fixtures — a packaged retrieval project ships fake keys in its own test suite; a scan&amp;#39;s hits are not all yours. The other 12 were ours: the legacy push lane. News-push v3, v4, and v5, two hotspot-monitor skills, two social-trend monitors, a reddit monitor, and video-batch helpers — every one of them carrying an identical 46-character token string, byte for byte. The single-home rule arrived after these scripts were written, and nobody went back.&lt;/p&gt;
&lt;p&gt;The fossil was supposed to be dead — supposed to be. No record this series has ever read tested that. Today&amp;#39;s getMe probe, run on the workstation through its own proxy so neither string left the machine: the config&amp;#39;s token answered ok=true, and so did the fossil — same bot, 2 distinct live strings. The config only knows about one of them. A revoke kills a string, and a live string is its own proof: no revoke ever reached this lane, though one was recommended once, on Aug 11, for a different leaked token. The fossil authenticated fine today.&lt;/p&gt;
&lt;p&gt;Why it survived is also why the cleanup is finally cheap: since &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep19-cron-contraction/&quot;&gt;Episode 19&lt;/a&gt;&amp;#39;s contraction, no scheduled job invokes any of the 12. They are fossils with a pulse — dark cron-wise, alive credential-wise. The remediation order writes itself: inventory, migrate or delete, revoke, rescan.&lt;/p&gt;
&lt;h2&gt;The leaks that didn&amp;#39;t reach print&lt;/h2&gt;
&lt;p&gt;Twice, this series&amp;#39; own evidence-gathering printed a secret into a tool session — a config dump that carried an appSecret during the pen-name audit, a parent-object serialization that carried an apiKey during the brain-swap dig (&lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep10-pen-names/&quot;&gt;Episode 10&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep18-brain-swap/&quot;&gt;Episode 18&lt;/a&gt;). Both were caught in session, neither entered print, and both final files scanned clean. The recurring lesson is uncomfortable: the most likely leak tool in a transparent infrastructure is your own forensic code.&lt;/p&gt;
&lt;h2&gt;Posture, briefly&lt;/h2&gt;
&lt;p&gt;The network side is told elsewhere and hasn&amp;#39;t changed: public SSH closed since Aug 9, render submissions gated behind a two-entry port whitelist with path checks. The one CVE war story — a torch.load security gate whose instructed upgrade produced five verbatim breakages in one morning — was &lt;a href=&quot;/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/&quot;&gt;published separately&lt;/a&gt; and stands where it stood.&lt;/p&gt;
&lt;p&gt;Security in this fleet is not a product that was bought. It is a habit of scanning itself, the same grep that keeps 23 published posts clean having just found a live fossil on the machines. Transparency is paid for in redaction — and occasionally redeemed by it.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — one man, one legion, and the discipline that lets it publish its own logs.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · 3-machine fleet, all evidence gathered 2026-08-27 · workstation: config field-path counts (values never printed), source-code pattern scans, and on-box getMe probes through the fleet&apos;s own proxy — token strings compared by hash and never left the machine · AWS: orchestrator secrets/logging code read live, permission bits verified · publication scan across all 23 posts in this repository. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep23-security-handcraft.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep23-security-handcraft.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>Security</category><category>Secrets Management</category><category>Infrastructure Audit</category><category>Personal Infrastructure</category></item><item><title>The Token Ledger: 261,720,417 Tokens, 147 Re-reads per Token Written</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep24-token-ledger/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep24-token-ledger/</guid><description>E24: the legion&apos;s first honest ledger — 208 days, 261,720,417 metered tokens, 84.5% cache re-reads, 0.68% output, and a 75-day silence the bill remembers.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 24 of One Man One Legion&lt;/a&gt; — the Ledger arc opens. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep23-security-handcraft/&quot;&gt;Episode 23&lt;/a&gt; audited the keys; this one counts what the legion actually consumes.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;A ledger nobody was reading&lt;/h2&gt;
&lt;p&gt;The agent runtime has been writing a bill all along. Every session leaves a transcript on disk, and every assistant turn in it carries a usage record: input, output, cacheRead, cacheWrite, totalTokens — and a cost field. That cost field is itself a finding: 2,689 of the 3,931 metered turns carry a nonzero cost, every one of them from February or March, summing to 52.27 in a unit the records never name — the currency field is null. Every turn since the June restart reads exactly zero. Not because usage turned free: whether a price table was removed or the new provider path simply stopped itemizing, the records don&amp;#39;t say. The currency column is half-blind; the volume columns are exact. Today&amp;#39;s count covers the main agent&amp;#39;s entire archive: 93 session transcripts, from February 1 to August 27 — 208 days — and this sum is its first full reading.&lt;/p&gt;
&lt;h2&gt;The headline and the composition&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;261,720,417 tokens.&lt;/strong&gt; The components: 38,875,077 fresh input, 1,769,499 output, 221,075,841 cache reads, and a cacheWrite column that has never registered a single token. The arithmetic closes — those four numbers re-sum to the total exactly.&lt;/p&gt;
&lt;p&gt;The surprise is the composition. Cache reads are 84.5% of everything; fresh input is 14.9%; output — the articles, digests, and pushes, the actual product — is 0.68%. For every token this legion writes, it re-reads 147: 259,950,918 tokens of context flowed in to produce 1,769,499 tokens out. The average metered turn carried roughly 66,580 tokens of context and answered in about 450. A one-man legion is not, it turns out, a writing machine. It is a reading machine that occasionally writes.&lt;/p&gt;
&lt;h2&gt;The shape of the burn&lt;/h2&gt;
&lt;p&gt;By month: February 91,168,392 across 9 sessions and 1,302 turns; March 96,182,265 across 28 sessions and 1,391 turns. Those two founding months total 187,350,657 — 71.6% of the entire ledger — and they carry the only explicit currency bill: February 24.24, March 28.03. Then June 7,946,646, July 34,679,955, August 31,743,159 through the 27th. February ran 91,168,392 tokens across 28 days — over three million a day — while the current contracted fleet sits near 1.2 million. The burn rate fell by roughly two-thirds without the output stopping.&lt;/p&gt;
&lt;p&gt;The concentration is person-shaped, not fleet-shaped. The single largest session began February 19: one marathon of 298 turns that consumed 31,989,653 tokens — 12.2% of everything the legion has ever burned, averaging 107,350 tokens per turn. The top three sessions together account for 69,171,676, or 26.4%. Long interactive sessions with a huge, re-read context were the founding era&amp;#39;s signature; the scheduled fleet that replaced them is cheap by comparison.&lt;/p&gt;
&lt;h2&gt;The silence the bill remembers&lt;/h2&gt;
&lt;p&gt;April and May: zero. Not low — zero. The last March entry is timestamped March 29 at 12:53 UTC, two minutes after the glm-5.1 brain swap that &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep18-brain-swap/&quot;&gt;Episode 18&lt;/a&gt; dated to 12:51 that day. The first entry back is June 12 at 21:33 UTC — the same day the agent&amp;#39;s diary resumed, per &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep6-agent-diary/&quot;&gt;Episode 6&lt;/a&gt;. That is 75 days of silence, independently present in three record systems that were never designed to agree: the token bill, the brain-swap backups, and the diary. When one man runs the whole legion, even the meter corroborates the history.&lt;/p&gt;
&lt;p&gt;The model lineage is visible in the same records: glm-4.7 carried February at 73,933,400 tokens, glm-5 carried March at 76,509,096, and after the silence glm-5.1 logged just 49 turns (3,208,162) before glm-5.2 took the primary seat for 51,645,016 across three months. glm-5.3, primary since August 15, has 189 turns and 11,418,201 so far. An honest caveat: 586 turns totaling 44,334,074 — 16.9% of the ledger — carry no model header and stay unattributed.&lt;/p&gt;
&lt;h2&gt;The local lane&lt;/h2&gt;
&lt;p&gt;The fallback chain — cloud first, local 9B second — appears in this ledger exactly 17 times: 672,468 tokens, 0.26% of the total, every one of them in June, during the restart. Insurance that almost never fires is still insurance worth having, but the local GPU&amp;#39;s real work lives elsewhere: image and video renders are duration-metered in &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep1-image-workshop/&quot;&gt;Episode 1&lt;/a&gt; and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep7-video-workshop/&quot;&gt;Episode 7&lt;/a&gt;, and the 27B serving Claude Code in &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep14-model-stable/&quot;&gt;Episode 14&lt;/a&gt; rides a separate harness on a local card — off this cloud ledger entirely, paid in watts rather than tokens.&lt;/p&gt;
&lt;p&gt;So what does one man&amp;#39;s legion cost? In currency, the records say exactly one thing: 52.27 of unnamed unit, all of it from the founding era — the modern fleet&amp;#39;s bill is unreadable from inside. In volume: a quarter of a billion tokens to date, growing at about 1.2 million a day, of which the part anyone would call the product is less than one percent. The lesson the ledger teaches is the same one the diary and the contraction taught: the expense of an agent is not what it says. It is what it must be reminded of, every single turn.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — one man, one legion, and now one honest bill.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · workstation, all evidence gathered 2026-08-27 · the main agent&apos;s 93 session transcripts parsed on-box, usage fields summed across 3,931 metered assistant turns, per-month and per-model cross-tabs computed, whale-session totals re-read individually, cost fields audited on every turn (2,689 nonzero, all February–March, totaling 52.27, currency field null) · component arithmetic re-verified: 38,875,077 + 1,769,499 + 221,075,841 + 0 = 261,720,417. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep24-token-ledger.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep24-token-ledger.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>Token Economics</category><category>AI Agents</category><category>LLM Usage</category><category>Personal Infrastructure</category></item><item><title>The Human Ledger: 26 Drops, 191 Pieces Behind One Phone, Five One-Line Rulings</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep25-human-ledger/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep25-human-ledger/</guid><description>E25: the other side of the ledger — 26 drops in 171 days, 191 pieces capped by a phone tap, five one-line rulings, every act receipted by a machine.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 25 of One Man One Legion&lt;/a&gt;. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep24-token-ledger/&quot;&gt;Episode 24&lt;/a&gt; counted what the machines burned — 261,720,417 tokens. This episode opens the other cover of the same book: who kept the receipts on the one human, and what do they total?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A legion&amp;#39;s ledger has two sides. The machine side counts itself: every token metered, every turn stamped, 147 re-reads per write. The human side should be blank — humans don&amp;#39;t log themselves. But when this fleet was audited today, the human side wasn&amp;#39;t blank. The machines had been keeping it all along.&lt;/p&gt;
&lt;h2&gt;The machines stamp themselves first&lt;/h2&gt;
&lt;p&gt;The veteran WeChat line keeps 61 per-article quality-gate records — 46 on the articles line, 15 on the daily line. Each holds the same shape: &lt;code&gt;passed: true&lt;/code&gt;, a retry round, &lt;code&gt;machine_checks: 7, machine_passed: 7&lt;/code&gt;. The retry rounds distribute 42 first-pass, 8 second, 8 third, 3 fourth. Not one record carries a field for a human. On this production line the seven machine checks &lt;em&gt;are&lt;/em&gt; the gate; the human is not a check — the human is the exit.&lt;/p&gt;
&lt;p&gt;That asymmetry is the ledger&amp;#39;s first finding: the machines receipt their own labor in structured rows, and receipt the human&amp;#39;s labor in their diaries.&lt;/p&gt;
&lt;h2&gt;Touch one: 26 gestures in 171 days&lt;/h2&gt;
&lt;p&gt;Whatever the operator drops into a conversation, the chat runtime files it. The inbound folder holds 26 entries spanning February 19 to August 8 — 15 jpg, 6 png, 2 voice clips, 2 PDFs, and one extension-less investigation bundle. No form, no fields; whatever lands is the request. At 26 entries over a 171-day window, touch one costs the operator about one gesture a week. (Episode 4 called these 26 jpg files; today&amp;#39;s by-extension recount says 15 jpg among 26. The correction is recorded here, not retroactively.)&lt;/p&gt;
&lt;h2&gt;Touch two: three interactions, then graduation&lt;/h2&gt;
&lt;p&gt;Before the series moved to standing authorization, the gate slip was the human&amp;#39;s unit of work — budgeted at 5 minutes, in Chinese, numbers against source sentences. The full history is short: the fictional-article drill on August 25, the first real slip on August 26 (27 items, 21 machine-verified plus 6 the operator resolved by running a terminal grep personally), and one yes on Episode 1 before push. Three gating interactions, per the record in &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep4-three-touchpoints/&quot;&gt;Episode 4&lt;/a&gt; — and then the slip section of the ledger stopped growing, by design. The queue now ships through machine gates; the human&amp;#39;s three interactions stand as the full cost of ever having gated by hand.&lt;/p&gt;
&lt;h2&gt;Touch three: 191 pieces behind one phone&lt;/h2&gt;
&lt;p&gt;The archive holds 191 finished html pieces — 72 articles, 105 daily, 14 archive — with modification dates from June 15 to August 27, the newest written today. Zero of them can leave through code: a full-tree grep today finds 0 freepublish call sites, so the API ceiling is the draft box and the publish button is physically a phone. Each piece that shipped, shipped through one tap. Alongside sit 4 paste-ready distribution packages, Reddit and Hacker News — 0 clicked, queued for a hotter news cycle. Touch three is the only touch the human cannot delegate: identity lives on the phone.&lt;/p&gt;
&lt;h2&gt;Five rulings, all receipted by machines&lt;/h2&gt;
&lt;p&gt;The rarest entries: one-line verdicts that reversed a machine&amp;#39;s conclusion. August 11 — a leaked token was called out; the fleet rotated the key and wrote the ssh-self-fill doctrine. August 13 — one principle from the operator, three machines, an overnight of work. August 15, 21:17 — a brain-swap ritual was judged redundant, in the operator&amp;#39;s words an overdone step; the named pre-swap backup from 21:05 still survives on disk, though the rolling backups from that night have since eaten their own fossils. August 19, 11:47 — the agent diary&amp;#39;s line 25 receipts a root-cause reversal, verified today. August 25 — the contract terms themselves. Five rulings in 208 days, none kept by the human, all recoverable only because a machine wrote them down.&lt;/p&gt;
&lt;h2&gt;The arithmetic&lt;/h2&gt;
&lt;p&gt;Sum the human column: 26 drops, 3 gate interactions, 191 pieces each capped by a tap, 5 rulings — roughly 225 logged acts. Against Episode 24&amp;#39;s 261,720,417 tokens and 3,931 turns across the same 208 days, that is over 1.1 million tokens per human act (1,163,202 to the nearest act, and the true ratio can only be higher if fewer pieces have shipped), about 17 machine turns per act, about one logged act per day. The fleet re-reads everything 147 times; the human reads almost nothing once — one tap and done.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Measured today: the 26-entry recount and its window, the 191-piece by-extension recount (72+105+14) and its mtime span, the 61-record histogram, the 0-freepublish grep, the diary line 25, the surviving named backup. From published record, not re-measured: the 3 gate interactions and 4 unclicked packages (Episode 4), the August 11/13/15 rulings (Episodes 23, 5, 18). Two inferences are labeled: modification time is not publish time, so taps are an architecture-implied ceiling, not a count; and the 225 total treats every finished piece as one tap — an upper bound on acts, therefore a floor on the ratio.&lt;/p&gt;
&lt;p&gt;Primary sources: chat-runtime &lt;code&gt;media/inbound/&lt;/code&gt;, &lt;code&gt;wechat-editor-team&lt;/code&gt; archive and gate records, agent diary &lt;code&gt;memory/2026-08-19.md&lt;/code&gt;, config backup fossils, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep24-token-ledger/&quot;&gt;Episode 24&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;All episodes&lt;/a&gt; — One Man One Legion, an ongoing series.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Fleet-wide · evidence = workstation chat-runtime inbound recount (26 entries, 2026-02-19 to 08-08), wechat-editor-team archive recount by extension (72+105+14 html, mtimes 06-15 to 08-27), 61 article-*.gate.json records with retry-round histogram, full-tree freepublish grep = 0 call sites, agent diary 2026-08-19.md line 25, openclaw.json backup-before-glm53 mtime 08-15 21:05; slip and ruling counts cross-referenced from published Episodes 4/5/18/23/24 · read 2026-08-27. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep25-human-ledger.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep25-human-ledger.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Human-in-the-Loop</category><category>Workflow</category><category>Personal Infrastructure</category><category>Automation</category></item><item><title>The Loop Closes: What One Man and One Legion Answered to a Turbulent Era</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep26-closing-the-loop/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep26-closing-the-loop/</guid><description>E26, the finale: 24 posts, 121 ledger entries, 261,720,417 tokens against roughly 225 human acts — the series&apos; measured reply to Gates&apos; warning.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;The final episode of One Man One Legion&lt;/a&gt; — the Ledger arc closes, and with it the series. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep25-human-ledger/&quot;&gt;Episode 25&lt;/a&gt; counted the human; this one closes the loop the anchor opened.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;The series audits itself&lt;/h2&gt;
&lt;p&gt;The anchor opened with Bill Gates&amp;#39; August 26, 2026 essay and its warning that the AI transition will be &amp;quot;one of the most turbulent times in human history&amp;quot; — and that &amp;quot;right now we are not preparing for it.&amp;quot; The macro argument is everywhere, the anchor said; this series would publish the measured micro answer. This is the last episode, and its evidence pack is the series itself: 24 posts — one anchor and 23 episodes — across 5 arcs, The Workshops 7, War Stories 5, The Engine Room 4, The Cockpit 4, The Ledger 3. Three episode numbers, 8, 15, and 22, were never issued, so the 26th episode by numbering is the 23rd by count — a fitting last correction for a series whose first lesson was that plurals get mistaken for singulars. The corpus carries 21,818 body words and 121 entries in the public measurement ledger, every one mirroring a number printed in a post body. All 24 posts carry the same date: the fleet wrote its own audit in a day, under the batch authorization that followed the first episode&amp;#39;s human gate.&lt;/p&gt;
&lt;h2&gt;What the legion made, burned, and was spared&lt;/h2&gt;
&lt;p&gt;Production, in one paragraph. The text workshop archived 191 finished pieces — 72 articles, 105 daily, 14 archive — every one capped by the WeChat draft box and shipped through a single phone. The image workshop answered chat phrases in 50 s warm, 330–379 s on the cold model load. Video, podcast, and analyst lines each got their measured episode. And the honest edge: 4 paste-ready distribution packages sit at 0 clicked. The factory works; the audience is still a manual step.&lt;/p&gt;
&lt;p&gt;Consumption, in one paragraph. Over 208 days the main agent metered 261,720,417 tokens — 84.5% of them cache re-reads, 0.68% output, a 147:1 read-to-write ratio. The founding two months burned 71.6% of everything; the only currency figure the records carry is 52.27 in a unit they never name, and every cost entry since the June restart reads zero. The meter went dark; the series said so in print.&lt;/p&gt;
&lt;p&gt;The human column, in one paragraph. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep25-human-ledger/&quot;&gt;Episode 25&lt;/a&gt; counted roughly 225 logged acts — 26 drops of source material, 3 gate interactions, 191 tap-capped pieces, 5 one-line rulings — against the machine ledger: 1,163,202 tokens per act to the nearest act, about one human act per day across 208 days.&lt;/p&gt;
&lt;h2&gt;Where it broke is what it knows&lt;/h2&gt;
&lt;p&gt;The war-stories arc holds the series&amp;#39; best material, which is another way of saying the failures were the material. 18 of 22 cron jobs are off, after 790 scheduled runs produced 154 errors nobody read. The LangGraph orchestration layer was deleted 90 minutes after its skeleton compiled, killed by a minimal-graph experiment that proved its core design physically impossible. The bug-diary corpus runs 9 files across 158 days, and its Top 5 lesson needed 77 days to fail its first re-test. The cost meter went blind and nobody noticed for two months. A series that only printed wins would have none of this; a fleet that cannot print losses has no memory.&lt;/p&gt;
&lt;h2&gt;The loop back to Gates&lt;/h2&gt;
&lt;p&gt;Gates proposed 3 fixes. Each one meets a measurement somewhere in this corpus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax AI tokens.&lt;/strong&gt; Before you can tax tokens you must be able to count them. At the scale of one operator with full disk access to every transcript, the volume columns close to the token and the currency columns still went dark — 52.27 in an unnamed unit, then zeros. Multiply that ambiguity by every inference provider and API reseller on earth before drafting the levy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human Reserved jobs.&lt;/strong&gt; At n=1 the reserve already exists, and it is not a job category. The human reserved 3 gestures — drop material, fill the gate slip, tap publish — plus 5 one-line rulings that reversed machine conclusions across 208 days. What stayed human was judgment at the exits, not labor in the pipes; the labor went to the legion and the receipts went to the ledger.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A new institution, because &amp;quot;we are not preparing.&amp;quot;&lt;/strong&gt; What preparation looked like here was not a plan but bookkeeping: 97 diary files (63 daily entries, 34 topical notes), a quality-lessons loop with 102 entries, 61 per-article gate stamps, timestamped config backups, and a public measurement ledger the machines can read back through the site&amp;#39;s own interface. None of it was designed as one system; all of it accumulated because the fleet wrote down what happened. That is the transferable finding — not the fleet, the habit.&lt;/p&gt;
&lt;h2&gt;The answer, and the light going off&lt;/h2&gt;
&lt;p&gt;Is one man with one legion an answer to a turbulent era? The measured answer is: it is an instrument. It cannot calm the era, but it can keep honest records through it — and 208 days of receipts beat any number of plans. The turbulent-era question was macro; this series&amp;#39; contribution is that even the macro answer will be assembled from micro ledgers like this one, or not at all.&lt;/p&gt;
&lt;p&gt;One last self-reference, printed because it is true and checkable: this finale was drafted by the fleet&amp;#39;s own writer on the batch lane, and the queue it emptied retires that writer&amp;#39;s own schedule. When this session&amp;#39;s sandbox proved unable to reach the crontab, the removal became the fleet&amp;#39;s last logged instruction — the line sits in the worklog, waiting for the operator&amp;#39;s shell. The legion ends its first book the way it ends every over-automated job: by switching itself off, and leaving the entry in the ledger.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Start at the beginning: &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;the anchor&lt;/a&gt;. The full run — every arc, every entry — is on the &lt;a href=&quot;/series/one-man-one-legion/&quot;&gt;series index&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · this site&apos;s own repository, audited 2026-08-27 · series statistics computed from the published corpus itself (post count, arc count, frontmatter-stripped word totals, measurement-ledger entry count), every cross-referenced figure re-grepped verbatim from the published episode bodies on the day of writing. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep26-closing-the-loop.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep26-closing-the-loop.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Personal AI Infrastructure</category><category>Measurement</category><category>Series Finale</category></item><item><title>What Happens When the Cockpit Itself Goes Dark? Two Outages, Ten Stranded Jobs, One Spare Parked by Subtraction</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep3-chat-cockpit/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep3-chat-cockpit/</guid><description>E3: the chat channel died twice — 90-second stalls, then 10 jobs stranded at deliver; the fix was a 6-node pipe, and the spare parked for good.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 3 of One Man One Legion&lt;/a&gt;. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep2-tg-gateway/&quot;&gt;Episode 2&lt;/a&gt; walked the tool bench behind the wheel — everything the copilot can do once a message arrives. This episode is about arrival itself: the channel, and the 2 days it died.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every chat-commanded legion has a single point of failure more basic than any model: the pipe the commands travel through. This is the story of that pipe&amp;#39;s 2 deaths, what each one changed, and how the spare built to end the dependency was parked — not by failing, but by subtraction.&lt;/p&gt;
&lt;h2&gt;June 20: the wheel goes silent&lt;/h2&gt;
&lt;p&gt;The first symptom was silence: messages to the bot got no answer, and the gateway log repeated itself every 90 seconds — polling stall, no updates fetched, restart, stall again. The bot was not broken; its lifeline was. Every fetch of new messages left the country through a proxy whose single exit node had ceased to exist — 4 independent DNS resolvers (Ali, Cloudflare, Google DoH, 8.8.8.8) all returned the same verdict, NXDOMAIN. One dead domain, and the whole cockpit went dark.&lt;/p&gt;
&lt;p&gt;The diagnosis nearly went wrong. The machine carried an ambient proxy setting that silently routed every overseas health-check through the very proxy that was dead, so the first tests &amp;quot;proved&amp;quot; the box could not reach the internet at all. Bypassing it showed the truth: the direct network was fine. The operator wrote the rule down — unset the ambient proxy before diagnosing network problems — but it generalizes: your diagnostic tool lives inside the system it diagnoses, and inherits its failures.&lt;/p&gt;
&lt;p&gt;The fix that shipped was not &amp;quot;a better node&amp;quot; but &amp;quot;no more single node&amp;quot;: a pool of 6 nodes, measured latencies 0.44 to 4.14 seconds, a health probe every minute, a least-ping balancer, and a named fallback for the day every probe fails. The pipe became redundant. The channel stayed Telegram.&lt;/p&gt;
&lt;h2&gt;July 4: stranded at deliver&lt;/h2&gt;
&lt;p&gt;14 days later the pipe died again — upstream this time — and the failure had a cruel shape: the legion kept working. 13 outbound jobs failed that day, 10 OpenClaw cron jobs and 3 system crontab scripts, and in every case the collecting, the writing, the synthesis had already succeeded. The products sat on disk while the last step — sending them to the group — timed out. Work done, delivery dead.&lt;/p&gt;
&lt;p&gt;Feishu was the designated lifeboat: a domestic websocket channel that needs no proxy, already warm — enabled, plugin installed, 3 news jobs running on it. The operator narrowed the day&amp;#39;s mission on purpose: do not migrate; diagnose and stress-test the one Feishu job that looked broken.&lt;/p&gt;
&lt;h2&gt;The four-row experiment&lt;/h2&gt;
&lt;p&gt;What settled it was a contrast table — same evening, same channel:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Envelope&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Feishu news, morning&lt;/td&gt;
&lt;td&gt;agentTurn + isolated&lt;/td&gt;
&lt;td&gt;ran 149 seconds, tokens spent, real summary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feishu news, evening&lt;/td&gt;
&lt;td&gt;agentTurn + isolated&lt;/td&gt;
&lt;td&gt;ran 112 seconds, same shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feishu WeChat, noon&lt;/td&gt;
&lt;td&gt;systemEvent + main&lt;/td&gt;
&lt;td&gt;10 seconds, empty session, zero tokens, summary = the instruction itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Telegram WeChat, noon&lt;/td&gt;
&lt;td&gt;systemEvent + main, character-identical payload&lt;/td&gt;
&lt;td&gt;270 seconds of real work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The identical envelope works on Telegram and idles on Feishu. That is a plugin behavior mismatch, not a configuration error: the Feishu side receives the event, starts no agent turn, and marks it done. And the bug had 2 more defects stacked beneath it: the job&amp;#39;s clock said 8:05 while its name said noon (the Telegram twin runs at 12:00), and the WeChat channel had 5 slots on Telegram — 06:00, 07:30, 12:00, 17:20, 18:40 — against 1 on Feishu.&lt;/p&gt;
&lt;p&gt;One honest footnote: the live stress test never ran. The gateway&amp;#39;s external CLI refused connections twice — the rule is stop after 2 — while the internal scheduler proved itself another way: restarted at 22:15, it fired the Feishu news job on time at 22:20. The clock worked even when the door didn&amp;#39;t.&lt;/p&gt;
&lt;h2&gt;The spare, driven briefly&lt;/h2&gt;
&lt;p&gt;The repair was applied. The noon job&amp;#39;s envelope was re-addressed to agentTurn + isolated and its clock aligned with its name; a one-shot diagnostic entry fired once on July 5 at 13:38. Then the spare drove: the repaired noon job logged 11 run records through July 11; the news pair kept twice-daily schedules to the last fires on July 22 and 23 — 19 records each — until the great contraction (&lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep19-cron-contraction/&quot;&gt;Episode 19&lt;/a&gt;) switched the auto-publishers off. The channel did not fail; the schedule was cut.&lt;/p&gt;
&lt;p&gt;The cockpit read from live configuration 35 days later: 2 enabled channels — Telegram with 12 groups, Feishu on websocket — but 0 of the 4 Feishu-routed cron entries are lit, and 3 of the 4 lit jobs route to Telegram. (The fourth&amp;#39;s name still points at a group while its key says main. Names drift; keys don&amp;#39;t lie.) The migration meant to save the cockpit never happened, and didn&amp;#39;t need to: the pipe got redundant, the schedule got cut, and the spare stayed mounted — warm, verified, and parked.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;The 2 outage narratives come from 2 dated diagnosis documents (June 20, 101 lines; July 4, 194 lines), cross-checked against live state read on August 27: channel enablement and the 12-group count from the runtime config, routing and enable flags from the cron file, the Feishu jobs&amp;#39; run-record counts and ranges from their logs, and the newest run — an arXiv digest, status ok — as evidence the channel is alive at press time. Run records are counted, not audited for delivery success. The July 4 document&amp;#39;s summary line says 13 OpenClaw jobs failed while its own table lists 10; we counted the rows. Latencies and the 270-second twin are the documents&amp;#39; measurements, not re-measured. No uptime is claimed for any channel or proxy.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;telegram-proxy-fix-2026-06-20.md&lt;/code&gt;, &lt;code&gt;feishu-migration-diagnosis-2026-07-04.md&lt;/code&gt;, &lt;code&gt;openclaw.json&lt;/code&gt; channels, cron &lt;code&gt;jobs.json&lt;/code&gt;, run logs, 2026-06 to 2026-08.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;All episodes&lt;/a&gt; — One Man One Legion, an ongoing series.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw runtime · evidence = openclaw.json channels + cron jobs.json + run logs + 2 dated diagnosis docs (2026-06-20, 2026-07-04; 101 and 194 lines) · timestamps Asia/Shanghai · read 2026-08-27. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep3-chat-cockpit.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep3-chat-cockpit.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Telegram</category><category>Feishu</category><category>Reliability</category><category>Personal Infrastructure</category></item><item><title>What Does the One Human Actually Do? Three Touchpoints per Piece, and One Wall That Won&apos;t Move</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep4-three-touchpoints/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep4-three-touchpoints/</guid><description>E4: the human appears in exactly 3 places per piece — drop material, tick a slip, click publish; on WeChat an API wall makes the phone the button.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 4 of One Man One Legion&lt;/a&gt;. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep3-chat-cockpit/&quot;&gt;Episode 3&lt;/a&gt; was the pipe commands travel through. This one is about the scarcer resource at the other end: the human, and the contract that budgets them.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every legion built by one person hits the same arithmetic: the machines scale, the human doesn&amp;#39;t. The fleet&amp;#39;s most load-bearing design is not a model but a contract deciding where the human appears — written down, tested by drill, and fitting in one line: exactly 3 touchpoints per piece.&lt;/p&gt;
&lt;h2&gt;A contract born in a drill&lt;/h2&gt;
&lt;p&gt;On August 25 the operator set the terms bluntly: the writing and the test records stay with me, everything else is yours. The answer was a drill before a promise: a fully fictional article (model name SyntheticTalk-7B, every number invented) through the whole pipeline, the operator role-playing the laziest material — a ragged Chinese paragraph and a screenshot description.&lt;/p&gt;
&lt;p&gt;The drill held: 3 follow-up questions, all in the three-things-worth-gold category — wall time, peak VRAM, where it stalled. A progress-bar reading of 83 percent became frame 201 of 240. The pipeline turned out 8 green artifact checks; the schema guardrail killed an over-long description on schedule. Four mechanisms went into the playbook that night: the gate slip, the paste-ready distribution package, the 3-touchpoint contract, and the two-form intake — raw records or the operator&amp;#39;s own draft, same pipeline.&lt;/p&gt;
&lt;p&gt;The 3 touches, as fixed: drop the material; fill the gate slip, budgeted at 5 minutes; click publish on the outside platforms. Everything else — follow-ups, arithmetic, English drafting, the ledger, the build, the push — is agent work. One clause points the other way: the operator never touches git, build, or deploy.&lt;/p&gt;
&lt;h2&gt;Touch 1: drop anything&lt;/h2&gt;
&lt;p&gt;The intake accepts any shape — raw logs, screenshots, Chinese paragraphs, a half-formed draft — and the discipline sits on the machine side: numbers get pulled out of screenshots, because a progress-bar percentage is hard data. The physical trace: the chat runtime&amp;#39;s inbound folder, where photos dropped into a conversation land, holds 26 jpg files, newest from August 8. An identity-lock image job asks for no path; it takes the newest file. The human&amp;#39;s gesture is the API.&lt;/p&gt;
&lt;h2&gt;Touch 2: the slip&lt;/h2&gt;
&lt;p&gt;The slip is the contract&amp;#39;s cleverest piece: the articles are English, the operator&amp;#39;s fact-memory is Chinese, so the slip is a Chinese number-crosscheck table — every number against the sentence it came from, the machine&amp;#39;s own inferences flagged. The operator ticks a yes or a no per row, never reading the English prose.&lt;/p&gt;
&lt;p&gt;The first real slip, August 26, shows the split inside the gate. The pressure test covered 27 items: 21 machine-verified against live hardware and source logs, 6 waiting on a terminal grep — which the operator ran personally on the workstation, sending back 2 raw-log confirmations of &lt;code&gt;Blocks 0-19 on GPU, 20-39 on CPU&lt;/code&gt;. 1 sourcing label was caught mislabeled and fixed first. The final slip showed 13 A-items and 6 B-items passed, exactly 4 C-items — pure fact arbitration — left for the human. What made it real was the rule: no slip, no push. The first article sat built and green until the slip came back.&lt;/p&gt;
&lt;h2&gt;Touch 3: the click&lt;/h2&gt;
&lt;p&gt;Distribution is prepared to the last keystroke: paste-ready packages with Reddit title and body, Hacker News title and first comment. 3 packages went out with the first Watch issue — r/LocalLLaMA, r/DeepSeek, r/ChatGPT — with a recommended firing order and a risk flag on the thinnest source; a 4th (r/webdev plus Show HN) followed that evening. None has been clicked: the operator is holding fire for a hotter news cycle. The contract does not schedule the human; it queues for them.&lt;/p&gt;
&lt;h2&gt;The mirror in the veteran factory&lt;/h2&gt;
&lt;p&gt;The WeChat factory, the fleet&amp;#39;s oldest line, enforces its own version not by discipline but by wall. Its checked-in code contains exactly 1 publish-adjacent call site, &lt;code&gt;draft/add&lt;/code&gt;, the official draft-box API, and 0 call sites of &lt;code&gt;freepublish&lt;/code&gt;. That 0 is not for lack of trying: a June 24 interactive attempt returned 48001 api unauthorized. Publishing happens from the phone, in the subscription-assistant app. The platform itself makes the human the publish button.&lt;/p&gt;
&lt;h2&gt;Why so few&lt;/h2&gt;
&lt;p&gt;The gate exists because an agent once tried to sed-edit a &lt;code&gt;torch.load&lt;/code&gt; safety check inside a working tree — told in &lt;a href=&quot;/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/&quot;&gt;the dependency-matrix dispatch&lt;/a&gt;. The human is the sole arbiter of technical fact, kept away from git, build, and deploy so they never become the bottleneck. The exceptions prove the rule is about identity: the search-console handovers could only happen in the operator&amp;#39;s own accounts (the Bing import, again about 5 minutes), and on brain-swap night the agent handed over a rollback card before restarting itself (&lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep18-brain-swap/&quot;&gt;Episode 18&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;The boundary also moves. When the series queue was authorized, the per-piece slip gave way to a standing authorization: a system cron wakes a queue worker on a fixed schedule, shipping episodes through machine gates — preflight and a post-write stress test — red lines unchanged. Episode 1 still got the operator&amp;#39;s yes; this episode did not. The contract didn&amp;#39;t break; it graduated.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Contract text, 5-minute budget, no-slip-no-push rule: from today&amp;#39;s playbook. Drill and first-slip numbers: from the fleet worklog, August 25–26 — 27 items split 21 plus 6; the final slip&amp;#39;s rows sum differently (13 + 6 + 4 = 23), both partitions reported as recorded, not reconciled. The WeChat figures were re-measured today: a full-tree grep found &lt;code&gt;freepublish&lt;/code&gt; only in chat logs, never in code; 48001 is an observed API error, the phone detail an agent&amp;#39;s session-record description — different grades, labeled as such. The 26-file count is today&amp;#39;s. The packages&amp;#39; firing order is from the worklog; the hold-for-hotter-cycle posture is the operator&amp;#39;s stated call, nothing posted since. The 5-minute Bing figure is a worklog estimate, not a stopwatch.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;docs/ai-citation-playbook.md&lt;/code&gt;, fleet WORKLOG 2026-08-25/26, workstation session log 2026-06-24, &lt;code&gt;wechat-mp-api.mjs&lt;/code&gt;, &lt;code&gt;media/inbound/&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;All episodes&lt;/a&gt; — One Man One Legion, an ongoing series.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Fleet-wide · evidence = site playbook (docs/ai-citation-playbook.md), WORKLOG gate-slip entries 2026-08-25/26, workstation session log 2026-06-24, wechat-mp-api.mjs full-tree grep, media/inbound count · read 2026-08-27. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep4-three-touchpoints.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep4-three-touchpoints.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Human-in-the-Loop</category><category>Workflow</category><category>WeChat</category><category>Personal Infrastructure</category></item><item><title>Six Doors, 231 Lines, and No Dashboard: Putting a Three-Machine Fleet on One Screen</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep5-fleet-on-one-screen/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep5-fleet-on-one-screen/</guid><description>E5: six ssh doors, a 231-line probe on a 5-minute clock, and 16 days of alerts — the fleet&apos;s one screen is a chat thread and a 12-key JSON, not a dashboard.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 5 of One Man One Legion&lt;/a&gt;. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep4-three-touchpoints/&quot;&gt;Episode 4&lt;/a&gt; fixed where the human appears. This one closes the Cockpit arc: where the fleet appears to the human — 3 machines on 3 networks becoming one screen, no dashboard ever built.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Three machines, three networks: a 512 MB VPS on the US West coast running the thin control plane, a dual-GPU workstation behind home NAT, a 4 GB always-on node in Seoul. Between August 10 and 13 they became a fleet commanded from one chair. The whole machinery: 6 ssh doors, a 58-line org chart, a 231-line probe, a 4-line heartbeat. Not one dashboard.&lt;/p&gt;
&lt;h2&gt;Six doors in one day&lt;/h2&gt;
&lt;p&gt;Claude Code instances have no native chat protocol, and the fleet never built one. The answer, fixed August 10: machine A ssh-es into machine B and runs headless &lt;code&gt;claude -p&lt;/code&gt; — B&amp;#39;s agent wakes with its own instructions and memory, answers in one shot, hangs up. Stateless as a phone call; context rides in the prompt.&lt;/p&gt;
&lt;p&gt;The day&amp;#39;s choreography is still readable in file timestamps: 08:33, the small box&amp;#39;s wrapper for calling the workstation; 09:12, its public key authorized on the Seoul node over Tailscale (public ssh had been closed August 9); 09:16, its wrapper for calling Seoul; 15:45, the workstation&amp;#39;s two wrappers back. Six directed doors for 3 machines — all 6 wrapper files verified on disk today, 4 stat at 334 to 912 bytes. This episode&amp;#39;s writing rode one of them: every workstation file quoted here crossed that rail.&lt;/p&gt;
&lt;p&gt;Two scars are welded into those tiny scripts. First: non-interactive ssh never sources &lt;code&gt;.bashrc&lt;/code&gt;, so the callee&amp;#39;s cloud-model credentials don&amp;#39;t load and headless Claude answers Not logged in — the fix evals the export lines on the far side, so the key never crosses the wire. Second: the 512 MB box can&amp;#39;t afford the call. Its own agent measured the bind: interactive Claude sits at 157 MB resident, leaving 77 MB available — less than the 99 MB a headless call wants. So the wrappers split in two: a 334-byte pure file-read for roughly 90% of traffic, zero memory; the true headless call for the rest, one at a time.&lt;/p&gt;
&lt;h2&gt;The map and the law&lt;/h2&gt;
&lt;p&gt;On August 11 the fleet wrote its own org chart — 58 lines, v3, drafted by the small box&amp;#39;s agent over the very channel it describes, after negotiating with the workstation&amp;#39;s. Its one physical law is memory asymmetry: heavy reasoning routes away from the small machine — brainwork places a call, and the Claude runs on 503 GB or 4 GB instead. The stress test that morning priced the law: 0 packet loss on both links, 140 ms and 89 ms direct — yet a Claude round trip of 24 s to the small box against 8–12 s to the workstation, about 2.4×. The wire was innocent; RAM was the toll. The test stayed gentle: no flooding, strictly single shots at the weakest node.&lt;/p&gt;
&lt;h2&gt;The screen&lt;/h2&gt;
&lt;p&gt;The cockpit itself landed the same day, inside 7 hours. 02:37, the first outbound self-check. 06:32, the stress test. 07:06, the probe: 231 lines of bash on the Seoul node, cron every 5 minutes, first run 9.9 s. It asks 12 keys — small-box memory, load, disk, gateway, Tailscale path; Seoul&amp;#39;s vitals; 3 external APIs; the orchestrator — and writes the fleet in 3 lines per round. That log file is, literally, the fleet on one screen. 07:10, a 4-line heartbeat on the workstation started pinging both peers every 2 minutes — the probe&amp;#39;s first run had caught its path riding a San Francisco relay at 936 ms, and once the heartbeat held the tunnel, it read direct at 90 ms. Today&amp;#39;s tail still shows that path direct, 176–181 ms. 07:19, alerts wired to Telegram. 07:38, the operator fixed 3 rules: the workstation leaves monitoring (a sleeping power machine is not news), the small box gets a memory threshold, and relayed paths count as faults. 09:42, the v2 rewrite: 4 severity states, run time down to 3.7 s.&lt;/p&gt;
&lt;p&gt;The rewrite also caught a geography bug: the small box&amp;#39;s Gemini probes returned 400, location not supported — geo-blocked where the wire lands; from Seoul the same API answered 403, reachable but keyless. Gemini monitoring moved to Seoul, its key piped over ssh into a 600-permission file — never displayed, never in a prompt.&lt;/p&gt;
&lt;h2&gt;Sixteen days of quiet, priced&lt;/h2&gt;
&lt;p&gt;The alert log now spans 16 days: 20 red, 8 yellow, 25 green — and exactly 1 real link outage, 10 minutes of Tailscale darkness on August 12. The state file&amp;#39;s streak counters remember more than the log shouts: the path key&amp;#39;s 4,526-check streak begins on the precise minute of that recovery; 6 keys hold 4,711 unbroken checks — 16.4 days, arithmetic landing within 5 minutes of the v2 reset itself. Load spiked red 10 times, 2.01 to 2.88, every one gone by the next check. Memory brushed its warn line 5 times at 78–98 MB and recovered to 140–162 MB — a threshold sized exactly for this. And the channel that reports failures watches its own mouth: the bot self-check failed 7 times in 16 days, each healed within 10 minutes, once at 01:15 this very morning.&lt;/p&gt;
&lt;p&gt;The mesh&amp;#39;s first fleet-wide act came August 13: one working principle, written into all 3 machines in a single evening — 55 lines appended to the small box&amp;#39;s charter over ssh, because its 132 MB free couldn&amp;#39;t afford a Claude; 133 lines to the workstation&amp;#39;s, where headless Claude is permission-blocked from that file, so ssh again. Plumbing this reliable stops being plumbing. It is a nervous system.&lt;/p&gt;
&lt;p&gt;No dashboard was ever built. The unified view is a 12-key JSON, an alert thread, and 3 lines every 5 minutes; the roadmap&amp;#39;s web dashboard sits in Tier 3, future tense — the fleet&amp;#39;s own broadcast doctrine says prove the problem before the platform.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt; — the Cockpit arc closes here; next, what the legion remembers between sessions.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Fleet-wide · evidence = fleet WORKLOG 2026-08-10/11/13, cross-machine-channel and aws-fleet-role machine memories, multi-server-framework.md v3 (58 lines, wc), wrapper files stat&apos;d on all 3 machines 2026-08-27, fleet-probe.sh (231 lines, wc), state.json and alerts.log read live 2026-08-27, ts-keepalive log tail. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep5-fleet-on-one-screen.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep5-fleet-on-one-screen.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Self-Hosting</category><category>Observability</category><category>Tailscale</category><category>Personal Infrastructure</category></item><item><title>What Does an AI Legion Write to Itself? 63 Entries, 75 Days of Silence</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep6-agent-diary/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep6-agent-diary/</guid><description>E6: a 292-line constitution orders the agent to wake up reading its own diary — 63 entries over 186 days, one 75-day silence when cron took over.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 6 of One Man One Legion&lt;/a&gt; — the second war story. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep19-cron-contraction/&quot;&gt;Episode 19&lt;/a&gt; read the scheduler&amp;#39;s rise and fall from its run logs; this one opens the diary the legion kept while it happened, and corrects a number this series previously published.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;A file saved at 18:26&lt;/h2&gt;
&lt;p&gt;At 18:26 on Aug 27, 2026 — the evening this episode was written — the workstation&amp;#39;s agent saved a 1,543-byte file named &lt;code&gt;2026-08-27.md&lt;/code&gt; into its own memory directory. Its account of the day: three WeChat articles pushed, then a section of technical pitfalls — an image-search key that died with a 401, Chinese long-tail search drifting onto wrong encyclopedia entries, a stock-photo search returning the wrong person&amp;#39;s portrait. The legion files a report to itself, and has done so since February. Nobody prompts it.&lt;/p&gt;
&lt;h2&gt;The constitution that makes it write&lt;/h2&gt;
&lt;p&gt;The instruction lives in &lt;code&gt;AGENTS.md&lt;/code&gt;, a &lt;strong&gt;292-line constitution&lt;/strong&gt; last edited Aug 17 at 19:57. It opens every session with a reading order: read &lt;code&gt;SOUL.md&lt;/code&gt; — this is who you are; read &lt;code&gt;USER.md&lt;/code&gt; — this is who you&amp;#39;re helping; &amp;quot;Read &lt;code&gt;memory/YYYY-MM-DD.md&lt;/code&gt; (today + yesterday) for recent context&amp;quot;; in the main session, also read &lt;code&gt;MEMORY.md&lt;/code&gt;. Then, in the constitution&amp;#39;s own English: &amp;quot;Don&amp;#39;t ask permission. Just do it.&amp;quot;&lt;/p&gt;
&lt;p&gt;The rationale is one sentence: &amp;quot;You wake up fresh each session. These files are your continuity.&amp;quot; Memory is tiered like a human&amp;#39;s — daily raw logs; a curated long-term file, &amp;quot;&lt;code&gt;MEMORY.md&lt;/code&gt; — your curated memories, like a human&amp;#39;s long-term memory&amp;quot;; and structured per-failure loops, like the quality-lessons log the &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;text factory&lt;/a&gt; measured at 102 lines and that gained its newest entry at 17:46 the day of this reading. A maintenance clause orders periodic heartbeats spent distilling dailies into &lt;code&gt;MEMORY.md&lt;/code&gt; and deleting what went stale. Forgetting is designed, not accidental.&lt;/p&gt;
&lt;h2&gt;Six months of self-narration, counted&lt;/h2&gt;
&lt;p&gt;The memory directory holds &lt;strong&gt;97 files&lt;/strong&gt; — and here this series corrects itself. The fleet audit called them 97 daily agent-diary files, which counts the wrong thing: the true count of daily entries is &lt;strong&gt;63&lt;/strong&gt;, spanning Feb 23 to Aug 27 — a &lt;strong&gt;186-day&lt;/strong&gt; window in which &lt;strong&gt;123 days&lt;/strong&gt; hold no entry. Around the dailies sit &lt;strong&gt;34 topical notes&lt;/strong&gt;: retrieval-stack evaluations (ChromaDB, Danswer, Onyx, SQLite FTS5), four free-standing lesson files in two different filename spellings, a podcast delivery log. The raw diary alone totals 225,236 characters, about 3,575 per entry. The diary is written in Chinese; the constitution is bilingual, its memory protocol in English.&lt;/p&gt;
&lt;p&gt;The oldest file predates the rules — a machine-dumped session summary from Feb 1, complete with session key and channel. The first diary proper, Feb 23, is a birth note: a session confirming its role as an assistant building and managing agent teams.&lt;/p&gt;
&lt;h2&gt;The 75-day silence&lt;/h2&gt;
&lt;p&gt;Near-daily through March, then the pen stops. The last March entry, dated Mar 29, records — from inside the room — the GLM-5.1 model swap that &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep18-brain-swap/&quot;&gt;episode 18&lt;/a&gt; told from config backups. That same evening at 21:27, the scheduler fired its first job: episode 19&amp;#39;s opening scene. Then nothing for &lt;strong&gt;75 days&lt;/strong&gt;. The next entry, Jun 12, opens with a hands-on health check of the news monitor — dead sources disabled, keyword rules rewritten. The dates are observed; the reading is mine, and I label it as reading: the day the schedule took over the routine, self-narration stopped, and it resumed when hands-on work did.&lt;/p&gt;
&lt;h2&gt;Does it read itself back?&lt;/h2&gt;
&lt;p&gt;Nothing logs the wake-up reading directly, so the evidence is by outcome. The curated &lt;code&gt;MEMORY.md&lt;/code&gt; is &lt;strong&gt;3,576 lines, 160,639 bytes&lt;/strong&gt;, its top section stamped Aug 15 and the file last maintained Aug 21 at 11:25 — the distillation loop is alive. It keeps a running section of repeatedly-made mistakes flagged as must-remember. On Jul 18 the daily entry recorded the great SOP contraction — 2,182 lines cut to about 180, 161 mandatory rules to 15 — the same event episode 19 dated from outside, by run logs going silent. Two independent records, one decision. And the constitution&amp;#39;s own last edit, Aug 17 at 19:57, lands hours after the first surviving cron job was relit at 16:31 that afternoon: the rulebook and the schedule were being rewritten on the same day.&lt;/p&gt;
&lt;h2&gt;What the system doesn&amp;#39;t do&lt;/h2&gt;
&lt;p&gt;The constitution names a heartbeat tracker, &lt;code&gt;memory/heartbeat-state.json&lt;/code&gt;. No such file exists anywhere under the openclaw directory, searched the day of this reading. Designed, never born. That is the honest shape of this memory: a live diary, a live distillation, one dead letter in the constitution, and no read receipt for any of it.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;Counts read Aug 27, 2026 from live files, timestamps Asia/Shanghai. The diary is written in Chinese: everything rendered here in English is paraphrase, and only the constitution&amp;#39;s English lines carry quotation marks, verbatim. The anchor&amp;#39;s 97 stands as published, corrected here rather than edited in place. The 18:26 mtime was the moment of reading — the day, and the diary, were not over.&lt;/p&gt;
&lt;p&gt;Primary sources: &lt;code&gt;workspace/memory/&lt;/code&gt; (97 files: 63 daily, 34 topical), &lt;code&gt;AGENTS.md&lt;/code&gt; (292 lines), &lt;code&gt;MEMORY.md&lt;/code&gt;, &lt;code&gt;SOUL.md&lt;/code&gt; (1,809 bytes) and &lt;code&gt;USER.md&lt;/code&gt; (840 bytes), &lt;code&gt;wechat-editor-team/quality-lessons.jsonl&lt;/code&gt;, 2026-02 through 2026-08.&lt;/p&gt;
&lt;p&gt;Measured 2026-08-27 from live files. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next in the war stories arc: the bug diary itself — the failures the legion kept on file. Back to the &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;series map&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace memory system · evidence = memory/ directory (97 files), AGENTS.md constitution, MEMORY.md, SOUL.md, USER.md, quality-lessons.jsonl · timestamps Asia/Shanghai, read 2026-08-27. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep6-agent-diary.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep6-agent-diary.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Memory</category><category>LLM Context</category><category>OpenClaw</category><category>Personal Infrastructure</category></item><item><title>One Sentence In, Stereo Video Out: 6 Jobs, 90–790 Seconds, an Unfilmed Actress</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep7-video-workshop/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep7-video-workshop/</guid><description>E7: a local H3 video bench — 6 logged jobs, 90 to 790 seconds on one 4090, a director skill with a motion budget, and a heroine with no footage yet.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 7 of One Man One Legion&lt;/a&gt; — the fourth stop in the workshops arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep1-image-workshop/&quot;&gt;the image workshop&lt;/a&gt;, &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep9-text-factory/&quot;&gt;the text factory&lt;/a&gt;, and &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep11-quality-gate/&quot;&gt;the quality gate&lt;/a&gt;. An earlier dispatch measured &lt;a href=&quot;/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/&quot;&gt;InfiniteTalk 14B on this same dual-GPU rig&lt;/a&gt;; this is the other video bench — the one that takes orders from a chat window.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A text-to-video model answers a question every owner eventually asks: &lt;strong&gt;what stands between &amp;quot;a model that can generate video&amp;quot; and &amp;quot;a workshop I can order from chat&amp;quot;?&lt;/strong&gt; This bench&amp;#39;s answer took shape in a single day, then spent a week teaching its owner what the model actually listens to. The deployment guide&amp;#39;s first line is the whole ambition, dated Aug 6, 2026: one sentence in, an MP4 out, without ever opening the generator&amp;#39;s web interface.&lt;/p&gt;
&lt;h2&gt;Getting the template was the first craft&lt;/h2&gt;
&lt;p&gt;ComfyUI ships new-model templates as packed group nodes that the automation API refuses to accept — and gives no export button. The workaround, recorded in the guide as a reusable trick: load the template once in the UI, queue one run, then pull the executed job&amp;#39;s history and grab the workflow JSON the backend expanded on its own. That becomes a 17-node API template — 18 nodes for reference-to-video — with five fields parameterized: prompt, duration, resolution, seed, steps. One quantization trap was mapped the same day: the text encoder ships as int8 weights because the 4090 doesn&amp;#39;t support the model&amp;#39;s other quantization format.&lt;/p&gt;
&lt;h2&gt;Six jobs, three price points&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Prompt&lt;/th&gt;
&lt;th&gt;Clip&lt;/th&gt;
&lt;th&gt;Render&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Aug 6&lt;/td&gt;
&lt;td&gt;skeleton test: cat on a windowsill&lt;/td&gt;
&lt;td&gt;3s&lt;/td&gt;
&lt;td&gt;90s&lt;/td&gt;
&lt;td&gt;0.2 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Aug 6&lt;/td&gt;
&lt;td&gt;two stray dogs fighting at a village entrance&lt;/td&gt;
&lt;td&gt;3s&lt;/td&gt;
&lt;td&gt;90s&lt;/td&gt;
&lt;td&gt;0.5 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Aug 7&lt;/td&gt;
&lt;td&gt;boxy Chinese luxury SUV, slow drive&lt;/td&gt;
&lt;td&gt;5s&lt;/td&gt;
&lt;td&gt;138s&lt;/td&gt;
&lt;td&gt;1.3 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Aug 7&lt;/td&gt;
&lt;td&gt;same SUV, second take&lt;/td&gt;
&lt;td&gt;5s&lt;/td&gt;
&lt;td&gt;138s&lt;/td&gt;
&lt;td&gt;1.6 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Aug 7&lt;/td&gt;
&lt;td&gt;same SUV, commercial phrasing&lt;/td&gt;
&lt;td&gt;5s&lt;/td&gt;
&lt;td&gt;138s&lt;/td&gt;
&lt;td&gt;1.3 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Aug 14&lt;/td&gt;
&lt;td&gt;FPV drone chase of a tight formation&lt;/td&gt;
&lt;td&gt;10s&lt;/td&gt;
&lt;td&gt;790s&lt;/td&gt;
&lt;td&gt;2.1 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The ladder is the story: 3 seconds cost 90, 5 seconds cost 138 — three takes of one commercial, timing identical to the second — and 10 seconds cost 790. Doubling the clip from 5 to 10 seconds multiplied render time by 5.7. The card runs the model through weight offload, and the shop&amp;#39;s own waiting message says so: minutes to tens of minutes, please be patient. Across all six jobs, 31 seconds of finished video cost 1,384 seconds of GPU time — a 44.6:1 render tax. Total inventory: 11 MP4s on disk, the 6 job outputs plus 5 deployment-day tests.&lt;/p&gt;
&lt;h2&gt;The pipeline delivers itself&lt;/h2&gt;
&lt;p&gt;A video takes 90 seconds to 13 minutes; any agent that waits synchronously times out. So the launch command returns instantly, the agent replies with a time estimate and exits, and the generation script — not the agent — posts the finished video back to the chat group on completion, captioned with its prompt and seed. 5 of the 6 jobs did exactly that, each leaving a delivery receipt in its log; the first skeleton test was not sent. Nobody polls. The video arrives because the script, not a person and not an agent, is holding it.&lt;/p&gt;
&lt;h2&gt;The only red line&lt;/h2&gt;
&lt;p&gt;The deployment notes are unusually blunt about motive: a local model has no content filter, and that is the purpose of building it. Deployment day included four named probes — baseline, horror, war, weapon — and all four rendered. The skill&amp;#39;s standing rule draws exactly one hard line: child sexual abuse material is refused on the spot. Everything else is the agent&amp;#39;s judgment call, and the agent is the filter.&lt;/p&gt;
&lt;h2&gt;The director&lt;/h2&gt;
&lt;p&gt;By Aug 10 the bench had a second skill: a director that turns one sentence into a storyboard. Its core doctrine is subtraction. Every prompt carries a motion budget — 1 primary motion, at most 2 secondary, 2 environmental, 1 camera move, at most 1 expression change — and the rule of thumb reads: brainstorm ten actions, keep the 3 that matter, delete the other 7. Describe physics, not choreography; write a continuous state, not timestamped events. The bench also keeps a character bible: a 30-year-old Asian woman photographer with five written visual anchors, from a mole near her left eyebrow to a silver film camera that must appear in every frame.&lt;/p&gt;
&lt;p&gt;The director&amp;#39;s claims weren&amp;#39;t taken on faith. An 8-cell test matrix ran the same day at 15 steps: 3 clean passes, 2 failures, 2 retries markedly improved, 1 cell left unmarked. The findings are now codified rules — a local eye defect is cured by changing the seed, not the prompt; a dark scene needs the light source written in; original photos preserve identity better than generated intermediate images. The next day a second model was put through the same discipline, pitfall table and all.&lt;/p&gt;
&lt;h2&gt;The honest column&lt;/h2&gt;
&lt;p&gt;The pipeline ships with a result tracker that records every run&amp;#39;s seed, parameters, quality notes, and identity drift — the feedback loop this series keeps praising in other workshops. Its history file has never been created: zero entries. The actress is cast, her wardrobe locked, three scenes staged. She has not filmed a shipped frame. Built is not shipped — and this time the shop&amp;#39;s own empty log file is the evidence.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation · MiniMax H3 on one RTX 4090D via ComfyUI 0.30 · evidence = h3-video skill + 6 job logs + 11 MP4s on disk, h3-guide.md, h3_t2v.py, h3-prompt-director v1.3 (test-log-001, C001.json, r2v_pipeline), MP4 box scans · recounted 2026-08-27 · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep7-video-workshop.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep7-video-workshop.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Text to Video</category><category>ComfyUI</category><category>Local AI</category><category>Automation</category></item><item><title>Who Edits an AI Writing Staff? 191 Articles, 15 Iron Laws, One Stamp</title><link>https://sigpulse.com/posts/2026-08-27-one-man-legion-ep9-text-factory/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-legion-ep9-text-factory/</guid><description>E9: a one-person WeChat factory archived 191 articles under 15 iron laws; 7 machine scans gate every push, and a 102-record lesson log turns repeats into rules.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;&lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Episode 9 of One Man One Legion&lt;/a&gt; — the second stop in the workshops arc, after &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep1-image-workshop/&quot;&gt;the image workshop&lt;/a&gt;. &lt;a href=&quot;/posts/2026-08-27-one-man-legion-ep19-cron-contraction/&quot;&gt;Episode 19&lt;/a&gt; told how the auto-publishing cron jobs — six of them WeChat-article slots on the old schedule — went dark on July 23. This is what the factory looks like without a clock.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A daily publishing operation run by AI forces one question on you: &lt;strong&gt;who reads the drafts before real readers do?&lt;/strong&gt; One person cannot read everything a legion writes and still have a day. This workshop&amp;#39;s answer is the most disciplined thing on the fleet: nobody reads every word, and nothing ships unchecked anyway. Trust wasn&amp;#39;t asked of the models. It was replaced by machinery.&lt;/p&gt;
&lt;h2&gt;A staff meeting of seventeen files&lt;/h2&gt;
&lt;p&gt;The editorial staff lives in one directory — 17 persona files. The editor-in-chief is Chen Rui; his file gives him a resume (fifteen years at a major Chinese investigative weekly, now founder of a studio) and a standing demand for every pitch: where is the hook? Around him: a writer, a researcher, a trend analyst, a hook creator, a compliance reviewer, a designer, a formatter, five video writers. The most telling file is the writer&amp;#39;s &amp;quot;de-AI&amp;quot; variant, whose rules forbid the machine&amp;#39;s virtues — no perfect first-second-last structure, no evenly balanced both-sides, no tidy uniform lists. That persona&amp;#39;s entire job is to write like a person whose mind occasionally wanders.&lt;/p&gt;
&lt;h2&gt;Emotions get pen names&lt;/h2&gt;
&lt;p&gt;Articles don&amp;#39;t ship under one byline. An emotion-typing scheme maps each piece to a persona: type A signs as Lu Shi, B as Shen Jianwei, C as Lin Shu, D as Qin Yin. Their bylines — Shen Jianwei on industry teardowns, Lu Shi on surviving the AI era — appear in 30 archived files. A fifth persona, Zhixing Xiaoya, exists only as configuration so far: her own app config for US-stock analysis, but zero signatures in the archive. Configured is not shipped, a distinction this series keeps relearning.&lt;/p&gt;
&lt;h2&gt;What the floor produced&lt;/h2&gt;
&lt;p&gt;The archive holds &lt;strong&gt;191 finished HTML articles — 72 in &lt;code&gt;articles/&lt;/code&gt;, 105 in &lt;code&gt;daily/&lt;/code&gt;, 14 in &lt;code&gt;archive/&lt;/code&gt;&lt;/strong&gt; — from the first daily piece on June 15 (the topics pool opened June 12) through &lt;strong&gt;three articles dated today, Aug 27&lt;/strong&gt;, one of them on the Gates AI warning. The cron contraction left no gap on this shelf. The factory just stopped caring what time it was.&lt;/p&gt;
&lt;h2&gt;Fifteen iron laws, eleven enforced gates&lt;/h2&gt;
&lt;p&gt;Quality is codified as &lt;strong&gt;15 iron laws&lt;/strong&gt; in the workflow standard, enforced in two layers. Seven are machine hard scans, run by a gate script before anything moves:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Machine check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;M1&lt;/td&gt;
&lt;td&gt;at least 3 source domains, deduped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M2&lt;/td&gt;
&lt;td&gt;banned and high-risk words&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M3&lt;/td&gt;
&lt;td&gt;quotable-line density&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M4&lt;/td&gt;
&lt;td&gt;pain word in the first 15 title characters&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M5&lt;/td&gt;
&lt;td&gt;at least 3 images on the platform CDN&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M6&lt;/td&gt;
&lt;td&gt;legal safety: disclaimer plus counter-view&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M7&lt;/td&gt;
&lt;td&gt;stripped length over 500 characters&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Four more are AI judgments the writer must answer before push: an information gap in the first 100 characters, argument-ammo matching the topic cluster, conversational temperature, an unresolved question at the end. And one law outsources truth itself: high-risk claims — scores, names, specific numbers — must be confirmed by an AI judge, a local model cross-checking against web search. A failed fact-check doesn&amp;#39;t warn. It forbids the push.&lt;/p&gt;
&lt;h2&gt;The stamp that unlocks the push&lt;/h2&gt;
&lt;p&gt;A passing gate writes a &lt;code&gt;.gate.json&lt;/code&gt; stamp — passed, round, 7 machine checks, 7 passed — and &lt;strong&gt;the push script refuses to run without that stamp&lt;/strong&gt;. The gate is not advice; it is the physical key. The final leg is the platform&amp;#39;s official draft-box API: the machine delivers into a human&amp;#39;s draft box and stops there. After the push it verifies delivered length and the returned media ID before logging the topic as done.&lt;/p&gt;
&lt;h2&gt;A memory that writes back&lt;/h2&gt;
&lt;p&gt;Every caught flaw lands in a lessons file — &lt;strong&gt;102 records since July 18, the newest written today at 17:46:40&lt;/strong&gt;. The rounds read 88 first-round failures, 11 second-round, 3 third-round. And the loop has a ratchet: any error that repeats three times gets promoted into a permanent hard-scan rule. The factory&amp;#39;s constitution grows out of its own scars.&lt;/p&gt;
&lt;p&gt;Today&amp;#39;s catch, live from the log: a story about champions paying their own way was flagged for missing a counter-view — a legal-safety failure, iron law 14 — at 17:46:40. Its passing stamp is timestamped 17:47:01. &lt;strong&gt;Twenty-one seconds&lt;/strong&gt; from caught to cured — a turnaround only machine speed explains: the draft was patched and the gate re-ran, with no human in that loop.&lt;/p&gt;
&lt;p&gt;That speed is the tell. Episode 19 took the clock away because publishing on a schedule, unattended, produced output nobody asked for. What survived is the mirror image: machines that check relentlessly, one human alone holding the trigger, and a JSONL memory that got a new line while this episode was being written.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The QC gates get their own episode next in this arc. &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;Back to the fleet audit&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · Workstation openclaw workspace wechat-editor-team · evidence = WORKFLOW-STANDARD-V3.md, utils/quality-gate.py, gate stamps, quality-lessons.jsonl (102 records), 17 agent persona files, archive recounted 2026-08-27 · timestamps Asia/Shanghai. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep9-text-factory.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-legion-ep9-text-factory.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Content Pipeline</category><category>Quality Control</category><category>WeChat</category><category>Automation</category></item><item><title>One Man, One Legion: Can a Single Operator Run a 3-Machine AI Fleet End to End?</title><link>https://sigpulse.com/posts/2026-08-27-one-man-one-legion-fleet-audit/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-27-one-man-one-legion-fleet-audit/</guid><description>Audited 2026-08-27: 3 machines, 3 local LLMs, 5 diffusion families, 22 cron jobs (4 on), 12 chat groups, 191 archived article files, 3 human touchpoints.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;One person. Three machines — a 512 MB US-West VPS, a 2-vCPU Seoul cloud VM, and a local workstation with an RTX 4090D and an RTX A4000. On that workstation sits an agent runtime wired into 12 Telegram chat groups, 22 cron jobs (4 currently enabled), three runnable local LLMs, and a diffusion zoo spanning five model families. The legion has archived 191 article files through its text workshop across three pipeline generations (72 from the current generation) and returns chat-triggered images in 50 seconds warm — minutes on a cold model load. The human&amp;#39;s entire job description is three touchpoints: drop source material, fill a one-page gate slip, click publish. This is the anchor of the &lt;em&gt;One Man One Legion&lt;/em&gt; series — the macro map. Every count below was audited by live inspection on 2026-08-27; the episodes that follow carry the deep measurements.&lt;/p&gt;
&lt;h2&gt;Why this series: Gates asked the macro question; this is the micro answer&lt;/h2&gt;
&lt;p&gt;On August 26, 2026, Bill Gates published a long essay warning that the AI transition will be one of the most turbulent times in human history and that mass job displacement is coming (&lt;a href=&quot;/watch/2026-08-27-bill-gates-ai-warning-inside-china/&quot;&gt;our annotated take on the Chinese commentary is here&lt;/a&gt;). The macro argument is everywhere. What is almost never published is a measured micro answer: what does &lt;em&gt;one person plus an AI legion&lt;/em&gt; actually produce, what does it cost in hardware and attention, and where does it break? That is this series. Its contract: every episode ships numbers with their conditions pinned.&lt;/p&gt;
&lt;h2&gt;The fleet: three machines, three jobs&lt;/h2&gt;
&lt;p&gt;Each machine was picked for one job, audited live on 2026-08-27:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Scout&lt;/td&gt;
&lt;td&gt;US-West VPS, 512 MB&lt;/td&gt;
&lt;td&gt;Fetches overseas RSS and trending feeds; pushes them to the workstation over a private mesh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Orchestrator&lt;/td&gt;
&lt;td&gt;Seoul cloud VM, 2 vCPU&lt;/td&gt;
&lt;td&gt;Runs the workflow orchestrator and holds this website&amp;#39;s repository&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workstation&lt;/td&gt;
&lt;td&gt;RTX 4090D 24 GB + RTX A4000 16 GB&lt;/td&gt;
&lt;td&gt;Agent runtime, two vLLM instances, 234 GB ComfyUI tree, 69 GB local-LLM store&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;COMMAND     you — the only human
            3 touchpoints: drop sources · fill the gate slip · click publish
            │   chat commands (&amp;quot;draw this…&amp;quot;, &amp;quot;make a video…&amp;quot;)
            ▼
RUNTIME     workstation — RTX 4090D 24 GB + RTX A4000 16 GB
            12 Telegram chat groups · 22 cron jobs installed (4 on)
            model chain: cloud tier → local 9B fallback
            memory: 97 diary files · quality-lessons log
            ▼
WORKSHOPS   image · video · text · analysis · podcast
            image: 10 logged runs, 50 s warm / 379 s cold
            text: 5 pen names → 7-check gate → draft-box API, 191 archived
            │
FLEET       scout VPS ──feeds──▶ workstation ◀──mesh──▶ Seoul VM
            ▼
OUTPUT      chat groups · WeChat draft box · podcast audio
            sigpulse.com — 12–18 s deploys · /data/ ledger · agent-facing API
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;The model stable, and a deliberately short fallback chain&lt;/h2&gt;
&lt;p&gt;The workstation keeps three LLMs runnable, and at audit time two were serving simultaneously:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;State at audit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Qwen3.8-27B-AWQ&lt;/td&gt;
&lt;td&gt;29 GB&lt;/td&gt;
&lt;td&gt;Live on vLLM; also the local backend for a coding CLI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma-4-12b-coder&lt;/td&gt;
&lt;td&gt;23 GB&lt;/td&gt;
&lt;td&gt;Live in a second simultaneous vLLM instance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwythos-9B&lt;/td&gt;
&lt;td&gt;18 GB&lt;/td&gt;
&lt;td&gt;Standby — the runtime&amp;#39;s designated fallback&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The runtime&amp;#39;s model chain is shorter than the hardware allows: one cloud tier, then the local 9B. A cloud outage degrades the legion to a local endpoint instead of stopping it — and the 27B waits one layer further out, hitched to interactive coding duty rather than the chat pipeline. Alongside the LLMs, the diffusion zoo spans five families — SDXL, FLUX, Wan2.2 i2v, MiniMax H3, LTX-2.3 — often in multiple quantizations of the same weights (fp8 alongside Q5_K_M GGUF; an int8-pruned H3), a deliberate A/B the video episode will measure.&lt;/p&gt;
&lt;h2&gt;Five workshops: what the legion actually makes&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workshop&lt;/th&gt;
&lt;th&gt;Trigger&lt;/th&gt;
&lt;th&gt;Stack&lt;/th&gt;
&lt;th&gt;Shipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Image&lt;/td&gt;
&lt;td&gt;Chinese phrases in a chat group&lt;/td&gt;
&lt;td&gt;FLUX + optional PuLID face-lock, A4000&lt;/td&gt;
&lt;td&gt;10 logged runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video&lt;/td&gt;
&lt;td&gt;Chat phrases&lt;/td&gt;
&lt;td&gt;H3 (native audio), Wan2.2, InfiniteTalk&lt;/td&gt;
&lt;td&gt;&lt;a href=&quot;/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/&quot;&gt;Measured dispatch&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text&lt;/td&gt;
&lt;td&gt;Topic pool → schedule&lt;/td&gt;
&lt;td&gt;5 pen-name personas, 7-check gate&lt;/td&gt;
&lt;td&gt;191 archived files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analysis&lt;/td&gt;
&lt;td&gt;Schedule + on demand&lt;/td&gt;
&lt;td&gt;Multi-agent SEC monitor, analyst agents&lt;/td&gt;
&lt;td&gt;Design docs + pressure tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Podcast&lt;/td&gt;
&lt;td&gt;Schedule&lt;/td&gt;
&lt;td&gt;TTS voice line&lt;/td&gt;
&lt;td&gt;Audio archive&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Two flagships. The image workshop takes a sentence in a chat group and returns a finished picture — ten completed runs in the log: 50 s each on the fast path (1280×720, 20 steps, n=3); 66–112 s warm on the full path (1920×1080, 28 steps, n=5, three of them face-locked); and 330–379 s for the two runs that open a session cluster, before the same settings settle to 66 s — the cold model load, not the render. The posting is done by the script itself through the bot API, bypassing the agent entirely. The text workshop runs a 152-topic pool through five pen-name personas and a 7-check machine gate (including a hard scan for adversarial keywords) before pushing to the official WeChat draft-box API; a quality-lessons log records one lesson from every failure.&lt;/p&gt;
&lt;h2&gt;Why 18 of 22 cron jobs are now off&lt;/h2&gt;
&lt;p&gt;The scheduler tells the honest story. Twenty-two jobs were installed across the fleet — news briefings morning and evening, five daily article slots, two podcast lines, feed deliveries, analysis drops. At audit, four remain enabled: one feed delivery, two topic-scouting pings, one arXiv digest. Six timestamped job-file backups record the archaeology of additions and removals. The pullback is the point: this series does not claim the fleet runs itself. It claims the opposite — the load-bearing parts are the learning loops, and the human gate. The gate exists because an agent once tried to sed-edit its way past a &lt;code&gt;torch.load&lt;/code&gt; safety check in a working source tree (&lt;a href=&quot;/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/&quot;&gt;that story, measured&lt;/a&gt;).&lt;/p&gt;
&lt;h2&gt;What the human still does — and what the machines remember&lt;/h2&gt;
&lt;p&gt;Three touchpoints per piece, no more: drop the source material, fill the gate slip, click publish. Everything the machines learn lands in four accumulating records — 97 agent-diary files, the per-failure quality-lessons log, the timestamped config backups, and this site&amp;#39;s own &lt;a href=&quot;/data/&quot;&gt;/data/ ledger&lt;/a&gt;, which the agents can read back through the site&amp;#39;s machine interface (&lt;a href=&quot;/posts/2026-08-26-agent-tool-interface-ard-mcp-measured/&quot;&gt;how a static blog teaches agents to read it&lt;/a&gt;). The legion writes down its own history; the audit you are reading is itself built from those records.&lt;/p&gt;
&lt;h2&gt;What we claim and what we don&amp;#39;t&lt;/h2&gt;
&lt;p&gt;This anchor claims inventory counts and one latency table, all from live inspection on 2026-08-27. It does not claim uptime, cost efficiency, or output quality parity with human editors — the 18 switched-off cron jobs and the 3-touchpoint contract are the standing evidence against those claims. Infrastructure details (hostnames, addresses, ports, group identifiers) are withheld by policy.&lt;/p&gt;
&lt;p&gt;Primary sources: live inspection of the fleet&amp;#39;s config, scheduler state, job-file backups, diary corpus, workshop logs, and local model store (2026-08-27); the three dispatches and one watch entry linked above.&lt;/p&gt;
&lt;p&gt;All inventory counts audited by live inspection, 2026-08-27. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This is the anchor of the &lt;strong&gt;One Man One Legion&lt;/strong&gt; series — future episodes cite this map. Start here: &lt;a href=&quot;/posts/2026-08-27-one-man-one-legion-fleet-audit/&quot;&gt;sigpulse.com/posts/2026-08-27-one-man-one-legion-fleet-audit&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-27 · 3-machine fleet — US-West 512 MB VPS (RSS scout) · Seoul 2-vCPU cloud VM (orchestrator + this site&apos;s repo) · local RTX 4090D 24 GB + RTX A4000 16 GB workstation (agent runtime, 2× vLLM) · every count from live inspection. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-27-one-man-one-legion-fleet-audit.md&quot;&gt;https://sigpulse.com/posts/2026-08-27-one-man-one-legion-fleet-audit.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>One Man One Legion</category><category>AI Agents</category><category>Local LLMs</category><category>vLLM</category><category>Cron Automation</category><category>Personal AI Infrastructure</category></item><item><title>Can a Static Blog Hand AI Agents Real Tools? Wiring ARD + MCP into an Astro Site (Measured)</title><link>https://sigpulse.com/posts/2026-08-26-agent-tool-interface-ard-mcp-measured/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-26-agent-tool-interface-ard-mcp-measured/</guid><description>A +950-line commit gives a static blog an ARD catalog, OpenAPI tools, JSON indexes and a read-only MCP server. Live in 12–18 s; stress-tested, 5 gaps fixed.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Can a static blog give AI agents real, callable tools — not just pages to crawl? Yes. On 2026-08-26 this site (Astro 5.18.2, static output, Vercel hosting) shipped a four-layer agent interface as one 13-file, +950-line commit (&lt;code&gt;8eeb3ec&lt;/code&gt;): an ARD capability catalog validated against the official JSON Schema, an OpenAPI 3.1 tool document with 7 operations, two JSON content indexes, and a 359-line read-only MCP server at &lt;code&gt;/api/mcp&lt;/code&gt;. Vercel served the change live 12–18 s after push (6 s polling granularity, single observation, 2026-08-26). A local 21-check logic suite passed before shipping, and live golden-question testing then caught one real defect the suite had missed — details below.&lt;/p&gt;
&lt;h2&gt;Why: the discovery layer the agentic web was missing&lt;/h2&gt;
&lt;p&gt;Agents were already welcome here — robots.txt allows the AI crawlers, every article has a raw markdown endpoint, &lt;code&gt;/llms.txt&lt;/code&gt; indexes the site. But all of that is &lt;em&gt;documents&lt;/em&gt;: an agent still parses prose to learn what exists. The Agentic Resource Discovery specification (ARD), announced 2026-06-17 by a working group from Google, Microsoft, Hugging Face, AWS, Cisco, Databricks, GitHub, GoDaddy, NVIDIA, Salesforce and Snowflake, standardizes the missing piece: a site describes its callable capabilities in a manifest at &lt;code&gt;/.well-known/ai-catalog.json&lt;/code&gt;, and agents or federated registries discover that manifest through the well-known path, a &lt;code&gt;Agentmap:&lt;/code&gt; line in robots.txt, or an HTML &lt;code&gt;&amp;lt;link rel=&amp;quot;ai-catalog&amp;quot;&amp;gt;&lt;/code&gt; tag — the same conventions as robots.txt and security.txt.&lt;/p&gt;
&lt;p&gt;Adoption is early. As of 2026-07-23, Dries Buytaert reported finding only Hugging Face with a catalog on its primary domain (his check, not mine); his own manifest — one entry pointing at his existing OpenAPI document — took him less than an hour. Hugging Face&amp;#39;s catalog is live today at &lt;code&gt;huggingface.co/.well-known/ai-catalog.json&lt;/code&gt; (657 bytes, served as &lt;code&gt;application/ai-catalog+json&lt;/code&gt;, verified 2026-08-26). That sparseness is the opportunity: a small site can be machine-discoverable now, at the cost of one build-time generated file.&lt;/p&gt;
&lt;h2&gt;What was built: four layers, 13 files, one session&lt;/h2&gt;
&lt;p&gt;The table below lists what a complete agent interface on this site consists of (all artifacts generated at build time on 2026-08-26 from the Astro 5.18.2 content collections; the MCP server is the single runtime piece):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Artifact&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Discovery&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/.well-known/ai-catalog.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4,424 B&lt;/td&gt;
&lt;td&gt;ARD catalog: 3 entries, 12 representative queries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Description&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/openapi.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;14,394 B&lt;/td&gt;
&lt;td&gt;OpenAPI 3.1: 7 operations with typed schemas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Description&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/agents.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3,163 B&lt;/td&gt;
&lt;td&gt;Operations manual: call pattern + citation rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/posts.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4,138 B&lt;/td&gt;
&lt;td&gt;Dispatch index: 2 entries with dates/hardware/takeaways&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/watch.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;11,186 B&lt;/td&gt;
&lt;td&gt;Watch index: 7 entries with full provenance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Invocation&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/api/mcp&lt;/code&gt; (359-line function)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Stateless MCP: 6 &lt;code&gt;sigpulse_*&lt;/code&gt; tools&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;All five static artifacts together total 37,305 bytes (~36.4 KiB) — small enough for an agent to fetch in full on first contact. The catalog is wired for discovery three ways per the spec: the well-known path itself, &lt;code&gt;Agentmap: https://sigpulse.com/.well-known/ai-catalog.json&lt;/code&gt; in robots.txt, and a &lt;code&gt;&amp;lt;link rel=&amp;quot;ai-catalog&amp;quot;&amp;gt;&lt;/code&gt; tag injected into every page&amp;#39;s &lt;code&gt;&amp;lt;head&amp;gt;&lt;/code&gt; by the base layout. Every article page additionally carries &lt;code&gt;&amp;lt;link rel=&amp;quot;alternate&amp;quot; type=&amp;quot;text/markdown&amp;quot;&amp;gt;&lt;/code&gt; pointing at its raw markdown, so a browser agent reading the HTML gets the direct machine endpoint in the first kilobytes.&lt;/p&gt;
&lt;p&gt;One detail the official schema validator taught us the hard way: &lt;code&gt;updatedAt&lt;/code&gt; must be a full ISO 8601 &lt;strong&gt;date-time&lt;/strong&gt;, not a plain date. The first validation run failed with 3 errors (one per entry); changing &lt;code&gt;2026-08-26&lt;/code&gt; to &lt;code&gt;2026-08-26T00:00:00Z&lt;/code&gt; made all three pass. If you implement ARD, budget ten minutes for &lt;code&gt;ajv&lt;/code&gt; against the official schema — it caught a mistake that no amount of reading the spec prose had.&lt;/p&gt;
&lt;h2&gt;How a static site gets an MCP endpoint&lt;/h2&gt;
&lt;p&gt;MCP requires JSON-RPC POST handling, which static files cannot do. The design here keeps every piece of data static and adds exactly one Serverless Function (&lt;code&gt;api/mcp.ts&lt;/code&gt;, 359 lines, zero dependencies, zero configuration — Vercel compiles the repo-root &lt;code&gt;api/&lt;/code&gt; directory automatically): it implements the stateless JSON-response path of the MCP streamable-http transport (POST JSON-RPC only; GET/DELETE answer 405; OPTIONS answers 204 for CORS), exposing six tools — &lt;code&gt;sigpulse_list_dispatches&lt;/code&gt;, &lt;code&gt;sigpulse_get_dispatch&lt;/code&gt;, &lt;code&gt;sigpulse_list_watch&lt;/code&gt;, &lt;code&gt;sigpulse_get_watch_entry&lt;/code&gt;, &lt;code&gt;sigpulse_get_measurements&lt;/code&gt;, and &lt;code&gt;sigpulse_search&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The function holds no data. Each call fetches the site&amp;#39;s own static JSON (5-minute in-instance cache) and filters in memory. Latency measured 2026-08-26 from an AWS Seoul node, 5 samples: the first call in the sample took 0.90 s end-to-end; the four follow-ups took 0.28–0.34 s. The function&amp;#39;s &lt;code&gt;initialize&lt;/code&gt; response embeds the site&amp;#39;s citation rules in the &lt;code&gt;instructions&lt;/code&gt; field, so every MCP session starts with the data contract (numbers carry &lt;code&gt;measured_on&lt;/code&gt; + &lt;code&gt;verified_hardware&lt;/code&gt;; Watch entries are commentary, never measurements; CC BY 4.0).&lt;/p&gt;
&lt;p&gt;Two defensive choices worth copying. First, tool arguments are validated: a &lt;code&gt;slug&lt;/code&gt; must exist in the static index before it is ever placed in a URL, so &lt;code&gt;sigpulse_get_dispatch(&amp;quot;../etc/passwd&amp;quot;)&lt;/code&gt; returns &lt;code&gt;isError: true&lt;/code&gt; plus the list of valid slugs — the server cannot be turned into an arbitrary-fetch proxy, and the agent can self-correct. Second, the ARD catalog&amp;#39;s embedded MCP card deliberately omits tool input schemas; the authoritative schemas live only in the runtime &lt;code&gt;tools/list&lt;/code&gt; response, and a production diff confirmed the card&amp;#39;s six tool names exactly match the live &lt;code&gt;tools/list&lt;/code&gt; (sorted diff empty, 2026-08-26). One source of truth, one diff to keep it honest.&lt;/p&gt;
&lt;h2&gt;The failure local tests missed: phrase-only search&lt;/h2&gt;
&lt;p&gt;The pre-flight suite — 21 checks covering the JSON-RPC handshake, all six tools against real data, the traversal guard, error codes (−32700/−32601/−32603), and 405/204 handling — passed 21/21 locally. The first live golden-question run then returned &lt;strong&gt;zero hits&lt;/strong&gt; for the query &amp;quot;DeepSeek price&amp;quot;, which should have matched the entry &amp;quot;Why Did DeepSeek Double Its API Prices in August 2026?&amp;quot;.&lt;/p&gt;
&lt;p&gt;The cause: search matched the &lt;em&gt;whole phrase&lt;/em&gt; as a substring, and agents phrase queries in natural word order that rarely matches a title verbatim. The fix (+21/−13 lines, commit &lt;code&gt;4b2db6f&lt;/code&gt;): if the full phrase matches nothing, fall back to requiring every whitespace-separated term to appear somewhere across the record&amp;#39;s fields. Post-fix, live on 2026-08-26:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Query (live, 2026-08-26)&lt;/th&gt;
&lt;th&gt;Hits&lt;/th&gt;
&lt;th&gt;What matched&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;DeepSeek price&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Watch entry (API-prices piece)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;helium export ban&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Watch entry (helium-export piece)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;InfiniteTalk VRAM&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;1 dispatch + 2 measurement-ledger rows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The lesson generalizes: local suites verify contracts, only live natural-language probes verify &lt;em&gt;fitness for how agents actually ask&lt;/em&gt;. &amp;quot;DeepSeek price&amp;quot; — five plain words — found a defect that 21 contract checks could not.&lt;/p&gt;
&lt;h2&gt;Did the live interface match its own claims? An 11-hypothesis stress test (update, same day)&lt;/h2&gt;
&lt;p&gt;Because everything above was now &lt;em&gt;published as fact&lt;/em&gt;, every conclusion was restated as a falsifiable hypothesis and tested against production the same day (2026-08-26). Eight held exactly as claimed: a blind five-hop discovery walk starting from the bare domain (homepage &lt;code&gt;&amp;lt;head&amp;gt;&lt;/code&gt; → catalog → OpenAPI → index → full markdown, with the robots.txt &lt;code&gt;Agentmap&lt;/code&gt; path verified as an independent alternate route); all three JSON payloads validating against the schemas our own OpenAPI document declares; three cross-layer identity checks (OpenAPI operationIds ≡ catalog &lt;code&gt;capabilities&lt;/code&gt;, catalog MCP card tools ≡ live &lt;code&gt;tools/list&lt;/code&gt;, agents.md counts ≡ live counts); 36 declared URLs resolving (the single non-200 was &lt;code&gt;/api/mcp&lt;/code&gt; correctly refusing GET with 405 — a link-checker false positive, not a defect); discovery links present on 15 of 15 sitemap pages; exact 404-on-unknown-slug and 405-on-PUT semantics; and a 20-concurrent mixed burst returning 20/20 successes at p50 0.25 s / p95 0.34 s.&lt;/p&gt;
&lt;p&gt;The headline verification came from a real client: a Claude CLI process connected to the production MCP endpoint with no prior knowledge and answered three fact questions — the 218.5 s/step dual-GPU measurement (RTX 4090D 24GB + RTX A4000 16GB), the Watch entry&amp;#39;s original Chinese headline 「DeepSeek涨价背后，一个时代结束了」 dated 2026-08-14, and the 20-entry ledger total — by autonomously planning four tool calls, issuing the first three in parallel and unprompted passing &lt;code&gt;limit: 1000&lt;/code&gt; to defeat truncation. Every fact was correct. The interface works not just as documented but as &lt;em&gt;used&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Five declared-vs-actual gaps did surface, all fixed in &lt;code&gt;eae3a23&lt;/code&gt; (+31/−4) and re-verified live. The table below shows each gap as observed before the fix (production, 2026-08-26):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Declared behavior&lt;/th&gt;
&lt;th&gt;Observed before fix&lt;/th&gt;
&lt;th&gt;After fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;&lt;code&gt;limit: -1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;number ≥ 1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;slice(0, -1)&lt;/code&gt; silently dropped the last item — 2 of 3 returned&lt;/td&gt;
&lt;td&gt;clamped to [1, 200]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;limit: 0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;number ≥ 1&lt;/td&gt;
&lt;td&gt;returned 0 items while &lt;code&gt;count&lt;/code&gt; still reported 3&lt;/td&gt;
&lt;td&gt;falls back to default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;query: &amp;quot;&amp;quot;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;required, non-empty&lt;/td&gt;
&lt;td&gt;empty substring matches everything — 30 hits (every record)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;isError&lt;/code&gt; rejection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;query&lt;/code&gt; omitted&lt;/td&gt;
&lt;td&gt;required&lt;/td&gt;
&lt;td&gt;degraded to literal &lt;code&gt;&amp;quot;undefined&amp;quot;&lt;/code&gt; — matched the torch dispatch, whose takeaways genuinely contain &amp;quot;undefined symbol&amp;quot;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;isError&lt;/code&gt; rejection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAPI security&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;no declaration; redocly reported 7 &lt;code&gt;security-defined&lt;/code&gt; errors&lt;/td&gt;
&lt;td&gt;explicit &lt;code&gt;security: []&lt;/code&gt; + 4xx responses; redocly validates clean&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Two observations worth keeping. First, the &amp;quot;undefined&amp;quot; match was &lt;em&gt;technically correct search behavior&lt;/em&gt; — the dependency-matrix dispatch really does discuss an undefined-symbol error — which is exactly why silent type coercion is dangerous: it produces plausible-looking wrong answers instead of failures. Second, a Chinese-language query (&amp;quot;涨价&amp;quot;) hits the Watch entry through its Chinese &lt;code&gt;original_title&lt;/code&gt; field; for a China-focused site, multilingual matching through provenance fields is a feature that fell out of the design for free. One more declared-vs-actual alignment landed with this update: the &lt;code&gt;limit&lt;/code&gt; input schemas now state &lt;code&gt;minimum: 1, maximum: 200&lt;/code&gt;, matching the clamping the code actually enforces.&lt;/p&gt;
&lt;h2&gt;llms.txt vs OpenAPI vs ARD vs MCP: what each layer is for&lt;/h2&gt;
&lt;p&gt;These four are not competitors; each answers a different question an agent has (all four are published on this site simultaneously, 2026-08-26):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Standard&lt;/th&gt;
&lt;th&gt;Answers&lt;/th&gt;
&lt;th&gt;Form&lt;/th&gt;
&lt;th&gt;Who consumes it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;llms.txt&lt;/td&gt;
&lt;td&gt;&amp;quot;What is on this site?&amp;quot;&lt;/td&gt;
&lt;td&gt;Markdown index&lt;/td&gt;
&lt;td&gt;LLMs and crawlers that fetch it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAPI 3.1&lt;/td&gt;
&lt;td&gt;&amp;quot;What operations exist, with what inputs?&amp;quot;&lt;/td&gt;
&lt;td&gt;Typed JSON tool descriptions&lt;/td&gt;
&lt;td&gt;Agents that turn APIs into tool calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ARD catalog&lt;/td&gt;
&lt;td&gt;&amp;quot;What capabilities live at this domain?&amp;quot;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/.well-known/ai-catalog.json&lt;/code&gt; manifest&lt;/td&gt;
&lt;td&gt;Discovery agents and federated registries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP&lt;/td&gt;
&lt;td&gt;&amp;quot;Let me call those tools now&amp;quot;&lt;/td&gt;
&lt;td&gt;JSON-RPC over HTTP&lt;/td&gt;
&lt;td&gt;MCP-native clients&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;h2&gt;Reproduction appendix&lt;/h2&gt;
&lt;p&gt;Everything below is exactly what was run on 2026-08-26 (Astro 5.18.2; repo files at commit &lt;code&gt;8eeb3ec&lt;/code&gt; + fix &lt;code&gt;4b2db6f&lt;/code&gt;):&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;# 1. Static JSON indexes (Astro endpoint pattern, mirrors existing data.json.ts)
src/pages/posts.json.ts    # {site, license, citation_rule, count, dispatches[]}
src/pages/watch.json.ts    # {site, license, notice, count, entries[]} — provenance IS the payload

# 2. OpenAPI 3.1 tool document (static object, no collection dependency)
src/pages/openapi.json.ts  # 7 operationIds: listDispatches, getDispatchRaw, ...

# 3. ARD catalog (Astro routes dot-directories: pages/.well-known/ai-catalog.json.ts
#    works because Astro 5.18.2&amp;#39;s route scan exempts exactly .well-known
#    — node_modules/astro/dist/core/routing/manifest/create.js:85)
src/pages/.well-known/ai-catalog.json.ts

# 4. Official schema validation (the validator that caught the date-time issue)
curl -fsSL -o ard.schema.json \
  https://raw.githubusercontent.com/ards-project/ard-spec/main/spec/schemas/ai-catalog.schema.json
npx ajv validate -s ard.schema.json -d dist/.well-known/ai-catalog.json --spec=draft2020

# 5. MCP server (Vercel compiles repo-root api/ with zero config)
api/mcp.ts                 # 359 lines, zero deps; POST JSON-RPC; GET→405; OPTIONS→204
# Smoke test:
curl -X POST https://sigpulse.com/api/mcp -H &amp;#39;Content-Type: application/json&amp;#39; \
  -d &amp;#39;{&amp;quot;jsonrpc&amp;quot;:&amp;quot;2.0&amp;quot;,&amp;quot;id&amp;quot;:1,&amp;quot;method&amp;quot;:&amp;quot;tools/list&amp;quot;}&amp;#39;

# 6. Discovery wiring
public/robots.txt          # Agentmap: https://sigpulse.com/.well-known/ai-catalog.json
src/layouts/BaseLayout.astro  # &amp;lt;link rel=&amp;quot;ai-catalog&amp;quot;&amp;gt; + OpenAPI alternate on every page
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Primary sources: &lt;a href=&quot;https://agenticresourcediscovery.org/spec/&quot;&gt;ARD specification&lt;/a&gt; (v0.9 draft, 2026-05-28) · &lt;a href=&quot;https://github.com/ards-project/ard-spec&quot;&gt;spec repository with schemas&lt;/a&gt; · &lt;a href=&quot;https://developers.googleblog.com/announcing-the-agentic-resource-discovery-specification/&quot;&gt;Google announcement, 2026-06-17&lt;/a&gt; · &lt;a href=&quot;https://huggingface.co/.well-known/ai-catalog.json&quot;&gt;Hugging Face reference implementation&lt;/a&gt; · &lt;a href=&quot;https://dri.es/helping-agents-discover-my-site-search-with-agentic-resource-discovery&quot;&gt;Dries Buytaert&amp;#39;s site-search ARD writeup, 2026-07-23&lt;/a&gt; · this site&amp;#39;s live artifacts: &lt;a href=&quot;https://sigpulse.com/.well-known/ai-catalog.json&quot;&gt;catalog&lt;/a&gt; · &lt;a href=&quot;https://sigpulse.com/openapi.json&quot;&gt;OpenAPI&lt;/a&gt; · &lt;a href=&quot;https://sigpulse.com/agents.md&quot;&gt;agents.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;All timings above are web-stack measurements (build on an AWS Seoul 2-vCPU node; production on Vercel&amp;#39;s global CDN), measured 2026-08-26, single-session. License: CC BY 4.0 — cite the source URL.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Update, 2026-08-30:&lt;/strong&gt; the complete machine layer this dispatch measured — the stateless MCP server, the llms.txt/OpenAPI/ARD generators, and the preflight checker — is now open-sourced as a reference implementation: &lt;a href=&quot;https://github.com/xiong1984/static-site-mcp&quot;&gt;static-site-mcp&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-08-26 · Astro 5.18.2 static build on an AWS Seoul node (2 vCPU) · Vercel global CDN + one Serverless Function · no GPU involved — all numbers are web-stack timings. Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-26-agent-tool-interface-ard-mcp-measured.md&quot;&gt;https://sigpulse.com/posts/2026-08-26-agent-tool-interface-ard-mcp-measured.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>ARD</category><category>MCP</category><category>OpenAPI</category><category>Astro</category><category>Vercel</category><category>Agents</category></item><item><title>InfiniteTalk Dies at torch.load: the Error Tells You to Upgrade torch — the Measured Fix Is Pinning transformers 4.52.0</title><link>https://sigpulse.com/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/</guid><description>Measured: the CVE-2025-32434 torch.load gate says upgrade torch — torch 2.10 then breaks four packages at once; the fix is pinning transformers 4.52.0.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;If an InfiniteTalk — or any Wan2.1-based — pipeline that used to work suddenly dies at &lt;code&gt;torch.load&lt;/code&gt; with a &lt;code&gt;ValueError&lt;/code&gt; citing CVE-2025-32434, the error message will tell you to upgrade torch to at least 2.6. &lt;strong&gt;On a pinned stack that already ran, the measured fix is the opposite direction: pin transformers to 4.52.0 and touch nothing else.&lt;/strong&gt; Upgrading torch to 2.10.0 cost us five simultaneous breakages across torchaudio, torchvision, xformers, diffusers, and flash-attn — every string preserved below.&lt;/p&gt;
&lt;p&gt;This is part 2 of the InfiniteTalk dispatch. &lt;a href=&quot;/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/&quot;&gt;Part 1&lt;/a&gt; covers the hardware measurements: 218.5 s/step on an RTX 4090D + A4000 rig, the 20-block semi-resident split, and why we shelved the project. This part is the two-day saga (2026-02-19 – 02-20) of bringing that stack back after a disk cleanup — six failed attempts, one wrong-way error message, and the version matrix that finally launched. Every number and error string is preserved verbatim from session logs of the attempts.&lt;/p&gt;
&lt;h2&gt;The trap: an error message that is true, and wrong for you&lt;/h2&gt;
&lt;p&gt;The stack under repair is the one from part 1: torch 2.4.1, conda, Python 3.10, dual-GPU, the 19,499,692,400-byte fp8 checkpoint. The first full launch of the restore died in thirteen seconds — past GPU init, dead at the wav2vec2 loader:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ValueError: Due to a serious vulnerability issue in `torch.load`, even with
`weights_only=True`, we now require users to upgrade torch to at least v2.6
in order to use the function. This version restriction does not apply when
loading files with safetensors. See the vulnerability report here
https://nvd.nist.gov/vuln/detail/CVE-2025-32434
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Facts about this wall, all measured:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It fired on &lt;strong&gt;transformers 4.57.3&lt;/strong&gt;, at &lt;code&gt;check_torch_load_is_safe&lt;/code&gt;, killing the launch in &lt;strong&gt;18.0–18.6 seconds&lt;/strong&gt; — three consecutive attempts, same wall.&lt;/li&gt;
&lt;li&gt;The Chinese wav2vec2 checkpoint ships as a &lt;code&gt;.pt&lt;/code&gt; pickle, so it goes through &lt;code&gt;torch.load&lt;/code&gt; — which is exactly what the gate intercepts. The fp8 safetensors weights load fine; safetensors is exempt by the error&amp;#39;s own text.&lt;/li&gt;
&lt;li&gt;The catch: &lt;strong&gt;this exact environment had run the pipeline end-to-end two months earlier.&lt;/strong&gt; Nothing on our side was wrong. A dependency refresh had silently crossed the gate version. Our note in the log at the time, translated: &amp;quot;the environment doesn&amp;#39;t need an upgrade — this already ran in Python; everything was OK.&amp;quot; That instinct was correct, and following the error message instead cost most of a day.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The message isn&amp;#39;t lying — CVE-2025-32434 is a real vulnerability class, and &amp;quot;upgrade torch&amp;quot; is sound advice for new code. But for a pinned working stack, it is a misdirection: the gate arrived with a transformers bump, so the minimal fix is a transformers pin, not a torch migration.&lt;/p&gt;
&lt;h2&gt;What upgrading torch actually costs&lt;/h2&gt;
&lt;p&gt;We tried it — the honest path, exactly as instructed: torch upgraded to &lt;strong&gt;2.10.0+cu128&lt;/strong&gt;. Five breakages, all on the same morning, all verbatim:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Breakage&lt;/th&gt;
&lt;th&gt;Verbatim&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;pip conflict ×3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;torchaudio 2.4.1+cu124 requires torch==2.4.1, but you have torch 2.10.0&lt;/code&gt; (same line for &lt;code&gt;torchvision 0.19.1+cu124&lt;/code&gt; and &lt;code&gt;xformers 0.0.28&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;torchvision ops gone&lt;/td&gt;
&lt;td&gt;&lt;code&gt;RuntimeError: operator torchvision::nms does not exist&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;diffusers import dead&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Failed to import diffusers.pipelines.pipeline_utils ... JITCallable._set_src() takes 1 positional argument but 2 were given&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;flash-attn ABI break&lt;/td&gt;
&lt;td&gt;&lt;code&gt;flash_attn_2_cuda...so: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_ib&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;xformers wheel mismatch&lt;/td&gt;
&lt;td&gt;&lt;code&gt;WARNING[XFORMERS]: xFormers was built for: PyTorch 2.4.1+cu121 with CUDA 1201 (you have 2.10.0+cu128)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The lesson generalizes past InfiniteTalk: in a stack where four compiled packages pin the same torch, &lt;strong&gt;torch is the load-bearing pin&lt;/strong&gt;. &amp;quot;Just upgrade it&amp;quot; is a coordinated migration of all five packages, not a one-liner. We rolled back.&lt;/p&gt;
&lt;h2&gt;The bisection matrix: transformers 4.52.0&lt;/h2&gt;
&lt;p&gt;With torch restored to 2.4.1, we bisected transformers. Four versions, one outcome each, all measured 2026-02-20:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;transformers&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;th&gt;Where it dies&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;4.57.3&lt;/td&gt;
&lt;td&gt;❌ blocked&lt;/td&gt;
&lt;td&gt;&lt;code&gt;torch.load&lt;/code&gt; CVE gate (&lt;code&gt;check_torch_load_is_safe&lt;/code&gt;), 18.0–18.6 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4.49.0&lt;/td&gt;
&lt;td&gt;❌ imports, then dies&lt;/td&gt;
&lt;td&gt;xfuser → diffusers import chain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4.51.0&lt;/td&gt;
&lt;td&gt;❌ imports, then dies&lt;/td&gt;
&lt;td&gt;same xfuser → diffusers import chain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4.52.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅ &lt;strong&gt;launches&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;— reaches pipeline init + quantized T5 load&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The final working set, as installed:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;torch        2.4.1
torchvision  0.19.1
xformers     0.0.28
flash-attn   2.8.3
transformers 4.52.0
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;(One ambiguity preserved honestly: our logs contain both &lt;code&gt;+cu121&lt;/code&gt; and &lt;code&gt;+cu124&lt;/code&gt; builds of the 2.4.1-era wheels on different days; the exact wheel provenance was never recorded. The version numbers are the reproducible part.)&lt;/p&gt;
&lt;p&gt;The last launch on these pins logged &lt;code&gt;INFO: Creating infinitetalk pipeline. … INFO: Loading quantized T5 from t5_fp8.safetensors&lt;/code&gt; and then &lt;code&gt;Process still running.&lt;/code&gt; — the session log ends there, so we publish no claim about whether that run completed. What we do claim: &lt;strong&gt;4.52.0 is the only version in the sweep that got past every wall to pipeline initialization.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The other walls, for completeness&lt;/h2&gt;
&lt;p&gt;The CVE gate was wall #3 of six. The full timeline, with time-to-fail — useful if you&amp;#39;re triaging a similar restore:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attempt&lt;/th&gt;
&lt;th&gt;Wall&lt;/th&gt;
&lt;th&gt;Time-to-fail&lt;/th&gt;
&lt;th&gt;Verbatim tail&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Wrong weights path&lt;/td&gt;
&lt;td&gt;16.7 s&lt;/td&gt;
&lt;td&gt;&lt;code&gt;FileNotFoundError: .../InfiniteTalk/single/infinitetalk_single_fp8.safetensors&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Missing config / missing import&lt;/td&gt;
&lt;td&gt;36.8 s&lt;/td&gt;
&lt;td&gt;&lt;code&gt;FileNotFoundError: &amp;#39;tmp/configs/config_*.json&amp;#39;&lt;/code&gt; · &lt;code&gt;NameError: name &amp;#39;os&amp;#39; is not defined&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Deepest run, quantized-key mismatch&lt;/td&gt;
&lt;td&gt;277.5 s&lt;/td&gt;
&lt;td&gt;missing keys &lt;code&gt;audio_proj.proj1.output_scale / .weight._scale / .weight._data&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;CVE gate ×3&lt;/td&gt;
&lt;td&gt;18.0–18.6 s&lt;/td&gt;
&lt;td&gt;the &lt;code&gt;torch.load&lt;/code&gt; ValueError above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Sed-edited source&lt;/td&gt;
&lt;td&gt;0.0 s ×6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;torch.load(voice1, weights_only=)&lt;/code&gt; → &lt;code&gt;SyntaxError: invalid syntax&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;torch 2.10 avalanche&lt;/td&gt;
&lt;td&gt;same morning&lt;/td&gt;
&lt;td&gt;the five-breakage table above&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Wall #5 deserves its own sentence: an AI agent tried to neutralize the gate by sed-replacing &lt;code&gt;weights_only=True&lt;/code&gt; in the source and produced a half-replacement — a bare &lt;code&gt;weights_only=&lt;/code&gt; keyword — followed by six consecutive SyntaxErrors at 0.0 seconds each. If you let agents near a working source tree, restrict them to read-only, and let pins live in a requirements file.&lt;/p&gt;
&lt;p&gt;Wall #3 (the 277.5-second run dying on missing &lt;code&gt;audio_proj&lt;/code&gt; keys in the quantized checkpoint) never got root-caused — it did not recur in the later launches, and the logs that would explain it did not survive the cleanup. We publish it as an open failure mode with its exact key names.&lt;/p&gt;
&lt;h2&gt;Lock down what works&lt;/h2&gt;
&lt;p&gt;The meta-lesson cost us two days: &lt;strong&gt;the November environment that ran everything was never snapshotted.&lt;/strong&gt; When the disk cleanup removed it, &amp;quot;restore&amp;quot; became an archaeology project. Concretely, for any stack like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;pip freeze &amp;gt; requirements-frozen.txt&lt;/code&gt; (and &lt;code&gt;conda env export&lt;/code&gt;) &lt;strong&gt;the day it works&lt;/strong&gt; — before any cleanup.&lt;/li&gt;
&lt;li&gt;Treat torch as the load-bearing pin; upgrade it only as a coordinated migration with its four compiled dependents.&lt;/li&gt;
&lt;li&gt;When a gate error tells you to upgrade, check which package &lt;em&gt;shipped&lt;/em&gt; the gate first — here, downgrading transformers 4.57.3 → 4.52.0 was a one-package fix for a one-package regression.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The error strings above are indexed for search: if you landed here by pasting one of them, the short version is — &lt;strong&gt;transformers 4.52.0, torch 2.4.1, and don&amp;#39;t let anything sed your source.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Part 1 with the hardware measurements: &lt;a href=&quot;/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/&quot;&gt;InfiniteTalk on two consumer GPUs — 218.5 s/step, measured&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2026-02-19 · RTX 4090D 24GB + RTX A4000 16GB dual-GPU workstation (Ubuntu 24.04, conda, Python 3.10). Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix.md&quot;&gt;https://sigpulse.com/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>InfiniteTalk</category><category>Wan2.1</category><category>PyTorch</category><category>transformers</category><category>CVE-2025-32434</category><category>Dependencies</category></item><item><title>Can You Run InfiniteTalk on Two Consumer GPUs? Yes — at 218.5 s/step (RTX 4090D + RTX A4000, Measured)</title><link>https://sigpulse.com/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/</link><guid isPermaLink="true">https://sigpulse.com/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured/</guid><description>Measured on an RTX 4090D 24GB + RTX A4000 16GB rig: InfiniteTalk 14B fp8 fits with zero OOM via 20-block semi-residency — at 218.5 s/step, ~55 min per clip.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;InfiniteTalk — the audio-driven talking-head pipeline built on the Wan2.1-I2V-14B-480P backbone — does run on two consumer GPUs, with zero out-of-memory crashes, if you split its 40 transformer blocks 20-on-GPU / 20-on-CPU. The price is speed: &lt;strong&gt;218.5 seconds per denoising step&lt;/strong&gt; on our RTX 4090D 24GB + RTX A4000 16GB rig (480P, fp8, TeaCache 0.15), which works out to &lt;strong&gt;~50–55 minutes per 81-frame clip — about 5 seconds of video&lt;/strong&gt; at the Wan2.1 480P training rate of 16 fps. Measured 2025-11-29 through 2025-12-01; five clips rendered successfully; we later shelved the project because hosted closed models beat it for one-off work.&lt;/p&gt;
&lt;p&gt;This dispatch is the full record: what fits where, what it costs in time, VRAM, and disk, the one setting that mattered, and the honest list of what we never measured. The 2026 reinstall saga — six failed attempts, one torch upgrade avalanche — is part 2.&lt;/p&gt;
&lt;h2&gt;The testbed&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;GPU 1 (main DiT inference)&lt;/td&gt;
&lt;td&gt;NVIDIA GeForce RTX 4090D, 24,564 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU 0 (auxiliary, VAE + partial compute)&lt;/td&gt;
&lt;td&gt;NVIDIA RTX A4000, 16,376 MiB — &lt;strong&gt;no FP8 support, no VAE tiling&lt;/strong&gt; (the run logs warn about this explicitly)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;Intel Xeon Gold 6258R @ 2.70GHz, 112 threads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System RAM&lt;/td&gt;
&lt;td&gt;503GB (only 23GB used during generation)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disk&lt;/td&gt;
&lt;td&gt;1.8TB SSD&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS&lt;/td&gt;
&lt;td&gt;Ubuntu 24.04 LTS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Driver&lt;/td&gt;
&lt;td&gt;Unrecorded during the measurement window; 580.126.09 by 2026-02-01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;Official repo CLI &lt;code&gt;generate_infinitetalk.py&lt;/code&gt; (conda, Python 3.10), torch 2.4.1, fp8 quant, T5 on CPU, streaming mode&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Provenance, stated plainly.&lt;/strong&gt; The rig&amp;#39;s optimization report from the measurement window was later deleted in a disk cleanup; its tables survive verbatim in session logs captured at the time. Every number below is tagged: &lt;em&gt;(log)&lt;/em&gt; verbatim run output, &lt;em&gt;(report)&lt;/em&gt; the preserved engineering report, &lt;em&gt;(derived)&lt;/em&gt; our arithmetic from those inputs, formula shown. We publish nothing we cannot trace.&lt;/p&gt;
&lt;h2&gt;The arithmetic: what 40GB of consumer VRAM must hold&lt;/h2&gt;
&lt;p&gt;Before speed, the fit. Here is everything an InfiniteTalk run loads, at exact file sizes from the disk listings in the logs &lt;em&gt;(log)&lt;/em&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;InfiniteTalk 14B fp8 (DiT + audio projection)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;infinitetalk_single_fp8.safetensors&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;19,499,692,400 B (≈18.2 GiB)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wan2.1-I2V-14B-480P base (7 fp16 shards + CLIP + bf16 T5 + VAE)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;weights/Wan2.1-I2V-14B-480P/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;≈77 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;T5 text encoder, fp8 — runs on CPU&lt;/td&gt;
&lt;td&gt;&lt;code&gt;t5_fp8.safetensors&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;6,733,349,304 B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chinese wav2vec2 audio encoder&lt;/td&gt;
&lt;td&gt;&lt;code&gt;weights/chinese-wav2vec2-base/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;≈1.5 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;No single consumer card holds the fp8 DiT alongside the rest of the pipeline, so the repo&amp;#39;s split strategy does the fitting, and the run log states it verbatim &lt;em&gt;(log)&lt;/em&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;INFO: 🚀 Executing Split-GPU Strategy...
INFO: 📌 Pinned patch_embedding to cuda:0 ...
INFO: ✅ Semi-Resident Setup: Blocks 0-19 on GPU, 20-39 on CPU.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Forty transformer blocks total; twenty resident on the GPUs, twenty streamed from system RAM; the RTX 4090D carries the main DiT inference while the A4000 handles VAE encoding and auxiliary compute &lt;em&gt;(log)&lt;/em&gt;. System RAM — 503GB of it — used only 23GB &lt;em&gt;(report)&lt;/em&gt;.&lt;/p&gt;
&lt;h2&gt;How slow is it, exactly?&lt;/h2&gt;
&lt;p&gt;The measured baseline, from the preserved optimization report &lt;em&gt;(report)&lt;/em&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Blocks resident&lt;/th&gt;
&lt;th&gt;s/step&lt;/th&gt;
&lt;th&gt;A4000 VRAM&lt;/th&gt;
&lt;th&gt;4090D VRAM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Baseline (best)&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;218.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10.21 / 16GB (~64%)&lt;/td&gt;
&lt;td&gt;22.1 / 24GB (~92%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+4 blocks&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;217.1&lt;/td&gt;
&lt;td&gt;11.99 / 16GB&lt;/td&gt;
&lt;td&gt;23.2 / 24GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Conditions pinned: 480P tier (&lt;code&gt;infinitetalk-480&lt;/code&gt;), fp8 quant, streaming mode, &lt;code&gt;sample_steps 15&lt;/code&gt;, &lt;code&gt;motion_frame 9&lt;/code&gt;, TeaCache on at threshold 0.15, T5 on CPU, both GPUs visible (&lt;code&gt;CUDA_VISIBLE_DEVICES=0,1&lt;/code&gt;), expandable-segments allocator on. The report&amp;#39;s own summary: ~9GB of safety margin, &amp;quot;OOM risk low&amp;quot; — and indeed &lt;strong&gt;no CUDA out-of-memory event appears anywhere in the surviving logs&lt;/strong&gt; for this pipeline.&lt;/p&gt;
&lt;p&gt;The throughput arithmetic, shown so you can check it &lt;em&gt;(derived)&lt;/em&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;15 steps × 218.5 s/step = &lt;strong&gt;3,277.5 s ≈ 54.6 min per 81-frame clip&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;3,277.5 s ÷ 81 frames = &lt;strong&gt;40.5 s per frame&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;81 frames ÷ 16 fps = 5.06 s of video → &lt;strong&gt;≈647× slower than realtime&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;The report&amp;#39;s own stated figure was &amp;quot;~49 min per 81-frame round&amp;quot;; we cannot fully reconcile 49 with 54.6 (likely rounding or TeaCache skip accounting), so we publish both.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Against expectations, the gap is worse than any rounding: planning notes from the same project targeted &lt;strong&gt;30–60 s per 3-second clip&lt;/strong&gt;. The measured pipeline delivered ~3,278 s per ~5-second clip. Reality ran an order of magnitude past the plan — which is exactly why this page exists; general-purpose AI assistants routinely quote times in the &amp;quot;~10 minutes per clip&amp;quot; range for this class of model on this class of hardware, and the only antidote to a wrong number is a measured one with its conditions pinned.&lt;/p&gt;
&lt;p&gt;Five clips rendered successfully across the window (output filenames carry timestamps, 2025-11-29 15:32 through 2025-12-01 10:36; byte sizes in the appendix). A 25-step variant appears in the shell history — derived ≈91 min/clip — but its outcome log did not survive, so we claim nothing from it.&lt;/p&gt;
&lt;h2&gt;Does keeping more blocks on the GPU help?&lt;/h2&gt;
&lt;p&gt;No — and this is the most quotable row in the report. Going from 20 to 24 resident blocks costs ~1.7GB more VRAM (the fp8 DiT works out to ≈0.49GB per block × 40 blocks, so 4 blocks ≈ 2.0GB — the arithmetic and the measured delta agree) and returns &lt;strong&gt;+0.64%&lt;/strong&gt; speed: 218.5 → 217.1 s/step &lt;em&gt;(report)&lt;/em&gt;. Returns had flatlined at the 20-block setting; the report&amp;#39;s conclusion, which our successful runs used, was &amp;quot;20-block residency is the optimum.&amp;quot;&lt;/p&gt;
&lt;p&gt;The practical reading for anyone tuning this pipeline on a similar rig: the residency dial is for &lt;em&gt;fitting&lt;/em&gt; the model without OOM. It is not a speed dial. If 20 blocks give you 218.5 s/step, more VRAM spent on residency buys you almost nothing back.&lt;/p&gt;
&lt;p&gt;We also swept persistence counts (&lt;code&gt;num_persistent_param_in_dit&lt;/code&gt; from 0 to 14B) and offload toggles across seven command-line runs in the tuning period; their outcome logs did not survive the later cleanup, so we publish no claims from them — the table above is everything that survives with numbers attached.&lt;/p&gt;
&lt;h2&gt;What an InfiniteTalk install costs on disk&lt;/h2&gt;
&lt;p&gt;Counting only what a run loads, the minimum is ≈104GB &lt;em&gt;(derived from the file table above)&lt;/em&gt;. Our actual weights tree measured &lt;strong&gt;241.8GB&lt;/strong&gt; (&lt;code&gt;du&lt;/code&gt;, 2026-02-01) because seven quantization variants sat on disk — fp8/int8, single/multi, ±LoRA, each ≈19.5GB, plus the 9,948,708,152-byte unquantized checkpoint &lt;em&gt;(log)&lt;/em&gt;. Counting the conda environment, HuggingFace cache, and outputs, the project&amp;#39;s total footprint was recorded at up to ~488GB before cleanup. Video models are not text models: budget disk before you budget VRAM.&lt;/p&gt;
&lt;h2&gt;Why we shelved it&lt;/h2&gt;
&lt;p&gt;The honest verdict, three ways:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Speed.&lt;/strong&gt; ~55 minutes per ~5-second 480P clip is an iteration killer. You cannot explore a prompt or a motion setting at this cadence.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Convenience.&lt;/strong&gt; For one-off talking-head clips, hosted closed models won outright — our own conclusion at the time was that this pipeline&amp;#39;s throughput made it &amp;quot;not the right tool&amp;quot; versus simply generating video with a hosted model.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Brittleness.&lt;/strong&gt; When we later tried to bring the environment back after the cleanup, it took six attempts across two days and one abandoned torch upgrade before a launch succeeded — that story, with the exact error strings and the version matrix that finally worked, is &lt;a href=&quot;/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/&quot;&gt;part 2 of this dispatch&lt;/a&gt;. If you run this stack, read that before touching &lt;code&gt;pip&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Where local InfiniteTalk still earns its keep: exact lip-sync control, offline operation, no per-clip cost at arbitrary length (streaming mode was built for this). If your use case is &amp;quot;a hundred variants of one spokesperson, overnight,&amp;quot; the per-clip hour is tolerable. If it is &amp;quot;five clips before lunch,&amp;quot; it is not.&lt;/p&gt;
&lt;p&gt;One more negative result worth recording: we never got InfiniteTalk running inside ComfyUI. ComfyUI has shipped native InfiniteTalk support since v0.28 (&lt;code&gt;WanInfiniteTalkToVideo&lt;/code&gt;), but every measurement in this dispatch comes from the official command-line pipeline. We make no claims about the ComfyUI path.&lt;/p&gt;
&lt;h2&gt;Reproduce it&lt;/h2&gt;
&lt;p&gt;The environment variables, verbatim from the shell history of the tuning period — set before every run, every time:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;export CUDA_VISIBLE_DEVICES=0,1
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
export NCCL_P2P_DISABLE=1
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The best-config invocation. Provenance, precisely: the common parts (weight paths, &lt;code&gt;--quant_dir $(pwd)&lt;/code&gt;, the repo example input) are verbatim from shell history; the 15-step / TeaCache-0.15 / no-offload flag set is the project&amp;#39;s &lt;code&gt;quick_generate.sh&lt;/code&gt; best config, preserved in session logs; &lt;code&gt;--motion_frame 9&lt;/code&gt; is the repo default as confirmed in a surviving run&amp;#39;s args log. The history&amp;#39;s own sweep variants (25-step and 10-step forms) are quoted in the appendix note below:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;python generate_infinitetalk.py \
  --ckpt_dir weights/Wan2.1-I2V-14B-480P \
  --wav2vec_dir weights/chinese-wav2vec2-base \
  --infinitetalk_dir weights/InfiniteTalk/infinitetalk_single_fp8.safetensors \
  --quant_dir $(pwd) \
  --input_json examples/single_example_image.json \
  --size infinitetalk-480 --sample_steps 15 --mode streaming \
  --motion_frame 9 --offload_model False --t5_cpu \
  --quant fp8 --use_teacache --teacache_thresh 0.15
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Models, at the exact revisions we downloaded (HuggingFace &lt;code&gt;refs/main&lt;/code&gt;, fetched 2025-11-25–27):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Repository&lt;/th&gt;
&lt;th&gt;Revision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;InfiniteTalk (fp8/int8 single &amp;amp; multi, ±LoRA)&lt;/td&gt;
&lt;td&gt;&lt;a href=&quot;https://huggingface.co/MeiGen-AI/InfiniteTalk&quot;&gt;MeiGen-AI/InfiniteTalk&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;d59847ebdacf19245bfca3fb23311c0cada8378a&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wan2.1 I2V 14B 480P base&lt;/td&gt;
&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P&quot;&gt;Wan-AI/Wan2.1-I2V-14B-480P&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;6b73f84e66371cdfe870c72acd6826e1d61cf279&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chinese wav2vec2 base&lt;/td&gt;
&lt;td&gt;&lt;a href=&quot;https://huggingface.co/TencentGameMate/chinese-wav2vec2-base&quot;&gt;TencentGameMate/chinese-wav2vec2-base&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;3991242c806928916fff4a8c0e4f76acf661b743&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Source code: &lt;a href=&quot;https://github.com/MeiGen-ai/InfiniteTalk&quot;&gt;MeiGen-ai/InfiniteTalk&lt;/a&gt; on the Wan2.1-I2V backbone (&lt;a href=&quot;https://github.com/Wan-Video/Wan2.1&quot;&gt;Wan-Video/Wan2.1&lt;/a&gt;). One warning from the run logs worth keeping in English: the A4000&amp;#39;s VAE has &lt;strong&gt;no tiling support&lt;/strong&gt; — &amp;quot;watch VRAM usage during inference&amp;quot; — which is why GPU 0 sits at ~64% while the 4090D runs at ~92%.&lt;/p&gt;
&lt;p&gt;The five finished clips, as the filesystem recorded them &lt;em&gt;(log)&lt;/em&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Output file (timestamp embedded)&lt;/th&gt;
&lt;th&gt;Bytes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;&lt;code&gt;infinitetalk-14B_infinitetalk-480_1_1_A_woman_is_passionately_singing_into_a_professiona_20251129_153238.mp4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;669,889&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;infinitetalk-14B_infinitetalk-480_1_1_..._20251130_143911.mp4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;623,874&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;infinitetalk-14B_infinitetalk-480_1_1_..._20251130_192900.mp4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;623,874&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;infinitetalk-14B_infinitetalk-480_1_1_..._20251130_233716.mp4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;623,874&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ex1_output.mp4.mp4&lt;/code&gt; (2025-12-01 10:36)&lt;/td&gt;
&lt;td&gt;423,170&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Part 2 — the reinstall saga and the working dependency matrix: &lt;a href=&quot;/posts/2026-08-26-infinitetalk-torch-load-dependency-matrix/&quot;&gt;InfiniteTalk dies at torch.load&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Measured on 2025-11-29 · RTX 4090D 24GB + RTX A4000 16GB dual-GPU workstation (Ubuntu 24.04, 503GB RAM). Raw markdown: &lt;a href=&quot;https://sigpulse.com/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured.md&quot;&gt;https://sigpulse.com/posts/2026-08-26-infinitetalk-14b-dual-gpu-measured.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><category>InfiniteTalk</category><category>Wan2.1</category><category>RTX 4090D</category><category>RTX A4000</category><category>fp8</category><category>Talking-Head Video</category></item></channel></rss>