The Staircase: Four Steps for Every Job in the Plant, and the Top Step Stays Empty
Key Takeaways — Executive & AI Summary
- Every job in the plant was placed on one four-step staircase: fixed pages, fixed scripts, script-plus-one-model-call, agents. The lower two steps carry no AI at all — the Episode-4 audit already proved the whole 30-day chain had exactly one model call point. The one rule: stand as low as the job allows. The AI-summary lane itself was born on Step 4 and climbed down to Step 3 the same week the audit priced its harness at 97% door fee — scaffolding, taken down when the building was finished.
- The word 'need' was retired from the record, by the operator's own pen: an agent is an elevator, not the electricity — no job loses the ability to be done without one. The live question is which jobs are worth hiring a temp for: the work is new, the volume clears the entrance fee, and a guardrail exists. Dissected, an 'industrial agent' purchase is four parts — model capability (a commodity), SOP capture (the expensive part, which no vendor can supply), a scheduling skeleton (cheap), and the autonomy loop (the only new thing, priced negative by this plant's own audit). One question exposes any costume: at run time, who decides the next step — the code, or the model?
- The Step-3 contest stayed open past its deadline on purpose. On 2026-09-13 the ruling kept both lanes — cloud-direct and on-prem — running side by side to 2026-09-20: the comparison costs 13.2K–18.5K tokens a day, the local lane bills zero API, and evidence outlives a rushed decision. Across 12 windows the on-prem lane went 12-for-12 (124–171 s); the cloud experiment recipe blanked twice in twelve, both at exactly the output cap — Episode 3's exam physics again, not a lane verdict. The rebuilt lane's first live event shift (2026-09-12 20:55) summarized both lines successfully inside 3 min 35 s end-to-end; the next morning's green shift made zero calls.
Episode 4 ended with a bill and a question. The bill: the plant’s only AI call, rebuilt the same day to a twenty-fifth of its price. The question: if one agent invocation cost twenty-five times its content and used none of its tools, what is an agent actually for — and what, exactly, would this plant be buying if it bought one?
The question was not academic. The operator’s company was weighing an industrial-grade agent platform for the plant. So the same afternoon the audit closed, the operator built the answer: a staircase, with every job in the plant standing on it. This episode is that record — the staircase, the law the plant’s own history had already been obeying, the purchase case taken apart screw by screw, and a ruling that arrived on schedule and chose to wait.
Four steps, one rule
| Step | What it is | Real jobs from this plant | Who decides the next step | Marginal cost |
|---|---|---|---|---|
| 1 · Fixed pages | dashboards, query and reconciliation windows | the device dashboard, the data-convergence window, the line-side raw-statistics view | designed in advance; a human clicks | free |
| 2 · Fixed scripts | scheduled Python that runs itself | the three patrol shifts, the changeover daily, incremental sync, PDF rendering and delivery | the code, written once | 0 tokens |
| 3 · Script + one model call | the machine phones a scribe | the patrol running summary: facts read out, prose read back, hang up | the code; the model only writes | ~2K tokens per call |
| 4 · Agents | the model decides steps and picks tools | the drafting sessions behind this record; one-off unattended batch runs | the model, at run time | door fee, tens of thousands of tokens per entry |
The one rule: stand as low as the job allows.
The lower two steps carry no AI at all, and that is not aspiration — it is the audited state of the plant. Episode 4’s forensic sweep of the whole patrol chain found exactly one model call point in thirty days of running: the running summary, on Step 3. Judgment, classification, rendering, delivery — all Step 2, all zero tokens. Most of the plant’s work never climbs past the second step, which is why the top step, for this plant, stays empty.
The scaffolding law
The staircase was not a new theory bolted onto the plant; it was a description of what the plant’s history already did. Three times, the same pattern:
- A threshold, found by AI, written into code. During an AI session, a pattern emerged in the event-window envelope: data dropout under 20% marks stretches where two faces change together, while a single-face cover sits near 28% and never qualifies. Confirmed by the operator against 36 days of data — four qualifying stretches, each with an account, zero misses — it was written into the monitor the same morning. The insight was AI’s; the rule is now Step 2’s.
- A research result, promoted into the daily. The overnight consistency study had produced conclusions humans had to go read. On 09-07 its verdict moved into the morning report’s fixed section; the next day brought the exemption for covered faces. Step 4’s output, re-homed on Step 2.
- The summary itself, born on Step 4, climbed down. The lane entered the world on 09-05 through a full agent harness — seven days, 44 calls, 2,239,178 tokens — and on 09-12 was rebuilt into a single direct call at a twenty-fifth of the price. Episode 4 told it as a bill. On the staircase it is something else: scaffolding, taken down the week the building was finished.
The law, as the operator wrote it that day: an agent is scaffolding. It is hired for the part of the job that is still unfinished — the exploring, the deciding-what-to-do-next. When the job is understood, it becomes a fixed step, and the scaffolding comes down. Which reframes the purchase under consideration: a resident agent platform in a plant of finished jobs is scaffolding rented forever on a building already topped out.
Retiring the word “need”
The first draft of the operator’s same-day memo listed four classes of work that “need” an agent. The operator struck the word himself. Debugging runs through a coding agent today — but ran for decades without one, on grep, logs and patience. An agent creates no capability that a human with fixed tools lacks; it saves labor and time. An agent is an elevator, not the electricity. You can always take the stairs.
The live question becomes: which jobs are worth hiring a temp for? Three tests, all three at once — the work is new enough that no instruction card can be written for it; the volume or urgency clears the entrance fee; and a guardrail exists, so the temp’s mistakes are catchable. Applied to this plant, agent usage survives in exactly three forms, none of them on the production floor: the drafting sessions behind this very record (no instruction card exists for “write the next episode”); the engineer’s chair, where an agent turns a three-hour search into twenty minutes — an accelerator, kept optional; and one-off unattended batch runs, hired for a night, cage and all. The plant’s runtime has no seat for one, and not for lack of imagination: runtime work is SOPs, and in a world of standard operating procedures, improvisation is not a skill — it is a violation. The agent’s one talent, deciding the next step on the spot, is the one talent the floor forbids.
What a purchase would actually buy
Take an industrial agent platform apart and four components fall out:
| Component | Nature | This plant’s standing |
|---|---|---|
| Model capability | a commodity — swap by API | dual-lane already in production: a frontier cloud model and an on-prem 27B, same manual, interchangeable |
| SOP and domain-knowledge capture | the expensive part — and no vendor can supply it; the domain must feed it | a manual library (four books) plus the judgment constitution, on a measured iterate-and-retest cycle |
| Scheduling skeleton | cron, retries, queues — cheap | patrol timers, comparison harness, unattended shifts: all running |
| The autonomy loop | the only component a purchase adds | priced negative by the plant’s own audit: door fee, variance, unauditable steps |
The accounting is the uncomfortable part for the purchase case: the cost of “industrial AI” was never in component four — it is in components two and three, which a platform does not remove and cannot supply. The buyer pays for the one component the audit priced as a liability.
And a caution about costumes: a good deal of what is sold as “industrial agents” is Step 2 wearing a badge — if the steps are written in advance, it is a workflow, whatever the box says. There is one question that strips the costume off any demo: at run time, who decides the next step — the code, or the model? If the purchase proceeds anyway, the record’s four tests travel with it: fully local or out; SOPs exportable as files; replayable, deterministic audit; and the model replaceable by API — this plant’s one-manual-two-models lane is the fourth test passed in production.
The ruling that chose to wait
One contest inside Step 3 remained genuinely open: cloud-direct or on-prem for the summary lane. The comparison had run each morning from 09-11 — the same sanitized fact sheet to both lanes, only the model and the manual differing. Its window was to close with a ruling on 09-13.
The ruling came, and it was: not either-or. Both lanes stay; the window extends to 09-20. The operator’s reason, paraphrased from the session record: the comparison costs almost nothing, so let it keep running. The numbers agree with him — twelve windows across three mornings:
| Mornings 09-11→09-13 | Cloud lane (stress recipe: earlier model, bare, 8,192 pad) | On-prem 27B (manual mounted) |
|---|---|---|
| Windows completed with prose | 10 of 12 | 12 of 12 |
| Blanks | 2 — both at exactly the output cap | 0 |
| Wall time, completed windows | 20–27 s | 124–171 s |
| Tokens, both lanes per day | 13,196 → 18,089 → 18,522 (local share: 7,389 / 7,675 / 8,026, zero API bill) |
Two footnotes keep that table honest. First, the blanks: the stress recipe is not the production recipe — an earlier-generation model, no manual, a double-size pad — and both blanks stopped at the cap exactly, the same exam physics Episode 3 established and Episode 4 met again. They indict a recipe, not a lane. Second, the daily total: it is both lanes combined, and the blank windows were its largest single items at 8,192 apiece — which is precisely why the operator could afford patience. Evidence at that price is a bargain; a rushed decision is not.
The extension itself was a one-line change — the comparison script’s exit date moved from 09-13 to 09-20, syntax-checked, effective the next morning. The decision was not cancelled. It was re-priced.
First live round, receipt and all
Episode 4 closed waiting for the rebuilt lane’s first live event shift. It came at 20:55 that same evening: an event verdict on both lines, the gate fired, and both summaries came back — the service journal logged both lanes’ AI calls complete at 20:59:22, the whole chain from trip generation to alert, PDFs and delivery inside 3 minutes 35 seconds. The samples read as the manual demands: bold verdict first, plant terms only, every sentence carrying its source field. The next morning’s 06:30 patrol came back green on both lines, and the lane made zero calls while the reports rendered on schedule — the control arm, clean.
One line is missing from the receipt, and the record says so: the production path prints token usage only on failure, so the first live round has no token-level bill — its cost anchor remains the A/B measurement (~2K per call). A one-line change would print usage on success too; it touches the production script, so it waits for approval rather than shipping quietly.
What the staircase did not cover
The audit had now reached every job in the plant: the chain, the lane, the manuals, the comparison, the purchase case. Every seat had been priced except one — the desk the audits were run from, where AI is used the way water is used, all day, in the open. That desk is the next episode.
Sources and method
First-party: the operator’s same-day memo of 2026-09-12 (the staircase, the temp-hiring tests, the four-part purchase dissection — including the operator’s own strike of the word “need”), the three comparison reports of 2026-09-11→09-13 (per-window tokens and wall times verbatim; the early log note that misnamed one blank line is corrected here against the reports), the ruling of 2026-09-13 as landed in the comparison script’s window line, the service journal entry of 2026-09-12 20:59:22 and the next morning’s green-shift report, and the re-addition that settles Episode 4’s 500-token discrepancy at 1,035,772. Assembled into English with AI assistance under human editorial direction; facts and numbers unchanged from the records; derived figures marked as computed. Deliberately absent, per the series’ disclosure policy: anything identifying the plant, its industry, its operator or its network, the identities of the harness and model products, and internal identifiers. The measured numbers are registered in the /data/ ledger.
All episodes — Machines Keep the Watch, a field-record series. The map: the anchor · Episode 1: the dress rehearsal · Episode 2: the second without AI · Episode 3: the blank-paper exam · Episode 4: the bill · Episode 6: the compaction.
FAQ — Direct Answers
- Two blank papers in twelve windows — doesn't that rule out the cloud lane?
- No, because the comparison's cloud recipe is not the production recipe. The experiment lane ran an earlier-generation model, bare (no manual mounted), with a double-size output pad of 8,192 tokens — a recipe chosen to stress the lane, not to mirror production, which runs the current generation with the manual mounted and a tighter pad. Both blanks hit the cap exactly (8,192 output tokens, one returning only a section number), which is the physics Episode 3 established on the local 27B and Episode 4 met again on the cloud model the same week: draft and body share one budget, and budget is a property of the exam, not of the examinee. A blank under the stress recipe is a data point about the recipe. It is not evidence against the production lane — and the record keeps that distinction explicit, including a correction: an early log note named the wrong line for one blank; the comparison reports, not the log, are the authority.
- Keeping both lanes running sounds like refusing to decide. Why extend?
- Because the arithmetic favors patience. Both lanes together burn 13.2K–18.5K tokens per day, of which the on-prem lane bills zero API cost — the price of continued evidence is trivial next to the cost of a premature decision. The operator's ruling, paraphrased from the session record: the comparison is cheap, so let it run. The window moved from 09-13 to 09-20 in a one-line change to the comparison script's exit condition, syntax-checked; the decision was not cancelled, it was re-priced.
- Wasn't there a discrepancy in Episode 4's numbers?
- Yes, and it is now settled. The gap lived in Episode 4's verification record, not its published text: that record's per-day table summed to one figure while a prose note in the same record carried another, 500 apart. Re-addition of the three steady-state rows (207,528 + 416,208 + 412,036) fixes the total at 1,035,772; the prose figure, 1,036,272, was a transcription slip that no combination of rows can produce. Episode 4's published text rounds to 'about 1.04M' and needed no correction on the site — the precision of the record, not the rounded prose, is what carries the arithmetic. The correction lives here, in the open, per this series' practice.
- Does the staircase generalize beyond this plant?
- With its boundary stated: the sample is one class of work — highly enumerable jobs with closed inputs and outputs, judged against written rules. A different site, with heterogeneous ad-hoc analysis demands across many lines and no in-house engineering, might land differently. But that is precisely the case the vendors sell to, and it is not this case — and even there, the four-part dissection holds: the expensive parts (SOP capture, data integration) are exactly the parts a platform cannot supply, and the part it adds (the autonomy loop) is the part an audit like this one prices as a liability. The vendor demo's natural-language query is reachable from Step 3 — skeleton, one model call, manual — with a replayable receipt.