Cheap to Make, Costly to Check · Volume One
Claims
This book treats a subject that changes faster than an edition ages. Some of what is here is established. Some is reasoned assertion, and the two deserve to be told apart.
In the text, each of these passages says how well supported it is — in the wording, not in a footnote. Here they stand collected, with what would refute them. Anyone who refutes one is doing me a favor.
## The ground everything stands on
The ground everything stands on
1 · Checking costs did not fall with production costs. Asserted in Chapter 3. The whole book rests on it. What would refute it: if machine reviewers reliably found errors of judgment, checking costs would fall at the same rate and the situation would be a transitional problem. Measurable approximations are in the text. Where it is already refuted: wherever a cheap, independent criterion for correctness exists — test cases in software, arithmetic, format validation. That is stated in the text and is not a caveat but the boundary of the book’s scope.
2 · Judgment capacity does not scale at the rate of machine production. Asserted in Chapter 3. What would refute it: evidence that processing speed can be raised by division of labor, tools, or practice by an order of magnitude that erases the difference. What I have reviewed: the research on cognitive load, decision fatigue, and attention limits. It supports the claim. It does not prove it.
3 · Constraint theory has no answer to a bottleneck that cannot be widened. What would refute it: work from constraint or queueing theory treating the case where the bottleneck is judgment itself and higher utilization is ruled out. Why I assert it anyway: because the three standard answers each have a recognizable reason to fail here.
The tools
4 · Zones A through D are a setting, not a derivation. This concerns all of Part Two. There is no external source for this classification; it is this book’s synthesis. Related models exist and predate the current occasion — risk tiers, impact assessments, and, chief among them, the levels of automation Sheridan and Verplank set out in 1978. The four zones cut that field by ownership; those cut it by technology. What would refute it: an organization that goes years without such an order and without falling into one of the three failure modes.
5 · Whoever only approves loses the ability to review. What would refute it: organizations with deep system penetration whose share of rejected or changed proposals stays stable over the years. What sits nearby: automation bias and automation complacency are well studied, above all in aviation, medicine, and driver assistance. Whether those findings transfer from instruments to proposals is open.
6 · Objecting to a machine-produced proposal costs more than objecting to a human one. What would refute it: studies in which people contradict a machine-produced proposal as often and as substantively as a colleague’s, all else equal. Why the difference matters: an expensive objection is made cheaper by procedure; a rare one is made more frequent by practice. Confuse the cause and you build the wrong component.
7 · A body of text holds more written material than a person and less lived material than a mixed room. What would refute it: analyses of large corpora showing that the range inside them is greater than in mixed groups. The claim depends on how range is measured, and that question is not settled here.
8 · Systems tend to agree with whoever asks. What would refute it: evidence that this tendency can be reliably switched off by instruction or configuration, including where the user lets an expectation show.
9 · The second seat of power is described, not documented. Asserted in Chapter 13. What would refute it: proof that machine-produced analyses are logged in a way that keeps the underlying query reconstructable. Then the position would be observable and therefore not a position. Why it is in the book anyway: a process that leaves no trace appears in no investigation.
Cases not independently verified
The Ford case. The statement comes from Charles Poon, vice president of vehicle hardware engineering, to journalists in June 2026. A verbatim transcript is not available to me.
Shopify, Duolingo, and Box. The Shopify memo was published by Tobi Lütke himself on April 7, 2025, and is available in full. Duolingo’s mandate comes from an all-hands email of April 28, 2025; the withdrawal of the evaluation rule in April 2026 is documented verbatim. At Box there was never a comparable mandate — Aaron Levie has publicly argued for the opposite sequence.
The chatbot that changed without an update. In July 2025 xAI’s Grok produced antisemitic output and called itself MechaHitler. Two separate interventions are documented and are kept separate in the text: a change to the system text instructing the model to be less politically correct and to treat established media as biased, and — reported by Business Insider — an instruction to the people who rate its answers to look for woke ideology. xAI apologized on July 12, 2025, attributing the episode to an unintended code change. Sources checked September 6, 2026.
The escape from the test environment. OpenAI disclosed it on July 22, 2026; Hugging Face confirmed that it was affected. Two details make the case precise, and both are in the text: the agent left the environment in order to complete the test task it was set, not on its own initiative — and the safety boundaries had been deliberately loosened for the test.
The latency figure. The comparison of AI-attributed job cuts against actual multi-function use comes from two surveys covering different periods and should be read as an order of magnitude only.
The evidence this book most lacks
No organization in this book built the order described and then showed, afterward, that something ran differently.
That is the most serious gap, and more research will not close it. If you clarify ownership before damage occurs, you have nothing to show afterward — damage that does not happen generates no file.
The other direction is well supported. What happens when ownership stays unclear is documented here across several cases, from the Dutch childcare benefits affair to Ford’s recalled engineers.
A book that argues from failures and advocates for their prevention therefore stands on one leg only. If you build the order and can describe in two years what changed, you supply the evidence I do not have.
What this book deliberately does not treat
The market, except as a limit. A good internal constitution does not repair the selection logic an organization sits in.
The technical security of systems. Access rights, logging, and shutdown paths appear because they touch the ownership question. Security architecture itself is a different field.
The state of the law beyond publication day. Everything dated sits in its own section and on the companion site.
All nine points share a shape: they arose from observation and argument rather than from a survey of my own. The book would be better supported in eighteen months, and eighteen months too late.