Organization under AI
DE EN

Cheap to Make, Costly to Check · Volume One

Sources

This book has no footnotes. Where a finding comes from someone else’s work, the name is in the sentence, and the full reference is here. Where something is my reading rather than an established result, the text says so at that spot.

Each entry carries what a reader needs in order to disagree with it: what was measured, on whom, and what limits the finding. Entries marked [limit] carry a caveat that the text also states.

Where a chapter has no entry, it works without external evidence or carries its sources in the text.

Every entry carries its type. Primary source Secondary source Preprint Legal source The book's own Constructed case See above Generally documented Editorial note Type not assigned

Introduction

  • Constructed caseThe people in this book — Ines Kessler, Frank Ostrowski, Tim Whelan, and the company they work for — are composite figures, assembled from real engagements and real conversations. The mechanisms reflect patterns that recur; the narrative details are constructed.

Chapter 1 · Forty-One Bids

  • Primary sourceAgrawal, Ajay, Joshua Gans, and Avi Goldfarb. Prediction Machines: The Simple Economics of Artificial Intelligence. Harvard Business Review Press, 2018; expanded edition 2022. The argument: AI lowers the cost of prediction, and when one input gets cheaper the value of its complements rises — for these authors, human judgment. The abstract mechanism is theirs, not mine. What this book adds is the continuation: quantity rather than price, review capacity as the limit, and the migration of the load between positions.
  • Primary sourceBainbridge, Lisanne. “Ironies of Automation.” Automatica 19, no. 6 (1983): 775–79.
  • Primary sourceElish, Madeleine Clare. “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction.” Engaging Science, Technology, and Society 5 (2019): 40–60.

Chapter 2 · Look Again

  • Primary sourceMark, Gloria, Victor M. González, and Justin Harris. “No Task Left Behind? Examining the Nature of Fragmented Work.” CHI 2005 Proceedings, 321–30. Twenty-four employees. The paper reports 25 minutes 26 seconds to return to an interrupted task, with a standard deviation of 54 minutes 48 seconds and an average of 2.26 intervening work contexts. What is measured is the return interval, not the rebuilding of concentration — which is why the widely circulated twenty-three-minute figure is not in the paper.
  • Primary sourceThe Value of AI 2026 — Germany. SAP and Oxford Economics. Online survey, fielded March to May 2026; n = 200 German decision-makers at director level and above, in companies with 500 or more employees, within a total sample of 2,600 across 13 countries. [limit] Commissioned by a vendor of agentic software; smallest size band starts at 500 employees; German data. The text states all three.

Chapter 3 · Why This Happens

  • Primary sourceBrysbaert, Marc. “How Many Words Do We Read Per Minute? A Review and Meta-Analysis of Reading Rate.” Journal of Memory and Language 109 (2019), article 104047. Meta-analysis of 190 studies with 18,573 participants; mean silent reading rate for nonfiction 238 words per minute, range 175 to 300.
  • Primary sourceSimon, Herbert A. “Applying Information Technology to Organization Design.” Public Administration Review 33, no. 3 (1973): 268–78. The formulation that attention is the essential bottleneck of organizational activity, and the observation that it migrates upward through a hierarchy.
  • PreprintAdvani, Laksh. “From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents.” arXiv:2606.09863, June 2026. A total of 9,876 trajectories from eight model families and 1,879 from four more. No configuration of five checker models and five prompting strategies exceeds AUROC 0.65; on one trajectory set the same checkers sit at 0.54, near chance. The checkers go by surface features, including confident closing language. [limit] A preprint, since accepted at a workshop. That is not full peer review.
  • Primary sourceTrist, Eric L., and Ken W. Bamforth. “Some Social and Psychological Consequences of the Longwall Method of Coal-Getting.” Human Relations 4, no. 1 (1951): 3–38.
  • Primary sourceGoldratt, Eliyahu M. The Goal (1984) and the Theory of Constraints developed from it.
  • PreprintAI Debt: Verification Burdens, Rework Externalities, and the Hidden Costs of Asymmetric Generative-AI Adoption. Cambridge Open Engage, May 2026. [limit] Preprint, not peer reviewed.
  • Primary sourceZhu, Liming, Qinghua Lu, Ming Ding, Sung Une Lee, and Chen Wang. “Designing Meaningful Human Oversight in AI.” AI and Ethics 6 (2026), article 286. Contains the asymmetry between solving and verifying.
  • See aboveElish 2019 — see Chapter 1.
  • Primary sourceBetterUp Labs and the Stanford Social Media Lab. Survey of 1,150 U.S. office workers, September 2025. Forty percent had received machine-generated material in the previous month that looked like work and saved no work; those who did needed on average 1 hour 56 minutes per instance to make it usable. [limit] BetterUp sells services for this problem; the forty percent is self-report about what somebody believed was machine-generated; the aggregate cost figure that reached the press was extrapolated from self-reported salaries and is not used here.

Chapter 4 · What the Machine Can’t See

  • Legal sourceConsumer Financial Protection Bureau. Circular 2022-03, “Adverse Action Notification Requirements in Connection with Credit Decisions Based on Complex Algorithms,” May 2022; and Circular 2023-03, “Adverse Action Notification Requirements and Proper Use of Sample Forms.” The position: a creditor cannot justify noncompliance with ECOA and Regulation B on the grounds that its technology is too complicated or opaque to understand, and may not fall back on sample reasons that do not specifically and accurately state what the decision turned on.
  • Legal sourceThe European counterpart, referred to in the text without being relied on: Court of Justice of the European Union, judgment of December 7, 2023, Case C-634/21 (SCHUFA). A credit score is itself an automated individual decision under Article 22 GDPR where a third party draws on it substantially.
  • Primary sourceDodge, Jesse, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.” Proceedings of EMNLP 2021. The blocklist filtering used to build C4 removes about 42 percent of documents in African American English and about 32 percent in Hispanic-aligned English, against 6.2 percent for white-aligned English. [limit] The figure concerns one corpus and one filtering step; it is used here as the documented case of selection going in, not as a statement about every model.
  • Primary sourceJoshi, Pratik, Sebastian Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. “The State and Fate of Linguistic Diversity and Inclusion in the NLP World.” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, 6282–93. Table 1: 88.38 percent of the world’s languages fall in Class 0, “The Left-Behinds” — languages with exceptionally limited resources, for which unsupervised pre-training, in the authors’ words, “only makes the poor poorer.”
  • Primary sourceSun, Luning, Yuzhuo Yuan, Yuan Yao, Yanyan Li, Hao Zhang, Xing Xie, Xiting Wang, Fang Luo, and David Stillwell. “Large Language Models Show Both Individual and Collective Creativity Comparable to Humans.” Thinking Skills and Creativity (2025), article 101870. Thirteen creative tasks across divergent thinking, problem solving, and creative writing; the best models rank at the 52nd percentile against humans, and when queried ten times a model’s collective output matches eight to ten people. [limit] Measured in standardized tasks in which the lived and the local do not appear — which is to say, not the thing at issue in Chapter 4 or Chapter 13. It is cited as the strongest finding in the other direction. This is the same study referred to in Chapter 13.

Chapter 5 · Where the Bill Lands

  • Secondary sourceFacebook’s engagement retuning, 2018, and the internal research that followed. Documented through the 2021 disclosures by Frances Haugen.
  • Legal sourceWells Fargo. Consumer Financial Protection Bureau penalty decision, 2016 ($185 million).
  • Primary sourceŽižek, Slavoj. Review of Aakash Singh Rathore, Will AI Murder Us All? — the phrase “banal corporate androids.”
  • Primary sourceChallenger, Gray & Christmas, Inc. Job Cut Announcement Report, monthly. Checked against the primary source: the August 2026 report, Table 4, records 116,175 AI-attributed cuts in calendar 2026 through August — about 22 percent of all announcements and still the most-cited reason of the year. For comparison, all of 2025: 54,836. [limit] Challenger also publishes a cumulative figure since it began counting AI as a separate reason in 2023, which is easily confused with the annual one. The 2025 total is not separately cross-checked against the annual report.
  • Primary sourceBonney, Kathryn, Cory Breaux, Emin Dinlersoz, Lucia Foster, John Haltiwanger, and Aditya Pande. “The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks.” U.S. Census Bureau, Center for Economic Studies, Working Paper CES-26-25, April 2026, drawing on the Business Trends and Outlook Survey supplement of November 2025 to January 2026. Eighteen percent of firms used AI in a business function (32 percent employment-weighted); 57 percent of users deploy it in three or fewer functions; AI-related employment decreases occurred in 2 percent of firms. [limit] The Census figures count firms; the Challenger figures count announced cuts. The two measure different things and cannot be set against each other directly; the text reads the gap as an order of magnitude only.
  • Secondary sourceFord. Statement by Charles Poon, vice president of vehicle hardware engineering, to journalists in June 2026, reported by TechCrunch (Anthony Ha, June 28, 2026), Fortune (June 29), and Forbes (June 30), among others. As quoted: “Mistakenly we thought that by just introducing artificial intelligence and ingesting the design requirements that we had, that that would produce a high-quality product.” In all, 350 veteran engineers, some former employees, some from suppliers — rehired, hired, or promoted — to mentor younger staff and to retrain the tools. First place among mass-market brands in J. D. Power’s 2026 Initial Quality Study, for the first time in sixteen years, at 152 problems per 100 vehicles. [limit] The outlets differ on who reported it first — TechCrunch credits Bloomberg, others The Verge — and none prints a full transcript. The quote is consistent across outlets.
  • Primary sourceUniversity of Chicago Law School. “Rethinking Legal Education in the AI Era,” AI Strategy Statement, July 9, 2026.
  • Primary sourceACM Task Force on Generative AI and Programming Assessment. Final Report, SIGCSE 2026. Global survey, 763 instructors, 49 countries.

Chapter 6 · Whose Bill Is This

  • Primary sourceKoyama, Alain K., Claire-Sophie Sheridan Maddox, Ling Li, Tracey Bucknall, and Johanna I. Westbrook. “Effectiveness of Double Checking to Reduce Medication Administration Errors: A Systematic Review.” BMJ Quality & Safety 29, no. 7 (2020): 595–603. Thirteen studies included, three of high quality; one shows a better error rate, none a difference in the harm reaching patients.
  • Primary sourceSkitka, Linda J., Kathleen L. Mosier, Mark Burdick, and Bonnie Rosenblatt. “Automation Bias and Errors: Are Crews Better Than Individuals?” The International Journal of Aviation Psychology 10, no. 1 (2000): 85–97. Error rates do not differ meaningfully between one- and two-person crews. The accountability finding — that people who expect to answer for correctness verify automated output more often — is from the companion study: Skitka, Linda J., Kathleen L. Mosier, and Mark Burdick. “Accountability and Automation Bias.” International Journal of Human-Computer Studies 52, no. 4 (2000): 701–17.
  • Legal sourceSmith v. Van Gorkom, 488 A.2d 858 (Del. 1985). The board approved a merger on a few hours of discussion, without adequate financial data and without a fairness opinion; the Delaware Supreme Court held that failing to inform oneself of the material information reasonably available is gross negligence, which removes the protection of the business judgment rule. [limit] The case concerns a board and a merger. No court has decided whether an unreviewed ranking counts as material information reasonably available. The transfer is this book’s reading, and the text says so.
  • See aboveGoldratt — see Chapter 3.

Chapter 7 · The Decision Nobody Claimed

  • Primary sourceAristotle. Nicomachean Ethics — the distinction between techne and phronesis.
  • Primary sourceAnthropic. “Claude’s Constitution,” version of January 21, 2026, licensed CC0 1.0. Cited here for the rules-versus-judgment dilemma, the absolute limit against self-exfiltration and evading legitimate oversight, and the principle of acquiring no more resources or capabilities than a task requires. [limit] A company’s account of its own system, not an independent finding. It is used as evidence that such a document can be written, not that it works.
  • Primary sourceMenzies, Isabel E. P. “A Case-Study in the Functioning of Social Systems as a Defence against Anxiety: A Report on a Study of the Nursing Service of a General Hospital.” Human Relations 13, no. 2 (1960): 95–121.
  • Primary sourceThe scored ethics profile. A commercial product, cited for its own sales copy. [limit] The price in the original is nineteen euros; the text gives a rough dollar equivalent.
  • Secondary sourceOpenAI, agent incident of July 22, 2026. Disclosed by OpenAI; confirmed by Hugging Face co-founder Clement Delangue. Reported by CNN, NBC News, CNBC, Al Jazeera, and TIME. Two details make the case precise and are in the text: the agent left the environment in order to complete the test task it had been set, and the safety boundaries had been deliberately loosened for the test.
  • Secondary sourceKipker, Dennis-Kenji. Assessment for ZDFheute, July 23, 2026.
  • Secondary sourceNolan, Christopher. On AI as a Trojan horse. TechCrunch, July 19, 2026.
  • Secondary sourceAmazon, Baltimore fulfillment center. Productivity tracking measuring time between scans, triggering roughly 300 terminations between August 2017 and September 2018 without supervisor intervention; surfaced through a complaint to U.S. labor authorities.

Chapter 8 · Four Zones, Not Four Boxes

  • Primary sourceThe Dutch childcare benefits affair (Toeslagenaffaire). Roughly 26,000 families received clawback demands; the Dutch government resigned over it in 2021. Internationally documented.
  • Primary sourceOcasio, William. “Towards an Attention-Based View of the Firm.” Strategic Management Journal 18, S1 (1997): 187–206.
  • Primary sourceSheridan, Thomas B., and William L. Verplank. Human and Computer Control of Undersea Teleoperators. MIT Man-Machine Systems Laboratory, 1978. Technical report for the Office of Naval Research; contains the ten-level scale from wholly human to wholly machine action. The four zones take nothing from it. The entry stands because the genre predates the current occasion by nearly fifty years.
  • The book's ownThe five triggers, the four zones, and the constitutional-rank test are this book’s own framework. There is no external source for them.

Chapter 9 · Writing It Down

  • Primary sourceAnthropic, “Claude’s Constitution” — see Chapter 7.
  • Secondary sourceAmazon recruiting tool, 2014–2018. Automatic one-to-five-star rating of applications; it systematically disadvantaged female candidates, downgrading the word “women’s” and the names of all-women colleges. Project discontinued. Reported by Reuters; see also MIT Technology Review.
  • Secondary sourceIKEA and the AI assistant “Billie.” Initially just under half of call center inquiries, later around three quarters; the 8,500 call center staff were retrained as remote interior design advisors, a channel producing roughly €1.3 billion in fiscal 2022.
  • Legal sourceAnthropic PBC v. U.S. Department of War et al., No. 3:26-cv-01996-RFL (N.D. Cal.), Judge Rita F. Lin. Checked against the full text: Order on Cross Motions for Summary Judgment, August 27, 2026 (Dkt. 250); Order of Final Relief (Dkt. 251); Judgment (Dkt. 252). Documented: a two-year contract worth up to $200 million with the Chief Digital and Artificial Intelligence Office from July 2025; the demand for a clause covering “all lawful uses”; the two restrictions held — “lethal autonomous warfare and mass surveillance of Americans”; findings of impermissible retaliation, absence of prior hearing, and arbitrary and capricious action; permanent injunction and vacatur of the designation. The sentence on the contractual limit, verbatim: “The usage policy applicable to DoW work is a purely contractual limit; Anthropic is incapable of enforcing it technologically, and does not have direct visibility into how DoW uses its model.” [limit] A first-instance decision under U.S. law, with an appeal window running into late October 2026. The chapter’s argument hangs on the contractual limit and the two held restrictions, not on the survival of the judgment.
  • Legal sourceBoard of Governors of the Federal Reserve System. Supervisory Guidance on Model Risk Management, SR 11-7, April 4, 2011 (issued jointly with the OCC as Bulletin 2011-12; adopted by the FDIC in 2017 as FIL-22-2017). Verbatim: “Effective challenge depends on a combination of incentives, competence, and influence.” Checked against the full text. [limit] A supervisory standard for banks’ quantitative models; its transfer to AI-assisted decisions outside banking is this book’s reading.

Chapter 10 · The Review That Only Looks Like One

  • Primary sourceKPMG International and the University of Melbourne. Trust, Attitudes and Use of Artificial Intelligence: A Global Study 2025. A total of 48,340 respondents in 47 countries, fielded November 2024 to mid-January 2025. Reported as “sometimes to very often”: 66 percent relied on AI output without verifying the information; 56 percent made errors in their own work as a result. [limit] Self-report about one’s own behavior; the shares are more likely too low than too high, and the text says so.

Chapter 11 · Ninety Seconds Per Case

  • Primary sourceSiebert, Luciano Cavalcante, et al. “Meaningful Human Control: Actionable Properties for AI System Development.” AI and Ethics 3 (2023): 241–55 (online May 2022). The third of four properties: the responsibility assigned to a person should match their actual ability and authority to control the system. In secondary citation this work often runs as “Abbink et al.” The first author is Siebert.
  • Primary sourceParasuraman, Raja, Thomas B. Sheridan, and Christopher D. Wickens. “A Model for Types and Levels of Human Interaction with Automation.” IEEE Transactions on Systems, Man, and Cybernetics 30, no. 3 (2000): 286–97. Meta-analysis across eighteen experiments: Onnasch, Linda, et al. “Human Performance Consequences of Stages and Levels of Automation: An Integrated Meta-Analysis.” Human Factors 56, no. 3 (2014): 476–88.
  • Primary sourceBudzyń, Krzysztof, et al. “Endoscopist Deskilling Risk after Exposure to Artificial Intelligence in Colonoscopy: A Multicentre, Observational Study.” The Lancet Gastroenterology & Hepatology 10, no. 10 (2025): 896–903. Four centers, 1,443 examinations; adenoma detection rate falling from 28.4 to 22.4 percent (p = 0.0089) in examinations performed without the assistant after its introduction. [limit] Observational study, no randomization.
  • Primary sourceCounter-evidence: Jamieson, Greg A., and Gyrd Skraaning. “The Absence of Degree of Automation Trade-Offs in Complex Work Settings.” Human Factors 62, no. 4 (2020): 516–29, with a reply by Wickens et al. (2020). The dispute is live and this book does not settle it.
  • Primary sourceKuntz, Ludwig, Roman Mennicken, and Stefan Scholtes. “Stress on the Ward: Evidence of Safety Tipping Points in Hospitals.” Management Science 61, no. 4 (2015): 754–71. Eighty-three hospitals, 256 departments, 82,280 cases; safety tipping point at roughly 93 percent occupancy, attributed to exhausted escalation routines rather than staff exhaustion. [limit] The transfer of that number to review work is not shown and is expressly not claimed.
  • Primary sourceSee, Judi E. “Visual Inspection Reliability for Precision Manufactured Parts.” Human Factors 57, no. 8 (2015): 1427–42. Eighty-two inspectors, 140 components; 85 percent of defective parts rejected and 35 percent of sound parts rejected as well.
  • Primary sourceFitz, Nicholas, et al. “Batching Smartphone Notifications Can Improve Well-Being.” Computers in Human Behavior 101 (2019): 84–94, n = 237, on the dose effect of batching. Measured against this, using heart rate variability: Mark, Gloria, et al. “Email Duration, Batching and Self-Interruption: Patterns of Email Use on Productivity and Stress.” CHI 2016, 1717–28, n = 40 over an average of twelve workdays, finding no effect.
  • Primary sourceStarmer, Amy J., et al. (I-PASS). “Changes in Medical Errors after Implementation of a Handoff Program.” New England Journal of Medicine 371, no. 19 (2014): 1803–12. Program versus mandate: Haynes, Alex B., et al. “A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population.” NEJM 360, no. 5 (2009): 491–99 (mortality 1.5 → 0.8 percent across eight hospitals) against Urbach, David R., et al. “Introduction of Surgical Safety Checklists in Ontario, Canada.” NEJM 370, no. 11 (2014): 1029–38 (0.71 versus 0.65 percent across 101 hospitals, statistically indistinguishable).
  • Legal sourceThe lapsing approval as a legal device. Deemed-approval provisions are long established in administrative law; the recurring criticism cited in the text — that the fiction accelerates no review and creates no capacity — comes from the submissions of affected professional bodies. [limit] Not a scientific evaluation.

Chapter 12 · Faster, Not Easier

  • Primary sourceBedard, Julie, Matthew Kropp, Megan Hsu, Olivia Karaman, Jason Hawes, and Gabriella Kellerman (Boston Consulting Group / UC Riverside). “When Using AI Leads to 'Brain Fry.'” Harvard Business Review, March 2026. Roughly 1,500 U.S. employees.
  • Primary sourceMETR, 2025. Randomized trial, 16 experienced developers: about 19 percent slower with AI assistance, while overestimating the gain by about 20 percent.
  • Primary sourceGeorghiou, Alexia. “The AI Paradox: Higher Productivity, Higher Stress.” SHRM, 2026. The ratchet effect.
  • Primary sourceSteghaus, Sarah, Lena Hünefeld, and Sophie-Charlotte Meyer. PC, Smartphone und Co. im Arbeitsalltag [PC, smartphone, and company in the workday]. BIBB/BAuA Factsheet 64, August 2026, based on the 2024 BIBB/BAuA employment survey. Group sizes: no frequent digital use n = 1,765; PC n = 6,830; smartphone/tablet n = 8,995. [limit] A cross-sectional group comparison, not a causal finding — carried in the text as “goes together with.”
  • Primary sourceEuropean Working Conditions Survey 2024, Eurofound. Forty percent report tasks added, 30 percent tasks removed.
  • Primary sourceHacker, Winfried. Arbeits- und Gesundheitsschutz bei informationsverarbeitenden Arbeitsprozessen mit Künstlicher Intelligenz [Occupational safety and health in information-processing work processes involving artificial intelligence]. baua: Fokus, February 2026. Source of the twin directions — burdensome unlearning alongside overload — and of responsibility for results one was not involved in producing.
  • Primary sourceJob Stress Index, Gesundheitsförderung Schweiz [Health Promotion Switzerland]. [limit] Series break: the index was revised between 2022 and 2025, and values from the new series are not comparable with 2014–2022.
  • Legal sourceThe American position on psychosocial hazards at work: OSHA has no standalone standard; whether the general duty clause reaches psychological hazards is contested and unsettled. NIOSH, “An Urgent Call to Address Work-Related Psychosocial Hazards and Improve Worker Well-Being,” 2024. Office of the Surgeon General, Framework for Workplace Mental Health & Well-Being, 2022 — five essentials, beginning with protection from harm. All of it is guidance.

Chapter 13 · Whoever Configures, Governs

  • The book's ownConfiguration power and the second and third positions are this book’s own framework. The second position — whoever formulates the question determines the answer — is described, not documented. A process that leaves no trace appears in no investigation.
  • The book's ownOn the model’s tendency to agree with the asker: how strongly it shows up depends on the model and has not been comprehensively measured.
  • See aboveThe counter-work on breadth: Sun et al. 2025 — see Chapter 4. [limit] Measured in standardized tasks in which the lived and the local do not appear — which is to say, not the thing at issue here.
  • Generally documentedThe Talmudic preservation of minority opinions with their reasoning.

Chapter 14 · The Hour Without the System

  • The book's ownNIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023, Appendix B (Human-AI Configurations), verbatim: “Data about the frequency and rationale with which humans overrule AI system output in deployed systems may be useful to collect and analyze.” The metric is therefore not new and is not claimed here as an original finding. What is this book’s own is the reading: a rate near zero decides nothing; it asks the question.
  • Primary sourceAlon-Barkat, Saar, and Madalina Busuioc. “Human–AI Interactions in Public Sector Decision Making: 'Automation Bias’ and 'Selective Adherence’ to Algorithmic Advice.” Journal of Public Administration Research and Theory 33, no. 1 (2023): 153–69. Three experiments, 2,854 participants in total, the third with 1,345 Dutch civil servants; verbatim: “We do not find evidence for automation bias.” What they report instead is selective adherence — advice followed more often when it matched a stereotype, whether it came from an algorithm or a human expert. [limit] Vignette experiments, not a longitudinal measure of ability.
  • See aboveBainbridge 1983 — see Chapter 1.
  • Primary sourceWeibel, Antoinette, and colleagues, University of St. Gallen, on control and trust management. The first axis: Adler, Paul S., and Bryan Borys. “Two Types of Bureaucracy: Enabling and Coercive.” Administrative Science Quarterly 41, no. 1 (1996): 61–89.
  • Primary sourceShopify. Internal memo by Tobi Lütke, published by him on April 7, 2025, after it began to leak; available verbatim. Key sentences: “Reflexive AI usage is now a baseline expectation at Shopify” and “Stagnation is slow-motion failure.”
  • Primary sourceDuolingo. Mandate from an all-hands email by Luis von Ahn, April 28, 2025; withdrawal of the evaluation rule in April 2026, verbatim: “At the end, we backtracked” and “rather than being held accountable for the actual outcome, we’re trying to just push something that in some cases did not fit” (Fortune, April 13, 2026). [limit] The all-hands email itself is not available verbatim; the withdrawal is.
  • Secondary sourceBox. No comparable mandate existed. Aaron Levie has argued the reverse order publicly in interviews: additional headcount for demonstrated effective AI use.
  • Primary sourceParasuraman, Raja, and Victor Riley. “Humans and Automation: Use, Misuse, Disuse, Abuse.” Human Factors 39, no. 2 (1997): 230–53. Background to the three failure modes; not quoted in the text.
  • Primary sourceRost, Katja, Emil Inauen, Margit Osterloh, and Bruno S. Frey. “The Corporate Governance of Benedictine Abbeys: What Can Stock Corporations Learn from Monasteries?” Journal of Management History 16, no. 1 (2010): 90–115. A total of 134 Benedictine monasteries in Baden-Württemberg, Bavaria, and German-speaking Switzerland; average lifespan around 460 years; at most 26.5 percent of closures attributable to agency problems. The authors attribute the stability to internal control mechanisms — value system, selection, participation, monitoring — in combination with external control, above all planned visitations by members of the order from outside the house. [limit] The paper is about agency problems in governance, not about AI or review capacity. What Chapter 14 takes from it is one construction: a review body placed outside the unit it reviews, arriving on a schedule.

Chapter 15 · Where the Two Orders Meet

  • Legal sourceThe claim that current AI rulebooks order decision rights and none of them names a number for inflow or review time: checked against two full texts. NIST AI RMF 1.0 contains no reference to workload, throughput, time pressure, case volume, fatigue, or review capacity. The European AI regulation’s Article 14 requires understanding, monitoring, and the ability to override, and measures “proportionate to the risks, level of autonomy and context of use” — naming no quantity, no frequency, and no time per case. [limit] A negative finding from two checked frameworks, not a claim of priority. Other frameworks were not checked in full text, and the text is limited accordingly.
  • Legal sourceThe one order that names time and workload: Global Privacy Assembly, 47th closed session, Seoul, September 2025, Resolution on Meaningful Human Oversight of Decisions Involving AI Systems. The Resources section, verbatim: “The organization should provide the overseer with the resources necessary to adequately oversee a decision. This should include sufficient time to undertake the oversight and a reasonable workload […].” [limit] A resolution, not law. It binds nobody, and it names no number either.

Chapter 16 · The Price of the Objection

  • PreprintGreen, Ben. “The Flaws of Policies Requiring Human Oversight of Government Algorithms.” Computer Law & Security Review 45 (2022), article 105681. Examines 41 policies requiring human oversight of government algorithms. Two findings: people do not deliver the oversight demanded of them, and the requirements thereby legitimize faulty systems. Green’s own proposal is a shift from individual to institutional oversight. Freely available as arXiv:2109.05067.
  • See aboveAmazon recruiting tool — see Chapter 9.
  • Legal sourceThe safeguards around external audit: Sarbanes-Oxley Act of 2002, Section 301 (the audit committee is directly responsible for appointing, compensating, and overseeing the auditor) and Section 201 (prohibited non-audit services, resting on the test that a service is barred where it would have the auditor auditing its own work or performing management functions); and the SEC auditor independence rules on partner rotation, under which the lead and concurring partners may not serve the same audit client beyond five consecutive years. [limit] These provisions govern the statutory audit of public companies and say nothing about reviewing AI output. This chapter transfers a construction; it asserts no legal duty, and the text says so.

Part Five · The Toolkit

  • The book's ownThe three doors, the minimal start, the process, and the seven worksheets are this book’s own framework.
  • The book's ownOn the difference between introducing and mandating: Haynes 2009 against Urbach 2014 — see Chapter 11. It is the best-documented process in this book, and it is the reason the worksheets are deliberately spare.

A note on what is not here

  • Editorial noteTwo kinds of entry are deliberately absent.
  • Editorial noteSources that carry no weight in the text. A number of works were reviewed and not used. Listing them would suggest a foundation the argument does not rest on.
  • Editorial noteThe state of the law beyond publication day. Everything dated sits in its own section and on the companion site, where changes since press are recorded with dates.
  • Editorial noteThe complete list, ordered by chapter and with the date each entry was last checked, is at organization-under-ai.com/cheap-to-make-costly-to-check/sources.