top of page
industry.jpeg
Human on the Loop Working Paper

As enterprises deploy autonomous AI agents, Human-on-the-Loop (HOTL) is touted as the ideal governance model. This paper proves that real-time human monitoring naturally breaks down as AI scales, establishing 4 core insights:

  • Physical Feasibility, Not Policy: Loop position is set by execution speed versus human response time; high throughput physically forces oversight out-of-the-loop.

  • Exhaustible Resource: Human attention is finite. Overloaded capacity leads to "insolvent" oversight, failing through silent rubber-stamping.

  • Low-Prevalence Decay & Blind Spots: Accurate AI makes human error detection harder, while supervisor AIs share identical blind spots. The human's true value is statistical verification independence.

  • The Scaling Dichotomy: Real-time monitoring costs $\Omega(T)$ in attention, whereas Constitutive Oversight (ex-ante rule-setting) costs $O(1)$—making it the only sustainable model at scale.

Human on the Loop Working Paper

No. 8

The evolution from generative AI to autonomous AI agents is bringing corporate management to a major turning point. Until now, many companies have operated under a structure in which humans execute work while systems and AI support them—a Human-in-the-Loop model, where humans remain continuously involved in business processes to make decisions and provide approvals.

However, in an era where AI agents can autonomously execute and coordinate activities across sales, marketing, analytics, customer service, software development, operations, and other functions, a management model in which humans individually approve and execute every decision will make it increasingly difficult to achieve both speed and scale.

VURA Capital Innovation defines this fundamental shift as the transition from Human-in-the-Loop to Human-on-the-Loop (HOTL) Management.

Human-on-the-Loop Management is a next-generation management model for the AI era in which AI autonomously executes tasks and operations, while humans define purpose, values, and capital allocation; oversee the system as a whole; and intervene only when necessary—controlling the system from outside the operational loop.

VURA Capital Innovation believes that this transition will become a core element of Enterprise Redefinition in the AI era—the redesign of business models and organizations in response to technological evolution.

Why Human-on-the-Loop Management, and Why Now?

AI agents are now evolving toward Closed-loop AI—systems that autonomously observe the outcomes of their own actions, identify issues, improve their approach, and execute again.

Traditional AI adoption has largely followed a Human-in-the-Loop model, with humans serving as the central hub for input, evaluation, and execution. Today, however, autonomous loops in which multiple AI agents collaborate to complete entire workflows are becoming practical realities.

Examples include AI coding agents autonomously identifying and fixing bugs, AI sales agents continuously optimizing proposals based on outcomes, and AI systems monitoring infrastructure and automatically recovering from failures.

When autonomous cycles operate at this extraordinary speed, a management model that requires humans to remain involved in every decision loop risks turning humans themselves into the organization's greatest bottleneck.

In the AI era, the role of management is no longer primarily to make individual operational decisions. It is to define purpose, values, and capital allocation, and to design and govern the overall system composed of both humans and AI.

Executives must therefore evolve from being individual decision-makers within the operational loop to becoming architects and overseers of the overall system in which AI agents autonomously create value.

Executives Must Move from Inside the Loop to Above the Loop

The Role of Management Is Shifting from Execution to Governance

VURA Capital Innovation views the evolution of industry and management models as follows.

Industrial Age | Human-in-the-Process

Humans as “Executors”

Since the Industrial Revolution, humans have been integral to the process itself: humans execute the work, and humans manage the work.

Digital Age | Human-in-the-Loop

Humans as the “Core of the Workflow”

With digitalization and the early adoption of AI, systems generate recommendations while humans remain responsible for approval and decision-making.

AI Agent Age | Human-on-the-Loop

Humans as “Governors”

AI agents autonomously execute and coordinate work, while humans define purpose, values, and capital allocation, overseeing the system from outside the operational loop and intervening when necessary.

The Evolution of Industry and Management Models

In Human-on-the-Loop Management, the essential roles of humans can be summarized into five areas:

1. Purpose — Defining Purpose
Determining the direction the company should pursue and the future it intends to create.

2. Values — Defining Values
Establishing the ethical principles and corporate values that guide AI behavior—creating the guardrails that prevent undesirable or uncontrolled AI actions.

3. Capital Allocation
Allocating financial, human, technological, and other resources to the areas where they can generate the greatest value and return.

4. Governance — Oversight and Governance
Establishing mechanisms to ensure that both AI and humans operate effectively, responsibly, and safely.

5. Exception Intervention
Making the final human decision when unexpected situations, exceptional circumstances, or significant risks arise.

Human-on-the-Loop: A Theory of Oversight Capacity, Its Failure Modes, and the Non-Delegable Residual in the Age of AI Agents Naoki Kadowaki VURA Capital Innovation Holdings, Inc. VURA Working Paper Series, No. 8 · ISSN 2761-011X · August 2026 · Version 1.3 Abstract. Firms deploying AI agents increasingly describe their governance posture as human-onthe-loop: the system executes autonomously while a human defines purpose, values and capital allocation, exercises governance, and intervenes at exceptions. This paper takes that proposition seriously enough to test it, and finds that it cannot be adopted as written. Three reclassifications follow. First, loop position is not a stance management adopts but an authority configuration whose feasibility is fixed by the ratio of a decision class's consequence latency to the overseer's measured end-to-end response latency; throughput, not policy, determines where a firm sits. Second, oversight is not a role but a consumable resource with a measurable capacity — uncommitted attention, interaction time, context re-acquisition cost, neglect time — and therefore with a solvency condition. Insolvent oversight fails not by refusing work but by accepting it, and its symptom is a rising override and non-response rate. Third, the human's irreplaceable contribution is not superior judgement but verification independence: a loop closed by a verifier statistically dependent on its generator does not converge on correctness, a claim that survives improvements in model capability only where an instrument maintains the overseer's detection probability as prevalence falls. The five roles then separate by scaling law. Constitutive oversight — purpose, admissible means, and the allocation of capital and attention across purposes — costs O(1) in agent throughput; control oversight admits no corresponding bound at a fixed detection guarantee, since under a fixed attention budget the share of consequential errors it removes falls as O(1/T), and strictly faster wherever the supervised system's reliability is improving, while holding that share at any positive constant costs attention growing as Ω(T) and strictly faster still. Routing exceptions selectively relocates that constraint to the recall of the routing rule rather than removing it. Any configuration financed by rising throughput therefore reaches a point beyond which only constitutive oversight remains solvent, which is what makes a reconstructed model implementable. The paper is a conceptual contribution carrying a formal capacity constraint and a measurement proposal; it generates no new data. It builds from the refutation set — human factors, human–computer interaction, agent evaluation, public administration and primary legal records — and closes with a protocol requiring a firm that asserts human oversight to state the quantities that would make the assertion true or false. Keywords: human-on-the-loop; supervisory control; oversight capacity; automation bias; low-prevalence effect; AI agents; real authority; verification independence; duty of oversight; attention allocation; enterprise redefinition JEL classification: D23, M12, M14, M54, L23, K22, O33, D83 Note on this version. Version 1.1 responds to referee comment on version 1.0. Propositions 5 and 12 are restated: Proposition 5 now carries an explicit condition on the origin of the exception stream, and Proposition 12 is stated in terms of oversight efficacy — the share of consequential errors that control oversight removes — after the referee observed that the earlier Ω(T) formulation presupposed a constant per-action error rate. A sensitivity analysis of the parameters transferred from other domains has been added as Section A.6 of Appendix A, and the declaration protocol of Appendix B is now tiered by the cost of each item. Claims resting on primary sources the research pass did not inspect have been narrowed to what the secondary record supports rather than merely flagged. Version 1.2 makes three further changes at referee request: the non-delegable residual of the title is now stated explicitly, in Section 1.3 and in full in Section 12, together with what falls 1 outside it; Section 7.4 extends the correlated-blind-spot argument to adversarially induced correlation, which an attacker can select for and which a design-time independence assessment will not capture; and the notation of the title concept has been checked for consistency throughout. Version 1.3 responds to a third round of referee comment with four changes: Proposition 12 is generalised to non-uniform review, so that the share of consequential errors control oversight removes carries the recall of the rule routing items to the overseer and triage is shown to relocate the constraint to that recall rather than to remove it; Section 4.7 is new and records that the capacity parameters and the detection probability are written as constants although all four drift adversely within a period, so that the static constraint is an optimistic bound; the corporate-law argument of Section 10.5 is bounded expressly to procedural adequacy, the threshold being bad faith rather than negligence and the claim concerning the accuracy of a representation about the control environment rather than any operational failure; and Appendix B now requires a minimum audit trail as a condition of a Tier 1 declaration, together with a residual uniform sample of the unrouted stream in Tier 2. 1. Introduction 1.1 The claim under examination The proposition this paper examines can be stated in one sentence. Artificial intelligence executes the work autonomously, the human defines purpose, values and the allocation of capital, exercises governance over the system as a whole, and intervenes at exceptions. Five roles are reserved to the human being; everything else is delegated to a machine that acts by default rather than on request. The arrangement is offered as a management model a firm may adopt: a stance, chosen at the level of policy and implemented downward through the organisation. The specific formulation examined here is not drawn from the academic literature. It was advanced by the author's own firm, VURA Capital Innovation Holdings, Inc., in a press release dated 29 June 2026 (VURA Capital Innovation Holdings, 2026). It is disclosed here, rather than in a footnote, because it creates an interest the paper must manage rather than assert away: a researcher examining his own firm's public claim has an incentive to vindicate it. The management adopted is structural: the paper is constructed so that the proposition can fail, and each of its thirteen propositions carries an explicit refutation condition. Section 11.7 discloses the interest in the form required of an academic disclosure, states the respects in which the paper's conclusions do not vindicate the firm's formulation, and records the citation discipline applied to the author's own prior working papers. The proposition deserves to be taken seriously, because it occupies the only remaining position between two arrangements that are each, on present evidence, untenable. Full delegation — a closed loop in which the system generates, checks and corrects its own work — is contradicted by the peer-reviewed self-verification literature reviewed in Section 7, which finds that intrinsic selfcorrection not merely fails to help but frequently degrades performance. Per-item human approval is destroyed by throughput: no arrangement in which a human authorises each action survives a system emitting hundreds of state-changing actions per task. Something between the two is not a rhetorical compromise but the only structurally available option, and the proposition correctly, if informally, distinguishes setting an objective from checking an output. It cannot, however, be adopted as written. Stated as five good roles it is a list of intentions containing no quantities, and each is bounded or contradicted by evidence assembled below. "Intervenes at exceptions" collides with the low-prevalence effect and with fourteen years of clinical override data that interface improvement did not move. "Executes autonomously" is unsupported at the level asserted: on the closest available measurements, the best systems complete under a third of the tasks in a simulated office and under a sixth of real, priced freelance projects, as Section 1.4 documents. "Exercises governance" presumes an information channel independent of the system Human-on-the-Loop · WP No. 8 1. Introduction 2 being governed, which most deployed arrangements do not possess. Above all, the proposition asserts a position without asserting the quantities that determine whether the position is occupied at all. Such a declaration is not weak oversight but an allocation of liability with no control content, and Section 9 documents what happens to the person to whom it is allocated. 1.2 Method: building from the refutation set The discipline of this paper is stated explicitly so that it can be checked. The paper does not argue for the five roles. It assembles the strongest evidence available against them — the human-factors literature on vigilance, complacency and automation bias; the human–computer interaction literature on reliance and explanation; the agent-evaluation literature; and the legal record of automated decision failures — and builds a theory out of whatever survives. Where the survivor resembles the original proposition it does so only under conditions that must be measured and declared; where it does not, the paper says so. This is the discipline of two earlier papers in this series (Kadowaki, 2026e, 2026g), cited as methodological precedent and not as evidence. The genre should be equally explicit, because the strength of a claim depends on it. This is a conceptual paper carrying a formal capacity constraint and a measurement proposal. It is not an empirical study: it runs no experiment, conducts no survey, and generates no new data. Every quantity is imported from a named source with its evidence grade attached, and none is re-derived, rounded or extrapolated. What is genuinely its own is analytical: six definitions that make loop position and oversight capacity measurable rather than declarative, thirteen propositions each with a refutation condition, a scaling result, and an audit protocol. One feature of the evidence base deserves naming at the outset, because it shapes every comparison the paper makes: its provenance is asymmetric. Of the strongest data points supporting agent autonomy, most are preprints or reports produced by parties with a direct commercial interest in the result — OpenAI's GDPval (Patwardhan et al., 2025), the Remote Labor Index produced by the Center for AI Safety with Scale AI (Mazeika et al., 2025), Sierra's τ²-bench (Barres et al., 2025), Anthropic's Project Vend (Anthropic, 2025a, 2025b). Of the strongest findings limiting autonomy claims, most are peer-reviewed — TheAgentCompany (Xu et al., 2025), τ-bench (Yao et al., 2024), MAST (Cemri et al., 2025), Huang et al. (2024), Stechly et al. (2025) — or are primary legal and regulatory records. That asymmetry cuts both ways, and the paper does not use it as a rhetorical instrument. Peer review is slower than the technology it evaluates, so a refereed result may describe a system generation no longer deployed while an interested preprint describes the system a firm is buying; and interested parties are often the only ones with access to frontier deployments, so excluding their reports would produce not a neutral evidence base but a thinner and older one. Nor are publication incentives uniformly optimistic: the largest meta-analysis of human–AI combination detected publication bias for its optimistic augmentation estimate but not for its pessimistic synergy estimate (Vaccaro, Almaatouq, & Malone, 2024). Section 11 returns to this as a live threat to the paper's conclusions. 1.3 Four contributions The first contribution is to reclassify loop position from a policy to a measured property. Definition 1 makes the classification a function of two latencies: the interval between a system's commitment to an action and the realisation of its consequence, and the human's end-to-end response latency. A configuration is on-the-loop only if the second is smaller than the first. Proposition 1 states that, holding human latency fixed, any increase in throughput or consequence speed shifts the achievable position monotonically toward out-of-the-loop irrespective of the policy asserted; Human-on-the-Loop · WP No. 8 1. Introduction 3 Proposition 2 states that where the human's information is strictly poorer and the decision must be taken inside the system's window, formal authority is exercised as ratification and real authority has already transferred (Aghion & Tirole, 1997). A firm does not choose to be on the loop. Throughput puts it there, or past it. The second contribution is to reclassify oversight from a role to a consumable resource with a solvency condition. Definitions 3 and 4 give oversight a capacity in terms of uncommitted attention, the mean interaction time to handle an exception to closure, the context re-acquisition cost of returning to displaced work, and the neglect time a supervised process tolerates unattended; solvency requires that the arrival-rate-weighted cost of exceptions not exceed uncommitted attention, and that the number of supervised processes respect a span bound. Propositions 3 to 5 draw the consequences: an insolvent exception channel fails not by refusing exceptions but by accepting them; adding an oversight duty without removing work reallocates capacity rather than adding any; and effective supervisory span, where the exception stream is generated by the supervised processes themselves and where it has been measured at all, is a single-digit number that tightens further where part of the stream arrives from outside those processes, and whose violation appears first as loss quality rather than as loss of throughput. The third contribution is the paper's deepest claim, and deliberately not the claim usually made on the human's behalf: the argument is not that the human judges better. Definition 5 defines verification independence statistically, as the condition that the probability a verifier misses an error is unchanged by the fact that the generator committed it. Propositions 10 and 11 follow: a loop closed by a non-independent verifier has an error rate bounded below by the shared error rate and can move away from correctness under iteration, and the marginal value of a human overseer is the product of the overseer's decorrelation from the system and their probability of detecting an error of the relevant class. This is the strongest available justification for the human's presence, and it survives improvements in model capability only where an instrument maintains the overseer's detection probability as prevalence falls: capability leaves decorrelation intact and erodes detection through Proposition 7. Independence, not judgement, is the scarce input, and it is not selfmaintaining. The fourth contribution is what makes the reconstructed model implementable rather than aspirational. Definition 6 separates constitutive oversight, which fixes the objective function of the automated system, from control oversight, which monitors executions and intercepts exceptions. Proposition 12 states the scaling result: the attention cost of the first is invariant in throughput, while the second admits no corresponding invariance at a fixed detection guarantee. Reviewing a constant sample bounds the cost of control oversight without bounding its failure, and the quantity that measures the failure is neither coverage nor absolute exposure but oversight efficacy η(T), the share of consequential errors that control oversight removes in a period at throughput T, which is the product of the coverage a fixed attention budget buys and the probability that an examined error is detected at the prevailing prevalence. Under a fixed attention budget η(T) is O(1/T) whatever the per-action error rate does, and by Proposition 7 falls strictly faster than that wherever the error rate is itself falling, since lower prevalence lowers the overseer's detection probability; holding η at any positive constant requires attention growing as Ω(T), and faster than linearly once that falling prevalence is accounted for. Reliability improvement therefore does not relieve the constraint but accelerates it. The result is stated for any review design and not merely for uniform review, which is the general model operated with a router that discriminates nothing: a firm that instead routes selectively raises the prevalence of the stream its overseers examine, and by Proposition 7 their detection probability with it, but cannot remove more consequential errors than its routing rule places in front of a human. Triage accordingly relocates the constraint from the overseer's attention Human-on-the-Loop · WP No. 8 1. Introduction 4 to the recall of the router rather than lifting it, and relocates it into a quantity the routed stream cannot reveal, since everything the firm observes of the rule is conditioned on having been routed. Any configuration financed by increases in throughput reaches a throughput beyond which control oversight can be afforded only by accepting an efficacy that is negligible and, in current practice, undeclared, and only constitutive oversight survives. The five roles therefore do not stand or fall together. Purpose, values and capital allocation are retainable at scale; contemporaneous monitoring and exception interception are not. The management question is not whether to sit on the loop but what can be placed in the O(1) class. Taken together the four contributions identify what the title calls the non-delegable residual: the part of the human's role that survives both the capacity constraint and any foreseeable improvement in the systems being supervised. It has three components resting on three different grounds — verification independence, which is structural; constitutive oversight, which is logical; and fiduciary accountability, which is doctrinal — and it is materially smaller than the five roles suppose. Section 12 states it in full, and states with equal precision what is not in it. 1.4 Scope conditions The theory is restrictive and its boundaries should be stated exactly. It concerns firms deploying automated decision systems that act by default — systems whose normal operation produces consequential actions without a human act as a precondition. It is silent on systems requiring affirmative human initiation, which are in-the-loop by construction under Definition 1 and whose problem is of a different kind: there the question is the quality of a decision certainly being made, not whether a reserved authority is exercised at all. The most serious scope limitation is evidentiary. The quantitative evidence on intervention latency, supervisory span and exception-channel saturation on which Sections 3 to 5 depend is drawn overwhelmingly from driving simulators and instrumented vehicles, teleoperated robotics, and clinical alerting. None is knowledge work. The transfer to a manager supervising an AI agent is a mechanism argument, not a measurement: attention is finite, interrupted work carries a reacquisition cost, and detection probability falls with prevalence, and these mechanisms are domaingeneral even where their constants are not. The constants almost certainly are not. Section 11 treats this as the paper's principal weakness, and the measurement proposals of Section 10 exist in part because the numbers a firm needs are not in the literature and must be generated locally. An honest statement of what AI agents currently do and do not do is owed at the outset, since the proposition asserts autonomous execution. On TheAgentCompany, a benchmark of office tasks in a simulated software company, the best system completed 30.3 per cent of tasks fully and scored 39.3 per cent with partial credit (Xu et al., 2025; peer-reviewed, NeurIPS 2025 Datasets and Benchmarks Track); the authors note that the tasks are "generally on the more straightforward side" because they must be automatically checkable, that no human baseline was collected, and that only two agent frameworks were tested. On OSWorld 2.0, a set of long-horizon computer-use workflows, the best agent reached 20.6 per cent binary completion and 54.8 per cent partial credit (Yuan et al., 2026; preprint, not peer-reviewed), with failure attributed not to poor graphical control but to agents losing track of constraints, missing information that arrives mid-task, guessing rather than asking, and skipping verification. On τ-bench, which scores customer-service interactions against the final state of the system of record rather than against what the agent said, state-of-the-art function-calling agents succeeded on fewer than 50 per cent of tasks and achieved pass^8 below 25 per cent in retail (Yao et al., 2024; peer-reviewed, ICLR 2025). Software engineering, where autonomy claims are strongest, shows a second problem. Under the official standardised bash-only harness in which every model is evaluated identically, the leading Human-on-the-Loop · WP No. 8 1. Introduction 5 system resolved 76.8 per cent of SWE-bench Verified instances in the February 2026 re-run (leaderboard artifact, reported in Willison, 2026), while vendor-scaffolded figures of approximately 95 per cent circulate which this paper's evidence base could not verify against any primary vendor source and which are therefore stated here as unverified. The roughly nineteen-point gap is itself the finding: measured agent capability is a property of the model plus a human-engineered scaffold. On real, previously commissioned and priced freelance projects, the Remote Labor Index recorded an automation rate of 2.5 per cent for the best agent in October 2025, rising to 15.8 per cent by July 2026 (Mazeika et al., 2025, preprint produced jointly by interested organisations; the 2026 figure from Center for AI Safety, 2026, a company blog post). GDPval, the most favourable serious measurement available, found a best win-or-tie rate of 47.6 per cent against human expert deliverables on one-shot, precisely specified, non-interpersonal tasks, with a grading instrument whose human inter-rater agreement was 71 per cent and whose automated grader agreed with humans only 66 per cent of the time (Patwardhan et al., 2025; preprint, produced by OpenAI). METR places the frontier at approximately 320 minutes of human-equivalent task length at a 50 per cent success rate, with a 95 per cent confidence interval of 170 to 729 minutes and a doubling time of about 89 days since 2024 (Kwa et al., 2025, preprint, as revised in METR, 2026). The conclusion the evidence supports is narrower than either camp usually states. Capability is improving fast, and the movements are large enough to survive the version and harness changes their own sources document: the Remote Labor Index moved from 2.5 to 15.8 per cent in nine months, and the time-horizon doubling since 2024 is now put at about three months. Neither, however, is a clean like-for-like series, and the qualifications belong with the figures rather than in a footnote. METR's revision expanded the task suite from 170 to 228 tasks and migrated the evaluation infrastructure from its in-house Vivaria to the open-source Inspect harness, with two previously evaluated models scoring statistically higher under Vivaria than under its replacement; only 5 of the 31 long tasks carry measured rather than estimated human baselines; and the doubling time since 2024 itself moved from 109 to 89 days between the two versions. The Remote Labor Index endpoints are different agents reported in different publications — 2.5 per cent for Manus in an October 2025 preprint, 15.8 per cent for a different model in a July 2026 company blog post. But capability is not autonomy, and the headline metric makes the point when read correctly. A 50 per cent success rate is a coin flip. Read correctly, the time-horizon metric is an argument for on-the-loop supervision, not against it — and the question is not whether such supervision is desirable but whether, at the throughputs that make agent deployment economic, it is arithmetically available. Table 6 sets each component of the autonomous-execution claim against the best measurement available for it, with the grade of that measurement attached. Claim Best measured evidence Grade Reading General office work TheAgentCompany: 30.3% full completion, 39.3% partial (best model), 175 tasks, no human baseline A Fewer than one task in three completed without help, on tasks the benchmark's authors call "more straightforward" and against no human baseline. Computerbased operations OSWorld 2.0: 20.6% binary completion on 108 long-horizon workflows B About one long-horizon workflow in five completed end to end; failures are of constraint tracking and verification, not of graphical control. Customer service τ-bench: under 50% pass@1; pass^8 under 25% in retail. τ²-bench: dual control costs 18–25 points A / B Reliability, not capability, is the binding constraint. Human-on-the-Loop · WP No. 8 1. Introduction 6 Claim Best measured evidence Grade Reading Software development SWE-bench Verified 76.8% under a standardised harness; SWE-bench Pro ~43.7%; vendor-scaffolded reports near 95% not independently verified F / B / C Supported for narrow, oracle-verified bug fixing only. Real paid delivered work Remote Labor Index: 2.5% (Oct 2025) rising to 15.8% (Jul 2026), humangraded, 240 priced projects B / C Improving fast; the level does not support autonomous execution. Expert-quality deliverables GDPval: 47.6% win-or-tie against human experts; grader self-agreement only 66– 71% B / C Closest to supported; the source's own economics assume human oversight. Closed-loop selfimprovement Huang et al.: every self-correction cell flat or worse. Stechly et al.: self-critique 5%→3% and 16%→2%; a sound external verifier reaches 38% and 37% A Contradicted by the peer-reviewed literature. Table 6. Claims of autonomous execution against measured evidence, with evidence grade. Grades: A peer-reviewed; B preprint; C vendor or company report; D grey literature; E primary legal record; F leaderboard. The paper proceeds as follows. Section 2 establishes the provenance of the phrase and the technical concept beneath it, and gives the two definitions on which the rest depends. Section 3 shows that loop position is determined by throughput rather than chosen (Propositions 1–2); Section 4 treats oversight as a consumable resource with a solvency constraint (3–5); Section 5 explains why more reliable automation is harder to supervise (6–7); Section 6 asks when oversight demonstrably works (8–9); Section 7 establishes that a loop cannot close itself (10–11); Section 8 states the scaling dichotomy and reconstructs the five roles (12); Section 9 analyses nominal oversight as liability allocation against three documented failures (13). Section 10 sets out the legal position and what a firm should measure, Section 11 the limitations, falsification conditions, competing explanations and interests, and Section 12 concludes. Appendices A and B give the solvency calculation and a declaration protocol. 2. What Human-on-the-Loop Is: Provenance and Definition 2.1 The phrase is doctrinal, the concept is technical A management literature that adopts a phrase without examining where it came from inherits commitments it has not inspected. The phrase "human on the loop" is not a term of art from human factors. It appears as no defined construct in the supervisory-control literature, has no canonical peer-reviewed definition, and organises no experimental paradigm. What does exist in that literature is the technical content the phrase gestures at, and that content is both older and more precise than the phrase. The technical anchor is the scale of levels of automation introduced by Sheridan and Verplank (1978) in a report on human and computer control of undersea teleoperators. Level 6 of that scale describes a configuration in which the "computer allows the human a restricted time to veto before automatic execution." That is the exact operationalisation of what doctrine later called being on the loop: the machine will act, the human may stop it, and the human's window is bounded. Level 7, in which the computer executes automatically and then necessarily informs the human, marks the boundary beyond which the arrangement is no longer a veto but a notification. Two things about that anchor should be stated plainly. First, the 1978 document is a Massachusetts Institute of Technology Man–Machine Systems Laboratory technical report distributed through the Defense Technical Information Service, catalogued at OCLC 8544670. It is grey literature. It was not Human-on-the-Loop · WP No. 8 2. What Human-on-the-Loop Is: Provenance and Definition 7 peer-reviewed, it carries no DOI, and the most-cited scale in the levels-of-automation tradition therefore originates outside the refereed record. Second, this paper does not quote the 1978 table from the original, which could not be retrieved directly; it cites the scale as reproduced in the peerreviewed article of Parasuraman, Sheridan and Wickens (2000), which attributes it openly to Sheridan and Verplank and which is the standard citation route in the field. The ten levels as they circulate in the secondary literature should be treated as verified-as-conventionally-reproduced rather than as verified against the 1978 original, and this paper uses only level 6, whose wording is stable across every reproduction consulted. The phrase itself enters through defence doctrine. The earliest authoritative primary use located is the United States Air Force Unmanned Aircraft Systems Flight Plan 2009–2047 of 18 May 2009, which states at page 15 that "Agile, redundant, interoperable and robust command and control (C2) creates the capability of supervisory control ('man on the loop') of UAS." The document is a service doctrinal publication, not a refereed work, but it is a primary record and it is decisive on one point: the Air Force itself glosses "man on the loop" as supervisory control, explicitly importing Sheridan's concept into doctrine. That gloss is the bridge between the two literatures, and it is documented rather than inferred. Two cautions attach to that claim. The first is temporal: it should be stated as "at least as early as 2009" and not as "first used in 2009," because the research underlying this paper did not examine the 2007 and 2011 Unmanned Systems Integrated Roadmaps, in which an earlier use may well be found. The second is evidentiary. One verification pass retrieved the Flight Plan in full text from the Defense Technical Information Center under accession ADA505168 and confirmed the sentence at page 15; a second pass, working from a different accession number, could not retrieve it, and the phrase is absent from the executive-summary text mirrored elsewhere online. The claim in this paper therefore rests on the successful retrieval. A reader who cannot reproduce it should treat the 2009 attribution as unconfirmed; the argument of Sections 3 to 12 depends on measured quantities rather than on etymology and is unaffected either way. 2.2 The taxonomy was fixed by Human Rights Watch, and the US Department of Defense declined it The three-way taxonomy that management writing now takes for granted was fixed not by an engineering body but by an advocacy report. In Losing Humanity: The Case Against Killer Robots, published on 19 November 2012 by Human Rights Watch and the Harvard Law School International Human Rights Clinic, the definitions are given as follows: "Human-in-the-Loop Weapons: Robots that can select targets and deliver force only with a human command"; "Human-on-the-Loop Weapons: Robots that can select targets and deliver force under the oversight of a human operator who can override the robots' actions"; and "Human-out-of-the-Loop Weapons: Robots that are capable of selecting targets and delivering force without any human input or interaction." The report gives no prior attribution for the taxonomy. Its evidentiary grade is grey literature produced by a party to an advocacy campaign; that does not make the definitions wrong, but a management paper citing them is citing a campaign document and should say so. The sharper finding is what happened next in the institution that would have been expected to adopt the vocabulary. United States Department of Defense Directive 3000.09, on autonomy in weapon systems, as reissued on 25 January 2023, does not use the loop vocabulary at all. Its operative standard is that such systems be designed to allow commanders and operators to exercise "appropriate levels of human judgment over the use of force," and its defined category for a system that acts under supervision is the "operator-supervised autonomous weapon system." The Directive Human-on-the-Loop · WP No. 8 2. What Human-on-the-Loop Is: Provenance and Definition 8 is a primary policy instrument and its silence is deliberate: the department that operates the systems the taxonomy was invented to describe declined the taxonomy. The change of term between the 2012 issuance and the 2023 reissuance is worth a sentence of its own. The earlier text spoke of a human-supervised autonomous weapon system; the later text speaks of an operator-supervised one. That is a tightening from the species to the designated role. It abandons the implication that the presence of any human being somewhere in the organisation constitutes supervision, and locates supervision in a specified position with training, authority and a station. Definition 2 below generalises exactly that move to the firm. The multilateral record points the same way. The Guiding Principles affirmed by the Group of Governmental Experts on Emerging Technologies in the Area of Lethal Autonomous Weapons Systems in 2019, annexed to the Group's report as CCW/GGE.1/2019/3, Annex IV, state at principle (b) that "accountability cannot be transferred to machines." The concern expressed across these instruments is not the presence of a human at a console at a particular instant. It is the retention of responsibility across a system's life cycle — design, testing, deployment and use — which is a very different requirement from the ability to press a button inside a window. Two conclusions follow for a management paper. First, it should not claim doctrinal support for the phrase "human on the loop." The phrase's authority is borrowed from a body of law and policy that either declines to use it or uses it as informal shorthand. Second, and more usefully, the doctrinal concern that does have authority — retention of responsibility across a life cycle rather than intervention at a moment — is a stronger and more demanding standard than the one the management version imports, and Sections 9 and 10 build on it. Table 1 sets the record out in order, separating what each source contributes to the technical concept from what it contributes to the phrase, and attaching to each the evidentiary status a management paper ought to carry with it. Source Date What it contributes Status Sheridan & Verplank, Human and Computer Control of Undersea Teleoperators, MIT Man–Machine Systems Laboratory 1978 Levels of automation; level 6 is "computer allows the human a restricted time to veto before automatic execution" — the technical content of the arrangement Technical report; not peer-reviewed. Cited through the route stated in Section 2.1. USAF, Unmanned Aircraft Systems Flight Plan 2009–2047, p. 15 18 May 2009 Glosses "man on the loop" as supervisory control, welding the doctrinal phrase to Sheridan's concept Primary document (DTIC ADA505168); retrieved in one verification pass, not in a second. Human Rights Watch & Harvard Law School IHRC, Losing Humanity 19 Nov 2012 Fixes the in-/on-/out-of-the-loop taxonomy; on-the-loop = force delivered "under the oversight of a human operator who can override the robots' actions" NGO report; gives no prior attribution for the taxonomy. US DoD Directive 3000.09 (reissuance) 25 Jan 2023 Declines the loop vocabulary; uses "appropriate levels of human judgment over the use of force" and "operator-supervised autonomous weapon system" Primary policy document; the phrase does not appear. CCW GGE on LAWS, Guiding Principles, CCW/GGE.1/2019/3 Annex IV 2019 "Accountability cannot be transferred to machines"; frames the duty as retention across a life cycle Non-binding international guidance. Table 1. Provenance of "human-on-the-loop": the technical concept and the phrase. Human-on-the-Loop · WP No. 8 2. What Human-on-the-Loop Is: Provenance and Definition 9 2.3 What human-on-the-loop means precisely, and the common mislabelling The following definitions convert the doctrinal phrase into measurable properties of a configuration. Definition 1 (Loop position). For a decision class d executed by an automated system, let τsys(d) be the interval between the system's commitment to an action and the realisation of that action's consequence, and let τH (d) be the human's end-to-end response latency — detection, orientation, decision and act — conditional on the human being available. The configuration is in-the-loop if the human's act is a precondition of the system's act; on-the-loop if the system acts by default and the human holds an effective veto or abort, which requires τH (d) < τsys(d); and out-of-the-loop if τH (d) ≥ τsys(d). Loop position is therefore a measured property of a configuration, not a declared policy. Definition 2 (Nominal oversight). An oversight arrangement is nominal when a human is designated as accountable for a system's outputs while at least one of the following fails: (i) they hold an information channel independent of the system's own inputs; (ii) their response latency is below the system's consequence latency; (iii) they possess unencumbered authority to disregard, reverse or halt; (iv) their uncommitted attention exceeds the arrival-rate-weighted cost of the exceptions routed to them. Nominal oversight is not weak oversight. It is an allocation of liability with no control content. The distinction these definitions enforce is load-bearing and is routinely lost. To be on the loop is for the system to act by default while a human holds an effective veto or abort. The default matters: absent human action, the machine proceeds. This is level 6 on the Sheridan–Verplank scale, and the word "effective" is doing work, because a veto that cannot be exercised inside τsys is not a veto but a formality. An arrangement in which a human reviews an output before it is sent is not this. It is in-the-loop under Definition 1, because the human's act is a precondition of the system's act; in the levels-ofautomation vocabulary it is level 5, where the computer executes the suggestion if the human approves. It is a different authority configuration, with a different failure mode, and it is described as on-the-loop in business writing with striking regularity. Worse, arrangements materially weaker than either are given the same label: an approval queue with a throughput target, in which the reviewer's realistic options are to approve or to fall behind, is neither a veto nor a decision, and it is nonetheless reported as human oversight. The mislabelling matters for a reason that is governance rather than terminology. It allows a firm to claim the governance posture of veto authority — a human who could have stopped this — while operating an approval queue whose observable behaviour, as Section 3 documents from the takeover-latency and override literatures, converges on ratification. The claim and the arrangement come apart, and they come apart in the direction that transfers real authority to the system while leaving formal accountability with the person. Definition 2 exists to name that state without euphemism. An arrangement failing any one of its four conditions is not a weaker version of oversight on a continuum; it is a different object, whose function is the allocation of liability, and Section 9 shows what that function does to the people it is allocated to. Human-on-the-Loop · WP No. 8 2. What Human-on-the-Loop Is: Provenance and Definition 10 2.4 Position of this paper within the series Figure 1. The three-layer architecture of the VURA Working Paper Series, with the present paper (No. 8) highlighted in Layer 3, enterprise management, alongside โ‘  Future Value Theory, โ‘ก Enterprise Redefinition, โ‘ข Enterprise Redefinition Observed, โ‘ฃ From Job Description to Purpose Description and โ‘ค Brain Capital Management; Layer 2 comprises โ‘ฅ The Self-Defining Society and Layer 1 โ‘ฆ Redefinition of Capitalism. Placement is architectural and is not a claim of logical dependence. This paper is the eighth in the series and sits in Layer 3, which treats enterprise management. Its relation to the earlier papers is best stated using the citation-classification discipline this series applies to its own work: every self-citation is classified explicitly as an axiomatic foundation, as observational evidence, or as a design proposal, and design proposals are never counted as support in any confidence rating attached to a claim. The discipline exists because a series that cites itself accumulates apparent corroboration without accumulating evidence, and the reader is entitled to know which of the two is on offer. Applying it here: Future Value Theory (Kadowaki, 2026a) and Enterprise Redefinition (Kadowaki, 2026b) supply the axioms on which the present treatment of purpose and capital allocation rests, and are cited as axiomatic foundations — that is, as premises adopted, not as findings established. Enterprise Redefinition Observed (Kadowaki, 2026c) supplies the observation-asymmetry finding, and is cited as observational evidence, the only one of the three categories that carries evidentiary weight in this paper. From Job Description to Purpose Description (Kadowaki, 2026d) supplies the role-design apparatus that Section 8 uses to reconstruct the five roles, and Brain Capital Management (Kadowaki, 2026e) supplies both the cognitive-capacity constraint that Section 4 formalises as oversight capacity and the three-layer disclosure architecture that Section 10 and Appendix B adapt into a declaration protocol. Both are cited as design proposals, and neither is counted as support for any empirical claim made here. Where Section 4's capacity argument needs evidentiary warrant it takes it from the peer-reviewed robotics and clinical-alerting literatures, not from the earlier paper that proposed the construct. Layer 1 — Era structure (macro): Redefinition Capitalism โ‘ฆ Redefinition Capitalism Layer 2 — Social structure (meso): Self-Defined Society โ‘ฅ Self-Defined Society Layer 3 — Enterprise management (micro) โ‘  Future Value Theory โ‘ก Enterprise Redefinition โ‘ข Enterprise Redefinition Observed โ‘ฃ From Job Description to Purpose Description โ‘ค Brain Capital Management โ‘ง Human-on-the-Loop (this paper) Human-on-the-Loop · WP No. 8 2. What Human-on-the-Loop Is: Provenance and Definition 11 2.5 The ironies, at their real evidentiary weight No paper on supervisory control can avoid Bainbridge (1983), and most that cite it overstate what it is. "Ironies of automation" is a five-page Brief Paper in Automatica 19(6), 775–779. It is an analytical essay. It has no sample, no design, no data and no statistics, and its very considerable authority is conceptual rather than empirical, deriving from the acuity of its argument and from several decades of citation. 

Presenting it as an empirical finding — "studies show that operators cannot monitor" — overstates the field's evidence base. This paper cites it as what it is: a theoretical statement of a problem, whose empirical content must be supplied from elsewhere. The quotations below were verified against a full-text mirror rather than the publisher's own file, and should be rechecked against the Automatica PDF before they are set in print. The irony most often quoted from the paper is the monitoring claim: that "it is impossible for even a highly motivated human being to maintain effective visual attention towards a source of information on which very little happens, for more than about half an hour." Two corrections are needed. The "about half an hour" is Bainbridge's summary of the then-existing vigilance literature in the Mackworth tradition, not a finding of her own, and attributing it to her as a result is a citation error. The best-sourced modern figure, and the one Section 4 uses, is from a peer-reviewed review by the field's leading authorities: detection accuracy in vigilance tasks declines by about 10 to 15 per cent after about 30 minutes, with most of that decline occurring within the first 15 (Warm, Parasuraman, & Matthews, 2008). The irony that actually transfers to AI-era management is a different one, and it is quoted far less often. Bainbridge observes that "if the decisions can be fully specified then a computer can make them more quickly, taking into account more dimensions and using more accurately specified criteria than a human operator can. There is therefore no way in which the human operator can check in real-time that the computer is following its rules correctly." The corollary she draws is the sentence a firm deploying agents should read twice: if the human is nonetheless required to follow the details, then "it is necessary for the computer to make these decisions using methods and criteria, and at a rate, which the operator can follow." That is not a complaint about human frailty. It is a statement that contemporaneous verification imposes a design constraint on the machine — on its method, its criteria and above all its rate — and that a system optimised for throughput has already spent the budget that verification would have required. Definition 1 is a formalisation of that observation, and Proposition 12's scaling dichotomy is its consequence. The field itself is not settled, and a paper that presents it as settled is not citation-grade. Endsley (2017), in a peer-reviewed review, names the "automation conundrum" — that as autonomy is added and reliability and robustness increase, operators' situation awareness falls and their ability to take over manual control declines — but treats it explicitly as a design problem, and devotes the article to interface features, levels of automation, adaptive automation and granularity of control as remedies. The most prominent critic of full autonomy in the field does not regard the difficulty as an inevitability. The strongest quantitative summary available is a peer-reviewed meta-analysis of 18 experiments (Onnasch, Wickens, Li, & Manzey, 2014), which finds clear benefits of a higher degree of automation for routine performance, weaker but present workload reduction in normal operation, and negative effects on failure performance and on situation awareness, with a critical performance boundary between automation that supports information analysis and automation that supports action selection: crossing from the former to the latter is where the costs appear. That boundary matters directly for the argument of Section 8, since the reconstructed model proposed there is in substance a proposal to place human effort on the analysis side of it. The meta-analysis reports its results by Human-on-the-Loop · WP No. 8 2. What Human-on-the-Loop Is: Provenance and Definition 12 significance rather than in standardised units, and no effect sizes are extractable from the accessible text; none is quoted here, and no reader should accept one quoted elsewhere without checking the source. Finally, the framework that supplies this paper's technical anchor is contested from within its own field. Wickens (2018), in a twenty-year retrospective on the stages-and-levels model he co-authored, responds to Kaber's concerns about the model's utility in design with "both simplifications and elaborations," and Jamieson and Skraaning (2018) argue that the levels-of-automation approach may not generalise to complex real work settings. The dispute is live and this paper does not resolve it. The position taken here is therefore narrower than the pessimistic reading and does not depend on it. Nothing in the argument requires that humans cannot monitor, that automation inevitably degrades its supervisors, or that Bainbridge's essay was empirically right. What the argument requires is quantities: response latencies against consequence latencies, exception arrival rates against uncommitted attention, detection rates against prevalence, and error correlation between a generator and its verifier. Those quantities are measurable, several of them have been measured, and where they have not been measured in the setting that matters, this paper says so and proposes how a firm might measure them itself. 3. Position Is Determined, Not Chosen 3.1 The argument in one move The proposition this section defends is simple to state and consequential enough to reorganise the rest of the paper. A firm does not choose its loop position. The ratio of the system's consequence latency to the human's response latency chooses it, and the firm's governance policy is at most a statement of intent about a variable it does not control. Definition 1, boxed in Section 2, fixes the terms. For a decision class d, τsys(d) is the interval between the system's commitment to an action and the realisation of its consequence, and τH (d) is the human's end-to-end response latency — detection, orientation, decision and act — conditional on availability. The configuration is on-the-loop only where τH < τsys, and out-of-the-loop where τH ≥ τ sys. Nothing in that definition refers to a policy document, an org chart or a role description. It refers to two measurable durations. Sections 3.2 and 3.3 assemble what is known about τH , including the component oversight designs routinely omit; Section 3.4 supplies τsys; Section 3.5 states the two propositions. One caution on sources. The driving literature is not offered as a model of corporate decision-making; task, stakes and sensory channel all differ. It is offered because it is the only literature in which human intervention latency under automation has been measured at scale, and its findings are used here for structural properties of the human response — magnitude relative to machine cycle times, variance, response to slack, and the separability of latency from quality — not as a transferable numeric constant. 3.2 How long human intervention actually takes The best available aggregate is Zhang, de Winter, Varotto, Happee and Martens (2019), a peerreviewed meta-analysis in Transportation Research Part F covering 129 studies and yielding 520 mean or median take-over time observations. The average mean take-over time was 2.72 s (SD = 1.45, n = 520), and mean take-over time across studies and conditions ranged from 0.69 s to 19.79 s: a central tendency of a few seconds, with a between-condition spread covering a factor of nearly Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 13 thirty. Three further findings matter more than that mean, and each is counter-intuitive in a way that damages a common oversight design. The first is the time-budget effect. A strong effect of time budget was found, with a higher mean take-over time — an average difference of 1.35 s — for a large time budget compared with a small one. More headroom did not produce a faster intervention; it produced a slower one. Humans consume the budget they are given. This is meta-analytic evidence against any oversight servicelevel agreement premised on engineered slack being converted by the human into a margin of safety. It is converted into elapsed time. The second is the mean–standard-deviation relationship. The correlation between the mean and the standard deviation of take-over times was strong, r = 0.82 (Spearman's rank correlation 0.73, n = 397). Two qualifications belong in the same breath as the figure, since the paper elsewhere insists on them. This is a between-study relationship among study-level means and standard deviations, so inferring the variance structure of any single deployment from it is an ecological inference; and a positive mean–standard-deviation relationship is close to mechanical in bounded, right-skewed distributions, of which response times are the standard case. The conclusion does not need the correlation to carry it. Mean take-over times ranged across conditions from 0.69 s to 19.79 s, and the authors state, as the next paragraph records, that collision risk is set by the tail rather than by the mean. One cannot design an on-the-loop configuration to a central tendency. The third is the authors' own limitation, which should be reported in substance rather than paraphrased away. Zhang et al. state that collision risk is not determined by the mean take-over time, but by outliers in the take-over time distribution: the meta-analysis measures the wrong moment of the distribution for safety purposes, and says so. They further note that take-over quality was not examined at all. And nearly all included studies are driving-simulator studies, so the behavioural validity of the corpus is an open question the authors flag rather than resolve. A reader who accepts that the risk is set by the tail will want to know where the tail is, and the honest answer is that this literature does not say. Zhang and colleagues report means, standard deviations and a range; they report no percentiles, and none is quoted here. What the reported moments do license is arithmetic, and it is offered as arithmetic and not as a finding. Treating the reported average mean of 2.72 s and standard deviation of 1.45 s as the parameters of a normal distribution places the ninety-fifth percentile at 2.72 + 1.645 × 1.45 ≈ 5.1 s and the ninety-ninth at approximately 6.1 s. That calculation assumes symmetry, which the same paper's own numbers contradict: a maximum of 19.79 s against a mean of 2.72 s, together with the strong positive relationship between means and standard deviations recorded above, indicates a right-skewed distribution, for which a normal approximation understates the upper tail. Fitting instead a lognormal distribution to the identical mean and standard deviation — the standard shape for bounded, right-skewed response times — moves the ninety-fifth percentile to approximately 5.5 s and the ninety-ninth to approximately 7.7 s. The normal figures are therefore floors under their own assumption, and the gap between the two shapes widens the further out one goes. Two limits must be stated with the exercise. Both sets of figures are consequences of an assumed distributional shape rather than observations, so neither is a percentile of anything. And neither is a percentile of individual take-over times in any event: the reported standard deviation is the dispersion of studylevel and condition-level means, so the arithmetic describes the spread of averages and not the tail of the human responses a design must actually survive, which will be wider still. The designrelevant quantity is thus a percentile the literature does not report, and a firm that sets an oversight response time from a published mean is specifying against the wrong statistic. That is precisely why Appendix A instructs a firm to report its own interaction time at the median and the ninetieth percentile rather than at the mean. Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 14 The simulator caveat makes the real-road datapoint disproportionately valuable. Gershon, Mehler and Reimer (2023), peer-reviewed in Traffic Injury Prevention, instrumented 14 drivers over 5,514 miles of hands-free Super Cruise and captured 265 system-initiated disengagements. Take-over time was 2.4 s hands-free and 1.8 s with a hand already on the wheel — consistent with the meta-analysis and, on their own, reassuring. The decisive number is different: the time to begin recovering normal on-road glance behaviour was 6.1 s, and glance transitions increased by 43 per cent. This licenses the distinction on which the rest of the paper depends. Grabbing control is not being in control. "Intervention initiated" and "supervisory competence restored" are different events, separated in the only deployed measurement available by a factor of roughly 2.5. Any τH defined as the moment the human's hand reaches the control surface prices the attentional handover at zero — and it is the attentional handover that determines whether the intervention is any good. Whether the intervention is any good turns out to be almost unrelated to how quickly it began. Louw, Markkula, Boer, Madigan, Carsten and Merat (2017), peer-reviewed in Accident Analysis & Prevention with N = 75 in a between-participants simulator design, manipulated visual information about the road and the automation state to induce out-of-the-loop states. Reducing visual information slowed take-over times but did not affect collision-avoidance timing, and long take-over duration did not predict collisions. The authors' summary is that take-over time and the timing and quality of the avoidance response appear to be largely independent. A latency service-level agreement can therefore be fully satisfied while intervention quality is unchanged: a firm reporting median response times as evidence of effective oversight is reporting a quantity the evidence says is orthogonal to what it is offered as evidence of. The fastest human intervention figure in the verified record comes from Dixit, Chand and Nair (2016), peer-reviewed in PLOS ONE, using California DMV autonomous-vehicle disengagement reports covering 460,097 autonomous miles: a pooled mean reaction time of 0.83 s. All three caveats belong in the same breath as the figure. The measurement excludes perception time, so 0.83 s is a motor response and not a detect–decide–act latency; only two companies reported exact measurements, the others giving upper bounds only; and these were trained, paid, professional safety drivers whose sole job was to watch, with no competing task — the most favourable possible case. The figure is a floor, not a representative value, and it should never be a design target for an overseer who has other work. The same paper supplies a finding more interesting than its headline. Reaction time increased with cumulative autonomous miles: Mercedes-Benz r = 0.122, p = 0.007; Google r = 0.161, p = 0.029. The authors interpret this as trust degrading response speed with exposure. Even in the population selected and paid to be maximally vigilant, τH was not a constant of the human but a function of operating history, and it moved in the direction that erodes the on-the-loop condition. A firm that measures τH once at deployment is measuring the most favourable moment in the system's life. 3.3 The knowledge-work overhead nobody counts: interruption and resumption In corporate deployments the overseer is not a safety driver whose only job is to watch, but a person doing other work who is pulled out of it by an exception. τH therefore contains a component the driving literature does not isolate: the cost of leaving one task and returning to it. The controlled measurements come from Monk, Trafton and Boehm-Davis (2008), peer-reviewed in the Journal of Experimental Psychology: Applied. Resumption lag was 949 ms (SD 283) in the uninterrupted baseline, rising to 1,157 ms, 1,398 ms and 1,639 ms after interruptions of 3 s, 8 s and 13 s respectively. Their second experiment established the functional form: resumption lag is logarithmic in interruption duration, fitted as y = 1032 + 189.4·log(x), R² = .989, asymptoting between 13 and 23 s. The cost lives in the switch, not the duration. A one-minute exception costs barely more Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 15 in resumption than a twenty-second one, so the arrival rate of exceptions, not their length, consumes the overseer. Their third experiment established what modulates the cost. Resumption lag was 1,322 ms when the interruption involved no task, 1,605 ms with a tracking task, and 1,789 ms with a high-demand nback task, the last producing roughly three times the error rate of the others (.06 versus .02 and .01). A supervisor pulled from one demanding exception into another pays a compounding cost, in errors as well as in seconds. The mechanism — a suspended goal must be re-retrieved from declarative memory through environmental cues after its activation has decayed — comes from Altmann and Trafton's (2002) memory-for-goals model, which contains no latency values of its own. The field magnitudes are larger by three orders of magnitude. Mark, Gonzalez and Harris (2005), a peer-reviewed CHI full paper, observed 24 knowledge workers in detail — seven managers, nine analysts, eight developers. The average length of time informants spent in central and peripheral working spheres was 11 min 4 s; 57.1 per cent of all working-sphere segments were interrupted; and when people did resume work on the same day, it took an average of 25 min 26 s. Two points of citation hygiene are required, and both are substantive. The figure of "23 minutes 15 seconds", almost universally attributed to this paper in management writing, does not appear in it; the paper's figure is 25 min 26 s. And that figure is conditional on same-day resumption — segments never resumed that day are excluded — which biases it downward relative to true resumption delay. The second hygiene case is more consequential. Mark, Gudith and Klocke (2008), also a peerreviewed CHI full paper with N = 48, is routinely cited as evidence that interruptions slow people down. It shows the opposite on time. Interrupted work was completed faster: 20.31 and 20.60 minutes in the same- and different-context conditions against a 22.77-minute uninterrupted baseline, F(2, 77.98) = 3.36, p < .05. The cost appeared elsewhere. Stress rose from 6.92 to 9.46 on a 1– 20 scale (F(2,92) = 12.15, p < .001); workload, frustration, time pressure and effort all rose significantly; and interrupted participants wrote shorter emails (31.49 words down to 29.17 and 30.16, p < .04). Used for the finding it actually supports, this is one of the most important results in the paper. The supervisor absorbs the cost of exception handling as strain and as compressed output, not as visible delay. The work still gets done, on time or sooner, which is precisely why an insolvent oversight regime does not announce itself: there is no queue to observe, no missed deadline to escalate, no throughput signal that degrades. The indicators a management system watches remain green while the quality of the human's contribution is silently traded away. Section 4.4 shows what that trade looks like once it has run to completion in a deployed exception channel. Table 2 collects the measured latencies on which Sections 3.2 and 3.3 rest, each figure placed beside the limitation its own authors state. Study Setting and N Measured latency Caveat stated by the authors Zhang et al. (2019) metaanalysis 129 studies, 520 observations, simulators Mean take-over 2.72 s (SD 1.45); range 0.69–19.79 s; +1.35 s for a larger time budget Collision risk is set by the tail, not the mean; quality not examined. Gershon, Mehler & Reimer (2023) Real roads, hands-free L2, 14 drivers, 265 disengagements Take-over 2.4 s hands-free, 1.8 s with a hand on the wheel; 6.1 s to begin recovering on-road glance behaviour Single vehicle platform; small driver sample. Louw et al. (2017) Simulator, N = 75 Take-over time and avoidance quality "largely independent" Simulator; induced out-ofthe-loop states. Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 16 Study Setting and N Measured latency Caveat stated by the authors Dixit, Chand & Nair (2016) 460,097 autonomous miles, California disengagement reports Pooled mean reaction 0.83 s; reaction time rose with cumulative miles Excludes perception time; professional safety drivers; only two firms reported exact values. Monk, Trafton & Boehm-Davis (2008) Laboratory, N = 12 and 36 Resumption lag 949 ms uninterrupted, rising to 1,639 ms after a 13 s interruption; 1,789 ms after a high-demand interruption Environmental cues may have aided resumption; no conclusions beyond one minute. Mark, Gonzalez & Harris (2005) Field ethnography, 24 knowledge workers Working-sphere segments 11 min 4 s; same-day resumption 25 min 26 s; 57.1% of segments interrupted Conditional on same-day resumption, so downward biased. Table 2. Measured human intervention latency. 3.4 What the system side of the ratio looks like Against these human latencies, consider two documented automated decision classes at opposite ends of the evidence-grade scale. The first is a primary regulatory record. The U.S. Securities and Exchange Commission's Administrative Proceeding, Release No. 34-70694, documents that on 1 August 2012 Knight Capital's system produced four million executions in 154 stocks for more than 397 million shares in approximately 45 minutes, and a loss of approximately $460 million. Four million executions over roughly 2,700 seconds is a derived rate of approximately 1,481 executions per second — an interexecution interval on the order of 0.7 ms. Against a human response latency measured in seconds, the ratio is roughly four orders of magnitude, and no governance policy or designated accountable officer alters that arithmetic. Section 9 returns to the case, where the point of interest is not that oversight was absent but that human operators were working the incident throughout and their first corrective action worsened it. The second is a preprint and should be labelled as such. OSWorld 2.0 (arXiv:2606.29537) reports approximately 318 tool calls per task against a median human completion time of about 1.6 hours, where the 2024 predecessor benchmark required roughly 30. The two figures are not a series and are not used as one here. OSWorld 2.0 selects long-horizon tasks by design, and its own maintainers state that scores across the earlier and later versions are not comparable, so the difference in tool calls is a property of the tasks selected rather than of agent behaviour over time; the later tasks are also much longer, so the two counts are not comparable per unit of work either. What the figure establishes is the action-to-attention ratio for one documented class of task, and that is all the argument requires. The dilemma it creates should be stated plainly, because it is the binding design constraint and it is not a capability question. If the human must review each action to retain accountability for the outcome, then at 318 actions per task the oversight costs more than the work, and the configuration is in-the-loop only in being slower and dearer than doing the task unaided. If the human reviews only the final output, they are accepting approximately 318 unreviewed state changes — file writes, form submissions, API calls, deletions — each of which may have committed a consequence the final output does not display. There is no third option: sampling is the second branch with a stated coverage fraction, and automated pre-filtering relocates the review to another automated system, which Sections 6 and 7 address. Model capability is not the binding constraint on oversight design. The action-to-attention ratio is. Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 17 Figure 2. The determination of loop position. Measured human latencies against system consequence latencies on a logarithmic time axis. Human side: take-over time 2.72 s mean, 0.69–19.79 s across conditions (Zhang et al. 2019); attentional recovery onset 6.1 s (Gershon et al. 2023); resumption lag 0.95–1.9 s (Monk et al. 2008); field resumption 25 min 26 s, conditional on same-day resumption (Mark et al. 2005). System side: approximately 0.7 ms between executions at the Knight Capital rate (SEC Rel. 34-70694), an agent tool call, an automated credit decision, a quarterly capital allocation. Shading marks τsys < τH , where an on-the-loop configuration is unavailable irrespective of policy. 3.5 Two propositions, and why formal authority survives the collapse of real authority Proposition 1 (Position determination). Holding the human's response latency fixed, any increase in the throughput or consequencespeed of an automated decision class monotonically shifts the achievable oversight position from in-the-loop toward out-of-the-loop, regardless of the governance policy asserted. Refutation condition. Refuted if a deployment is documented in which mean consequence latency fell below the measured human end-to-end response latency for a decision class and contemporaneous human approval nonetheless remained determinative of outcomes — i.e. approvals materially altered the action distribution rather than ratifying it. Proposition 1 is a statement about feasibility. It does not explain why a firm operating past that boundary continues to report that a human is in control, nor why the humans concerned continue, in good faith, to sign. The explanation lies in the economics of authority, and it is not a story about weakness of will. Aghion and Tirole (1997), Journal of Political Economy 105(1):1–29 — peerreviewed and theoretical, with no data — distinguish formal authority, the right to decide, from real authority, effective control over decisions, and make real authority a function of the structure of information. Where the principal is uninformed and the agent informed, their phrasing is that a poorly informed principal optimally rubber-stamps the subordinate's proposal by fear of picking a worse alternative. The word carrying the argument is optimally. Rubber-stamping is not a failure of diligence to be corrected by exhortation; it is the informed response of a decision-maker who knows The determination of loop position on-the-loop unavailable: τsys < τH 1 ms 10 ms 100 ms 1 s 10 s 1 min 10 min 1 h human response latency τH take-over time, mean 2.72 s (range 0.69–19.79 s) real-road take-over 2.4 s attentional recovery 6.1 s resumption lag 0.95–1.9 s field task resumption 25 min 26 s system consequence latency τsys Knight Capital: ~0.7 ms between executions one agent tool call automated credit decision quarterly capital allocation Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 18 that overruling a better-informed party raises the probability of a worse outcome. Formal authority survives intact. It is simply nominal. The determinants of the subordinate's real authority that Aghion and Tirole enumerate for a formally integrated structure are overload, lenient rules, urgency of decision, reputation, performance measurement, and multiplicity of superiors. Two of those — overload and urgency of decision — are not accidents that befall an on-the-loop configuration; they are what it constructs by design. Routing exceptions from a high-throughput system to a human who has other work is the definition of overload; requiring the decision inside τsys is the definition of urgency. The configuration manufactures, as a matter of architecture, the two conditions the model identifies as transferring real authority away from its formal holder. Proposition 2 (Rubber-stamping). Where the human's information about a decision is strictly poorer than the system's and the decision must be taken within τsys, the human's formal authority is exercised as ratification. Real authority has transferred to the system. Refutation condition. Refuted if, under those informational and temporal conditions, override rates and override accuracy are shown to be independent of the human's information disadvantage. There is a stronger version of this argument that the paper must confront honestly rather than deploy selectively. Dessein (2002), Review of Economic Studies 69(4):811–838, peer-reviewed and theoretical, shows that a principal prefers to delegate control to a better-informed agent rather than communicate with that agent, as long as the incentive conflict is not too large relative to the principal's uncertainty about the environment — and, decisively here, that keeping a veto right typically reduces the principal's expected utility unless the incentive conflict is extreme. The mechanism is that communication is lossy: under cheap talk the agent's information is coarsened in transmission, so the principal trades noise-in-communication against bias-in-objective, and delegation wins over a wide parameter range. This cuts against the paper's own subject matter and should be allowed to. Dessein is a strong argument that veto-based oversight — precisely the technical content of the on-the-loop configuration, as Section 2 established — is usually value-destroying relative to clean delegation plus objective alignment. A firm installing a human veto over a better-informed system, where objectives are broadly congruent, is on this analysis making itself worse off and calling it governance. But the honest reading runs in both directions. Dessein's veto result is explicitly conditioned on the incentive conflict not being extreme, and he leaves the extreme case open. Safety-critical deployment, and more generally deployment where the system's objective is a proxy that diverges from the principal's in the tail, is a candidate for exactly that carve-out. The result is a genuine constraint on when veto-based oversight is worth its cost, and it implies that configurations installed for reassurance rather than for tail risk are net-negative. What it does not license is the inference that veto authority is always value-destroying. It relocates the question to whether the misalignment is extreme — an empirical question about a given deployment, which Sections 6 and 7 take up in terms of information channels and verifier independence rather than incentive conflict. The section closes on the observation that motivates Section 4. A firm that keeps "a human in the loop" while raising throughput does not retain in-the-loop control. It silently converts in to on, and then on to out, because the ratio defining the position moves while the policy naming it does not. The conversion is invisible from the dashboard because approval rates do not fall; they rise, or hold at a level read as evidence that the system performs well. The overseer is still there, still signing, still Human-on-the-Loop · WP No. 8 3. Position Is Determined, Not Chosen 19 accountable — and, by Aghion and Tirole's own logic, still behaving optimally. What has been lost is not presence, diligence or formal rights, but the precondition under which any of the three would have had control content. 4. Oversight as a Consumable Resource 4.1 Oversight is not a role Section 3 established that loop position is determined by a ratio the firm does not control. This section makes the second of the paper's three theoretical moves: oversight is not a role that can be assigned, but a consumable resource with a measurable capacity, and that capacity can be exhausted. A firm that assigns oversight the way it assigns a job title has named a duty without provisioning the input the duty consumes. The input is attention, and the canonical statement of its scarcity is Simon (1971). The source is a chapter in an edited conference volume — Computers, Communications, and the Public Interest, edited by Martin Greenberger, Johns Hopkins Press, pp. 37–72 — and it is not a peer-reviewed article; it should be cited as what it is, an argument of exceptional influence rather than a tested finding. The passage, at p. 40, reads: "In an information-rich world, the wealth of information means a dearth of something else: a scarcity of whatever it is that information consumes. What information consumes is rather obvious: it consumes the attention of its recipients. Hence a wealth of information creates a poverty of attention and a need to allocate that attention efficiently among the overabundance of information sources that might consume it." Ocasio (1997), Strategic Management Journal 18(S1):187–206, peer-reviewed and theoretical, converts Simon's individual-level scarcity into a firm-level design variable. His central argument is that firm behaviour is the result of how firms channel and distribute the attention of their decisionmakers, structured through three linked principles: focus of attention, situated attention, and the structural distribution of attention through the firm's rules, resources and relationships. On this account the firm does not merely use attention. The firm is an attention-allocation structure. That reading licenses the move this section makes. If the firm is an attention-allocation structure, then introducing an automated system that emits candidate actions faster than attention can be allocated to them is not a productivity change. It is a change in the firm's constitution. It alters which situations decision-makers find themselves in, which issues reach them, and in what volume — the three variables Ocasio identifies as constitutive of firm behaviour. A deployment decision taken as an efficiency matter, at the level of a functional budget, is in fact a decision about where the firm's authority resides. The remainder of this section makes the capacity constraint explicit enough that the decision can be taken with its consequences visible. 4.2 The solvency constraint Definition 3 (Oversight capacity). The overseer's uncommitted attention A over an interval is the quantity of time in that interval not already claimed by other assigned work. Interaction time IT is the mean time to handle one exception to closure. Context re-acquisition cost c is the additional time to restore the working state of the task the overseer was displaced from. Neglect time NT is the interval a supervised process can proceed unattended before its expected performance falls below threshold. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 20 Definition 4 (Oversight solvency). An on-the-loop configuration with exception arrival rate λ is solvent iff λ·(IT + c) ≤ A and the supervised process count N satisfies N ≤ NT/(IT + WT) + 1, where WT is the queueing delay when several processes demand attention simultaneously. A configuration failing either condition is insolvent. Insolvency is not observed as refusal; it is observed as acceptance — override, dismissal and non-response rise until the constraint is satisfied by shortening IT to zero. One property of that statement should be flagged before the formula is read, because the notation invites the opposite reading. Definitions 3 and 4 write A, IT, c and NT as constants over the period, which is a simplification adopted for tractability rather than a claim about the overseer: each of the four drifts within a single period of duty, and in the direction that tightens the constraint. Section 4.7 assembles the measured instances and draws the consequences. ρ = λ·(IT + c)/A ≤ 1   and   N ≤ NT/(IT + WT) + 1 The first condition is written here as a utilisation ratio, which is the form in which it should be reported. Both sides of Definition 4's inequality are durations per period: λ is a count of exceptions per period and IT and c are times per exception, so their product is the time the exception stream demands in that period, while A is the time available in the same period. Their quotient ρ is therefore dimensionless, and solvency is the statement that it does not exceed unity. Reporting ρ rather than the two sides separately has a practical merit: it is the single number that says how far a configuration is from the boundary, and a firm that reports a value above one has reported that its oversight is insolvent without needing any further interpretation. The period must be stated and must be the same on both sides — a working day is the natural choice where exceptions arrive continuously — and the span bound is dimensionless in the same way, NT, IT and WT all being times per event. Each term is operational, and a firm that wished to could measure all six. The arrival rate λ is the count of exceptions routed to a named overseer per unit time, taken from the routing system rather than the design document; it is the one term firms almost always have and never publish. Interaction time IT is the elapsed time from an exception surfacing to its disposition being recorded, measured to closure and not to first touch — the distinction Section 3.2 drew between motor handover and restored competence applies here, and an IT measured to acknowledgement is measuring the wrong event. Context re-acquisition cost c is the additional time required to restore the displaced task, which Section 3.3 showed to be between roughly 0.95 and 1.9 s per switch in controlled conditions and, in observed knowledge work, on a wholly different scale. Uncommitted attention A is the term firms systematically overstate, because it is not headcount and it is not nominal availability. It is the time in the period not already claimed by other assigned work, and the measurement is subtractive: contracted time less every other duty carrying a delivery expectation. For an overseer who retains a full operational role that subtraction will frequently return a value at or near zero. That is not a defect in the test but its result. It says that no time has been provisioned for oversight, so ρ exceeds unity at any positive arrival rate and the configuration is insolvent by construction — which is a finding about the configuration, not a deficiency of the definition. The corresponding instruction is that A must be provisioned rather than measured: work is removed from the role until the time released covers λ·(IT + c), and the disclosure that evidences the provisioning is the list of duties removed when the oversight duty was added. This is Proposition 4 restated as an instruction. A is a quantity a firm allocates; it is not a residual a firm discovers. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 21 Neglect time NT is the only term that is a property of the supervised system rather than of the human: the interval a process can run unattended before its expected performance falls below threshold. Its analogue in an agentic deployment is the number of consequential actions taken before an error becomes costly to reverse, which is why Section 3.4's action counts bear on it. And WT, queueing delay, makes the second inequality bite: it is zero in the single-process case and grows with N, so span degrades faster than the naive ratio NT/IT suggests. The two conditions are less independent than their presentation as a pair suggests, and what relates them is an assumption about where the exception stream comes from. The span bound is derived on the supposition that each supervised process demands attention once per neglect interval, that is λ = N/NT. That is a substantive restriction on the deployment rather than a modelling convenience, and three regimes must be distinguished, because the model's applicability and the remedy for insolvency both differ across them. endogenous:  λ = N/NT   ·   exogenous:  λ = N/NT + λext ⇒  N ≤ (A/(IT + c) − λext)·NT In the first regime the stream is endogenous: every exception originates in a supervised process and none arrives from anywhere else. Then λ = N/NT holds exactly, and substituting it into the solvency condition λ·(IT + c) ≤ A returns N ≤ NT·A/(IT + c), which is the span bound up to the additive constant and the substitution of the queueing correction WT for the re-acquisition cost c. The two inequalities are one constraint written twice, and a firm in this regime that satisfies either satisfies both. The remedy for insolvency is correspondingly single-valued: reduce N or raise NT — supervise fewer processes, or make each of them safer to leave alone. In the second regime part of the stream is exogenous, arriving independently of the supervised processes: a customer escalation, a regulator's query, an incident reported by someone outside the system. The arrival rate is then λ = N/NT + λext, and substituting that into λ·(IT + c) ≤ A tightens the span bound to N ≤ (A/(IT + c) − λext)·NT. The two inequalities have now diverged, because solvency binds on the total while the span bound was derived from the endogenous part alone. The consequence is a usable result and should be stated plainly. Exogenous arrivals consume supervisory span directly, at an exchange rate of NT processes for each unit of exogenous arrival rate: a channel receiving one further escalation per neglect interval has lost the capacity to supervise one process, whatever the supervised processes happen to be doing. A firm therefore cannot infer its supervisable span from the behaviour of the supervised processes alone, and an oversight configuration sized against internal telemetry — which records what the system emits and not what arrives about it — will be systematically oversized. The third regime is mixed and non-stationary, and it is the realistic one. λext is typically bursty rather than smooth, and it is correlated with the failures the supervision exists to catch: a fault in the supervised processes generates escalations, enquiries and regulatory attention about that fault, so the exogenous term peaks at exactly the moment the endogenous term does. The model as written treats λ as a rate and therefore averages over that structure, which understates the constraint under bursty arrival — a configuration solvent in the mean can be insolvent throughout the interval in which the exceptions actually arrive. This is a limitation of the model rather than a result of it. The queueing delay WT is its sole representation of contention, it is not derived here from any arrival process, and no distributional assumption is offered from which it could be. The two conditions should therefore be read as necessary and not as sufficient, and a firm should state which of the three regimes it takes itself to be in. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 22 4.3 The measured span Span has been measured properly in one setting: supervisory control of multiple semi-autonomous robots. Crandall and Cummings (2007), peer-reviewed in IEEE Transactions on Robotics 23(5), with N = 12 participants supervising teams of two, four, six and eight robots, found interaction time of approximately 15–17 s and — the notable part — remarkably stable across team sizes, with only 2.34 s difference between the two-robot and eight-robot conditions. Neglect time rose steeply, from approximately 30 s at two robots to over 120 s at eight. This is the citable source for both fan-out formulations: FO = NT/IT + 1, and, with queueing, FO = NT/(IT + WT) + 1. Measured capacity fell between four and six robots. Effectiveness plateaued at six and degraded at eight, where operators lost more robots. The failure signature is the finding the present paper most needs: six-robot teams collected significantly more hard-to-reach objects than smaller teams but lost more robots. The capacity boundary announced itself as a quality and loss failure, not as a throughput failure. Throughput was still rising when quality had already turned. The replication and extension is Crandall, Cummings, Della Penna and de Jong (2011), peerreviewed in IEEE Transactions on Systems, Man, and Cybernetics — Part A 41(3):385–397, with N = 16 in the first study and N = 12 in the second. Average tokens collected peaked at about six robots; eight-robot teams performed significantly worse, driven by a twelvefold increase in robot losses between two and eight. The authors trace the mechanism explicitly to attention allocation — the decrease in system effectiveness in large teams is attributed to the way operators allocated their attention among the robots — rather than to any deficit of speed. The concrete signature was that roughly 40 per cent of robots in the eight-robot condition were sent out without collecting tokens in the final minute, against roughly 10 per cent in the four- and six-robot conditions. Their second study should discipline every claim that better tooling raises the span. Three interface modes were compared for supporting attention allocation. The automated mode achieved over 95 per cent adherence to optimal recommendations and significantly reduced perceived workload, and produced no statistically significant performance improvement whatever: F(2,33) = 0.50, p = 0.609. Sixty-seven per cent of users disliked it. Automating the attention-allocation layer improved compliance and comfort and did not improve outcomes. A tooling investment that moves the compliance metric and the workload metric without moving the performance metric has bought the appearance of capacity. The conceptual account of why the ceiling sits where it does comes from Olsen and Goodrich (2003), which introduced the fan-out construct. This is a PERMIS workshop paper and is not peer-reviewed; it contains no measured values and should be treated as a conceptual source. Its useful line is that context re-acquisition is probably the largest contributor to an upper bound on fan-out. The ceiling on supervision is set by the cost of rebuilding situational context, not by the mechanical cost of acting — which is precisely why the term c appears in Definition 3 and why IT measured to first touch understates the true cost. One negative result must be stated plainly. Cummings and Mitchell (2008), IEEE Transactions on Systems, Man, and Cybernetics — Part A 38(2):451–460, is frequently cited for a specific operatorutilisation ceiling above which supervisory performance degrades. The verification pass conducted for this paper could not obtain any quantitative result from it: the bibliographic record was confirmed at Crossref, but no figure from the paper could be independently checked. This paper therefore cites the record and quotes no number from it, and readers encountering a utilisation threshold attributed to that source elsewhere should treat it as unverified here. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 23 Study N Interaction time (IT) Neglect time (NT) Effective span Failure signature Crandall & Cummings (2007), IEEE T-Robotics 23(5); peer-reviewed 12 ≈15–17 s; stable across team size (2.34 s difference, 2 vs 8 robots) ≈30 s at 2 robots to >120 s at 8 4–6; plateau at 6 At 6: more hard-to-reach objects collected and more robots lost. Quality/loss failure before throughput failure. Crandall, Cummings, Della Penna & de Jong (2011), IEEE TSMC-A 41(3):385–397, Study 1; peerreviewed 16 — — Peak system score ≈6 robots 12× increase in robots lost from 2 to 8; ≈40% of robots idle-dispatched at 8 vs ≈10% at 4–6. Crandall et al. (2011), Study 2; peerreviewed 12 — — — Automated attention allocation: >95% adherence, lower workload, no significant performance gain (F(2,33) = 0.50, p = 0.609); 67% of users disliked it. Olsen & Goodrich (2003), PERMIS workshop; not peerreviewed — No measured values No measured values Framework only (FO = NT/IT + 1) Ceiling attributed to context re-acquisition rather than to the cost of acting. Cummings & Mitchell (2008), IEEE T-SMC-A 38(2):451–460; peerreviewed — Record verified; no quantitative result obtainable in verification. No figure quoted in this paper. Table 3. Measured supervisory span in the settings where it has been measured. 4.4 What insolvency looks like in the field The robotics studies measure the boundary under laboratory conditions. What an exception channel looks like after years of operation past that boundary is documented at very large scale in clinical decision support. van der Sijs, Aarts, Vulto and Berg (2006), a peer-reviewed systematic review in JAMIA 13(2):138–147 covering 17 papers, found override rates for drug safety alerts ranging from 49 to 96 per cent. Poly, Islam, Yang and Li (2020), a peer-reviewed systematic review in JMIR Medical Informatics 8(7):e15653 covering 23 articles, found override rates of 46.2 to 96.2 per cent. The comparison should be made explicitly, because it is the strongest single argument in this paper against interfacelevel remedies: across fourteen years of alert redesign, tiering, severity grading and usability work, the override range moved essentially not at all. Whatever is generating these rates is not an interface defect. The cleanest single-system measurement is Nanji, Slight, Seger, Cho, Fiskio, Redden, Volk and Bates (2014), peer-reviewed in JAMIA 21(3):487–491: 2,004,069 medication orders generated 157,483 alerts, an alert rate of 7.9 per cent, of which 82,889 — 52.6 per cent — were overridden. The same study supplies the finding that prevents the argument from becoming a complaint about clinicians. A subsample of 600 overrides was clinically reviewed, and 53 per cent were judged appropriate, with a range from 12 per cent for renal recommendations to 92 per cent for patient allergies. A high override rate is therefore not by itself evidence of human failure. Roughly half the time the human was right and the automation was wrong, which is evidence of a miscalibrated channel rather than Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 24 of an unreliable overseer. But the renal figure carries the opposite implication with equal force: in that specific low-salience, high-stakes category, the human was wrong about 88 per cent of the times they overrode. Both directions belong in the same paragraph, and an oversight regime designed on either one alone will be designed wrongly. The dose–response evidence comes from Ancker, Edwards, Nosal, Hauser, Mauer and Kaushal (2017), peer-reviewed in BMC Medical Informatics and Decision Making 17:36 — a retrospective cohort of 112 ambulatory primary care clinicians, 1,266,325 best-practice advisories and 326,203 drug interaction alerts. Note that a published Correction exists (PMID 31739801) and should be consulted before any figure is reused. Drug-alert acceptance overall was less than 1 per cent: the channel is, functionally, dead. Acceptance fell by 30 per cent for each additional reminder received per encounter (IRR = 0.70, p < .001). The path dependence is striking: where the first instance of an alert was overridden, 87.9 per cent of subsequent instances were overridden for advisories and 99.9 per cent for drug alerts, against 51.9 and 58.4 per cent respectively where the first instance had been accepted. And the finding that identifies the mechanism is a null: general workload showed no significant association with acceptance. It is alert volume per encounter, not busyness, that drives desensitisation. That distinction matters for design, because it means the remedy is not a lighter workload but a lower arrival rate — λ, not A. The precision side of the same equilibrium is measured by Drew, Harris, Zègre-Hemsey and colleagues (2014), peer-reviewed in PLOS ONE 9(10):e110274: 461 consecutive ICU patients over 31 days generated 2,558,760 unique alarms, of which 381,560 were audible — equal to 187 audible alarms per bed per day, roughly one every eight minutes around the clock. Of 12,671 annotated arrhythmia alarms, 88.8 per cent were false positives. The study could not report false negatives, a limitation the authors state. What these two studies jointly support should be stated carefully, because they are not one system and their figures do not compose. Drew et al. measure the arrival rate and the precision of a single saturated channel: 187 audible alarms per bed per day from continuous physiologic monitors, with 88.8 per cent of annotated arrhythmia alarms false. Ancker et al. measure, in a different channel with a different population, signal type, response modality and time base — ambulatory advisories and drug alerts routed to a clinician at an encounter, not a bedside monitor running around the clock — the dose–response by which acceptance falls as volume per encounter rises. Per bed-day and per encounter are not the same denominator, and nothing here is a natural experiment: there is no exogenous variation in the arrival rate, no counterfactual arm and no assignment mechanism, and neither study could measure its false-negative rate, which is precisely the quantity an insolvency claim would want. Taken together, what they are consistent with is the mechanism rather than a magnitude: a channel whose arrival-rate-weighted cost exceeds the attention available to it is satisfied in the only way available to it, by driving interaction time toward zero, which is what Definition 4 says insolvency looks like. Read that way, desensitisation is not an indiscipline to be corrected by training, reminders about professional responsibility or a stronger accountability framework. It is the behaviour the configuration selects for. Study Design and scale Override / acceptance Appropriateness van der Sijs et al. (2006), JAMIA 13(2):138–147; peer-reviewed systematic review 17 papers Override 49–96% Not pooled; authors note studies of the cognitive processes behind overriding are lacking Nanji et al. (2014), JAMIA 21(3):487–491; peerreviewed 2,004,069 orders → 157,483 alerts (7.9%) 82,889 overridden (52.6%) 600 reviewed: 53% appropriate; 12% (renal) to 92% (allergy) Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 25 Study Design and scale Override / acceptance Appropriateness Ancker et al. (2017), BMC MIDM 17:36; peerreviewed (see Correction, PMID 31739801) 112 clinicians; 1,266,325 advisories + 326,203 drug alerts Drug-alert acceptance <1%; acceptance −30% per additional reminder per encounter (IRR 0.70, p < .001) Not assessed; path dependence 87.9% (advisories) / 99.9% (drug) after a first override; workload not significant Poly et al. (2020), JMIR Med Inform 8(7):e15653; peer-reviewed systematic review 23 articles Override 46.2–96.2% 29.4–100% overall; drug–drug interaction 0–95% Drew et al. (2014), PLOS ONE 9(10):e110274; peerreviewed observational 461 ICU patients, 31 days; 2,558,760 alarms, 381,560 audible 187 audible alarms per bed per day 88.8% of 12,671 annotated arrhythmia alarms false positive; false negatives not measurable Table 4. Exception-channel saturation in deployed alerting systems, 2006–2020. 4.5 Vigilance is work, not idleness The solvency constraint assumes that watching costs something. That assumption is not obvious — the managerial intuition is that monitoring is a light duty which can be layered onto an existing role — and it is the assumption on which Proposition 4 turns. The evidence contradicts the intuition directly. Warm, Parasuraman and Matthews (2008), Human Factors 50(3):433–441, is a peer-reviewed narrative review rather than an experiment, and should be labelled as such; it reports no N of its own. Its conclusion is that converging evidence from behavioural, neural and subjective measures shows that vigilance requires hard mental work and is stressful. Vigilance tasks are resourcedemanding, not under-stimulating: NASA-TLX global workload scores for vigilance tasks fall at the upper end of the scale, transcranial Doppler measures show cerebral blood flow velocity declining in parallel with the performance decrement, and subjective measures show reduced task engagement with increased distress. 

On the magnitude of the decrement, they report that detection accuracy declines by about 10 to 15 per cent after only about 30 minutes, with most of the loss occurring within the first 15. Parasuraman and Manzey (2010), Human Factors 52(3):381–410, an integrative review and again not an experiment, establishes the conditions under which the failure appears. Automation complacency occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator's attention. That condition is the condition of essentially every real corporate deployment: the overseer is not sitting in front of a single monitored stream with nothing else to do. Three further claims from the same review are unusually strong. Complacency is found in both naive and expert participants and cannot be overcome with simple practice. Automation bias produces both omission and commission errors when decision aids are imperfect, occurs in experts as well as novices, and cannot be prevented by training or instructions. And it affects decision-making in teams as well as in individuals, so pairing overseers does not repair it. The managerial conclusion is sharp and is worth stating without qualification. A firm that assigns an oversight duty to an existing role without removing work from that role has not added oversight capacity. It has reallocated attention within a fixed budget, and it has done so in exactly the multiple-task-load condition under which the complacency literature predicts the duty will be discharged as acceptance. The standard corporate remedies — a training module, a written instruction to be careful, a second reviewer — are each addressed by name in Parasuraman and Manzey's findings, and none of them survives. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 26 4.6 A counterpoint: monitoring is not a neutral instrument It would be convenient to conclude that the answer to insolvent oversight is more oversight, better resourced. A brief counterpoint is necessary, because the evidence does not support that direction either. Falk and Kosfeld (2006), American Economic Review 96(5):1611–1630, a peer-reviewed laboratory experiment, found hidden costs of control: in the C5 treatment agents' mean effort was 12.19 under control against 25.11 where control was available but not used, and in C10, 17.53 against 22.99. This result must be cited alongside its replication. Ziegelmeyer, Schmelz and Ploner (2012), Experimental Economics 15(2):323–340, peer-reviewed, with N = 476 across four repetitions, found substantially smaller effects and conclude that hidden costs almost never statistically significantly outweigh the benefits of control. The defensible claim is therefore narrow: control can carry motivational costs whose magnitude is contested, appears smaller than first estimated, and shrinks as the control binds more tightly. The stronger caution comes from the audit literature. Bevan and Hood (2006), Public Administration 84(3):517–538, is a peer-reviewed documentary case analysis of a natural policy experiment in the English public health system — not a controlled study, with no N and no counterfactual, and it should be cited on those terms. Their typology is directly portable to oversight metrics: threshold effects, where performance bunches at the target and there is no incentive to exceed it; ratchet effects, where current performance is suppressed to avoid a tougher future target; and output distortion, where unmeasured quality is neglected. Their structural finding is that the measurement system itself went unaudited. Strathern (1997), European Review 5(3):305–321, an anthropological essay with no data, supplies the mechanism in a sentence: audit "does more than monitor — it has a life of its own that jeopardizes the life it audits." Taken together with Sections 4.3 and 4.4, this closes off the intensification route. Oversight cannot be made solvent by inspecting harder, because inspection consumes the resource that is already binding, distorts the behaviour it measures, and — on the Crandall et al. evidence — does not improve outcomes even when its own compliance metrics improve. That is why the constructive proposal developed in Section 8 relocates the human's authority to the class of decisions whose attention cost does not scale with throughput, rather than intensifying inspection of decisions whose cost does. Proposition 3 (Oversight solvency). An on-the-loop configuration whose exception load λ·(IT + c) exceeds uncommitted attention A does not fail by refusing exceptions; it fails by accepting them. Override, dismissal and nonresponse rates rise until the constraint binds. Refutation condition. Refuted if an exception channel operating persistently above solvency is documented in which per-exception handling quality did not decline. Proposition 4 (Substitution, not addition). Because attention is fixed in the short run, assigning an oversight duty to a role without removing work from that role does not add oversight capacity; it reallocates it. Refutation condition. Refuted if an oversight duty added to an unchanged workload is shown to produce detection performance equal to that of a dedicated overseer with equivalent training and information. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 27 Proposition 5 (Span). Where the exception stream is generated by the supervised processes themselves, the number of concurrent semi-autonomous processes a single human can supervise before quality degrades is bounded by NT/(IT+WT)+1, and in the only settings where it has been measured is a single-digit number. Where part of the arrival stream is exogenous to those processes, the bound tightens to N ≤ (A/(IT+c) − λext)·NT, so that exogenous escalations consume supervisory span directly and a firm cannot infer its span from the behaviour of the supervised processes alone. Exceeding the bound manifests first as loss quality, not as throughput loss. Refutation condition. Refuted if a controlled study measures effective supervisory span in a knowledge-work setting materially exceeding the measured robotics range without a corresponding increase in neglect time. 4.7 The parameters are not constants Definitions 3 and 4 write A, IT, c and NT as constants, and ρ is evaluated as though each held its value across the whole period. That is a simplification adopted for tractability, and the evidence already assembled in this paper contradicts it. Four measurements set out above and in Section 3 record one of these quantities moving within a single period of duty, as cognitive fatigue and cumulative interruption accumulate. None was collected to test the constancy of these parameters, and the four come from unrelated paradigms; but they move in one direction, and it tightens the constraint. The first is the vigilance decrement of Section 4.5. Warm, Parasuraman and Matthews (2008) report detection accuracy declining by about 10 to 15 per cent after only about 30 minutes, most of the loss occurring within the first 15, and vigilance tasks being resource-demanding rather than understimulating: what falls within a single watch is the detection probability entering Proposition 12's efficacy term and the derivation of Section 8.3.1. The second is the strongest, being a within-period decay measured directly against cumulative interruption. Ancker and colleagues (2017) found acceptance falling by 30 per cent for each additional reminder received per encounter, IRR = 0.70, p < .001, while general workload showed no significant association with acceptance: what degrades the response is not busyness but the number of interruptions already absorbed, which is the statement that these quantities are functions of the interruption count rather than constants of the person. The third is the demand effect of Section 3.3. Monk, Trafton and Boehm-Davis (2008) measured resumption lag rising from 1,322 ms after a no-task interruption to 1,789 ms after a highdemand one, the last carrying roughly three times the error rate, so that a supervisor pulled from one demanding exception into another pays a compounding cost and c depends on what they were pulled into as well as on what they were pulled out of. The fourth is the only instance from a deployed system: Dixit, Chand and Nair (2016) found reaction time increasing with cumulative autonomous miles — trust degrading response speed with exposure, on the authors' interpretation — in paid professional safety drivers whose sole task was to watch. Two consequences follow. The first is that the constraint as written is an optimistic bound. Every parameter drifts adversely: A falls as fatigue claims part of what was nominally available, effective IT and c rise with the interruptions absorbed, detection falls with time on watch, and no drift offsets another. A configuration evaluated on first-hour parameters may therefore satisfy ρ ≤ 1 comfortably and be insolvent by the end of a shift, and the static form of the inequality will not show it: the arithmetic returns the same verdict at the eighth hour as at the first, on inputs that no longer describe the overseer. The static solvency condition should accordingly be read as the best case for a configuration rather than the typical one. Human-on-the-Loop · WP No. 8 4. Oversight as a Consumable Resource 28 The second consequence is a measurement instruction, and it is cheap. A firm reporting ρ must state the point in the period at which the measurement was taken, and should report it at least twice, early and late; the difference between the two is the fatigue term. Obtaining it requires no theory of how the decay proceeds and no new instrumentation: the routing logs and workflow timestamps from which λ, IT and c are already drawn carry their own clock, so partitioning them by time within the period yields both estimates at no additional cost and satisfies the acceptance rule of Section 11.8. The drift is thereby made measurable even though this paper does not model its functional form, and a channel at ρ = 0.8 early and 1.4 late is insolvent in a way a single mid-period measurement records as solvent. What is not being claimed should be stated with equal clarity. No functional form for the decay is proposed here, and none of the four measurements licenses one. They are not commensurable: a laboratory vigilance decrement, a dose–response in clinical encounters, a resumption lag in milliseconds and a reaction time accumulated over driving miles have four populations and four time bases between them, and nothing is gained by combining them into a coefficient. What they support jointly is the sign of the drift, not its magnitude. A dynamic version of the constraint — in which A, IT, c and the detection probability are written as functions of elapsed time and of the cumulative interruption count, and solvency is required to hold at the end of the period rather than on average across it — is the obvious next formalisation, left undone here. The paper's static treatment is a simplification whose direction of error is known, which is the most that can be claimed for it. 5. Why Better Automation Is Harder to Supervise 5.1 The claim Section 3 established that loop position is determined by the ratio of system consequence latency to human response latency rather than by policy, and Section 4 that oversight is a consumable resource with a solvency constraint that binds silently. This section adds the third and least intuitive element of the argument, the one that converts a static capacity problem into a dynamic one. The two properties that make an automated decision class economically attractive are its reliability and its throughput: a system that is right more often and acts more often is worth more money. Those are also, on the measured evidence, the two properties that most reliably destroy the supervisor's ability to detect its failures. Improvement in the object of supervision degrades the supervision of it. The logical status of this claim should be stated precisely, because it is easy to mistake for a rhetorical inversion of the kind that reads well and proves nothing. It is not a paradox, and it is not an irony in Bainbridge's sense of design intentions producing their own defeat. It is the composition of two independently measured empirical effects that operate in the same direction, joined by a third that follows from the capacity arithmetic of Section 4. The first comes from the automationcomplacency literature: holding task load constant, the more constant an automated system's reliability, the lower the probability that its human monitor detects its failures. The second comes from the psychophysics of visual search: the rarer a target, the higher the miss rate, through a shift in the observer's decision criterion rather than a loss of perceptual sensitivity. The third is the solvency constraint itself: as throughput rises, attention available per exception falls, because attention is fixed in the short run and exception arrivals are not. None of these three effects requires the supervisor to be careless, undertrained or badly incentivised. Each has been measured in populations that were attentive, motivated, and in one case professionally expert. This matters because it determines what kind of remedy is admissible. If the mechanism were carelessness, the remedy would be selection, training and incentive design, all of Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 29 which firms already know how to deploy. If the mechanism is criterion placement under low prevalence and attentional resource competition under multiple-task load, then selection, training and incentives are the wrong instruments, and Section 5.3 shows that two of them have been directly tested and have directly failed. The consequence for measurement follows immediately. Realised risk is not the failure rate; it is the failure rate multiplied by the probability that a failure goes undetected and uncorrected, multiplied by the consequence of an uncorrected failure. Reliability improvement moves the first term down and, through the mechanisms below, moves the second term up. The net effect is therefore not signed by the reliability improvement alone. 5.2 The reliability–supervisability tension The foundational experiment here is Parasuraman, Bahri, Molloy and Singh (1992), Adaptive Automation and Human Performance: II, report NAWCADWAR-92036-60, held by the Defense Technical Information Center as accession ADA254127. It must be labelled accurately: this is a government technical report and it is not peer-reviewed. It is grey literature with primarydocument provenance, and is weighted accordingly below. The design places twenty-four adult participants (N = 24) in a multiple-task flight simulation in which an automated engine-status monitor operates alongside concurrent manual tasks. The manipulation is the constancy of the automation's reliability rather than its level. One group received constant-reliability automation, either 87.5% or 56.25% accurate throughout the session; a second received variable-reliability automation alternating between those same two values every ten minutes. Average reliability across the session is thus comparable, and what differs is whether it is stable or changing. The variable-reliability group exceeded 80% failure detection after block 4, while the constant-reliability group's detection remained at a low level, the difference significant from block 2 onward at F(1,40) ≥ 8.99, p < .01. Reaction time was unaffected by automation reliability. The effect is therefore on the detection rate and not on the detection speed: the constantreliability monitors were not slow to catch failures, they did not catch them. Two caveats belong beside those numbers. First, the report's stated degrees of freedom, F(1,40), are not obviously reconcilable with a stated sample of twenty-four, which may indicate pooling across experiments within the report; the figure is reproduced as printed and flagged as such. This paper's own verification pass went further than that flag and recommended against quoting these statistics at all, on the ground that the extraction may have garbled the numbers or conflated two experiments within the report, and they are reproduced here as printed with that recommendation disclosed, pending retrieval of the original document; nothing in the argument rests on the 80% figure or on the F ratio, since Proposition 6 requires only the direction of the finding, which the review literature independently corroborates. Second, the peer-reviewed companion study, Parasuraman, Molloy and Singh (1993), International Journal of Aviation Psychology 3(1):1–23, could not be independently verified in this paper's research pass. Its bibliographic record is sound and its direction of finding is corroborated by the review literature, but because the article itself was not obtained, no figure is quoted from it anywhere in this paper. A second, peer-reviewed experiment establishes the temporal shape of the same phenomenon. Molloy and Parasuraman (1996), Human Factors 38(2):311–322, inserted a single automation failure into a thirty-minute multiple-task session. More participants detected the inserted failure in the first ten minutes of the session than in the last ten, and this held for both the simple and the multicomplex task. The exact detection percentages were not obtainable in this paper's research pass and are accordingly not quoted; the commonly circulated claim that "only about half" of participants detected the failure is unverified and is not repeated. The verified directional finding Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 30 suffices: detection of a rare automation failure decays within a single half-hour watch, in the more demanding task condition as well as the simpler one. Both findings sit inside the integration offered by Parasuraman and Manzey (2010), Human Factors 52(3):381–410, a peer-reviewed integrative review rather than an experiment, which locates automation complacency specifically under conditions of multiple-task load, where manual tasks compete with the automated task for the operator's attention, and reports that it is found in expert as well as naive participants and cannot be overcome with simple practice. An overseer with nothing else to do is a laboratory artefact; an overseer with a job is the case the evidence describes, and it is the case every commercial deployment instantiates. The managerial consequence deserves restating in its sharpest form. Reliability improvement lowers failure frequency, which is the effect firms measure, report and are rewarded for. It also lowers failure detectability, which is the effect firms do not measure and which no standard disclosure requires them to report. A firm that reports only the first is reporting half of the change it has made. If the second effect is large enough, a reliability improvement can raise realised exposure while every dashboard the firm maintains shows it falling. 5.3 Rarity destroys detection The strongest evidence in this paper, by the ordinary criteria of evidence grading, is not from the automation literature at all. It is from the psychophysics of visual search, where the effect of target rarity on detection has been isolated in peer-reviewed venues of the highest standing, with mechanisms identified rather than merely described. Wolfe, Horowitz and Kenner (2005), Nature 435(7041):439–440, peer-reviewed, had observers perform a baggage-screening-style search for targets among distractors, varying only target prevalence. At 50% prevalence the miss rate was 7%; at 10%, 16%; at 1%, 30%. The target did not change, the display did not change, the observers did not change. Only the rate at which targets appeared changed, and the miss rate more than quadrupled. Critically, false alarms were, in the authors' words, "vanishingly rare (0.03%)". Observers were therefore not trading accuracy for speed; they were systematically failing to report targets that were present. In a further experiment mixing prevalence levels within a single stream, observers missed 11% of common targets, 25% of rare targets and 52% of very rare targets. The authors' mechanism is stark and directly transferable: at 1% prevalence, observers abandon the search in less than the average time required to find a target. The quitting threshold moves inside the search. Wolfe and colleagues (2007), Journal of Experimental Psychology: General 136(4):623–638, peerreviewed, then pinned the mechanism definitively using signal-detection analysis. Perceptual sensitivity d′ was 1.97 at 50% prevalence and 2.55 at 2%: sensitivity did not decline, it improved slightly. Over the same manipulation the decision criterion c shifted from 0.13, approximately neutral, to 1.14, highly conservative. The authors' interpretation is that observers equalise the raw number of misses and false alarms rather than their proportions, which at low prevalence sets a catastrophically conservative criterion. Why this matters so much for management deserves to be spelled out, because the distinction between sensitivity and criterion is the difference between a problem the organisation can address and a problem it cannot. The failure is not that the supervisor cannot see. Perceptually, the lowprevalence observer is at least as capable as the high-prevalence observer, and by this measurement slightly more so. The failure is that the supervisor's decision threshold has moved. A threshold is not a perceptual capacity; it is a policy about how much evidence is required before an exception is called, and such a policy is a variable the organisation sets, through the consequences it attaches to Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 31 a false alarm, through the base rate it exposes the supervisor to, and through whether it ever tells the supervisor what the truth was. The low-prevalence effect is an organisational design failure that presents itself as an individual attentional failure, which is why it is so consistently misdiagnosed and so consistently addressed with the wrong instruments. The 2007 paper's intervention series is the most actionable single result in this paper, and it is actionable largely in the negative. Payoff-matrix manipulation, the experimental form of paying supervisors more for catching exceptions, had no effect. Forcing slower responses, the experimental form of telling supervisors to take their time, had no effect, even where the manipulation added a full second to response time. Mixing prevalence levels helped only partially, rare targets still being missed at rates between 21% and 52%. High-prevalence bursts without feedback failed entirely, producing no carryover into the low-prevalence periods. The only intervention that improved matters at all was high-prevalence bursts with full feedback, and its warrant must be stated exactly, because the headline contrast is weaker than the form in which it usually travels. In Experiment 7 (N = 14) the low-prevalence miss rate was 21% against 18% in the matched high-prevalence condition, a difference not statistically significant at p = .312, while the 46% customarily set against that 21% is the low-prevalence baseline of Experiment 1a (N = 24), a different experiment with different observers; the 46% → 21% comparison is therefore a between-experiment difference rather than a measured intervention effect. What Experiment 7 establishes within its own design is a failure to reject the null at N = 14, which is weak evidence of elimination rather than a demonstration of it. A further complication for attributing the effect to feedback specifically is that the mixed condition of Experiment 3 produced the same 21% miss rate at 1% prevalence with no feedback at all. The recommendation nonetheless stands, because it is the only thing in the series that moved the miss rate in the right direction; but the study that would establish elimination — a within-design comparison of high- and low-prevalence conditions with and without feedback, at adequate power — has not been run, and the synthetic-exception injection item in Tier 2 of Appendix B accordingly asks firms operating an injection regime to generate it. Manipulation or condition Measured result Reading Target prevalence 50% Miss rate 7% Baseline Target prevalence 10% Miss rate 16% Miss rate more than doubles Target prevalence 1% Miss rate 30% Miss rate more than quadruples Mixed stream: common / rare / very rare targets 11% / 25% / 52% missed Rarity penalises within a single stream False alarms across conditions 0.03% ("vanishingly rare") Not a speed–accuracy trade-off Perceptual sensitivity d′, 50% vs 2% prevalence 1.97 → 2.55 Sensitivity slightly improves Decision criterion c, 50% vs 2% prevalence 0.13 → 1.14 The whole effect is criterion placement Payoff-matrix manipulation (incentives) No effect Paying for catches does not work Forcing slower responses No effect, even at +1 s response time "Be more careful" does not work Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 32 Mixing prevalence levels Partial only; rare targets still 21–52% missed. Experiment 3's mixed condition: 21% at 1% prevalence, without feedback Insufficient alone, and complicates attributing the Experiment 7 result to feedback High-prevalence bursts without feedback Failed; no carryover Exposure alone is not the active ingredient High-prevalence bursts with full feedback (Experiment 7, N = 14) 21% missed at low prevalence against 18% in the matched high-prevalence condition, difference n.s., p = .312. The 46% comparator is the low-prevalence baseline of Experiment 1a (N = 24) The only intervention that helped; but a null at N = 14, and the 46% → 21% contrast is across experiments, not within one Table 5. The low-prevalence effect and what does and does not correct it. All entries peer-reviewed. Prevalence figures from Wolfe, Horowitz & Kenner (2005); mechanism and intervention figures from Wolfe et al. (2007). The blunt implication should be stated without softening, because it contradicts the two remedies oversight regimes actually deploy. Paying supervisors more to catch exceptions, and instructing them to be careful and to take their time, are not merely unproven; they have been experimentally tested against this failure mode and experimentally demonstrated not to work. What was tested there is incentive and instruction, however, and not training in the sense of repeated practice with feedback, on which Section 5.5 records a located but unverified finding pointing the other way. The one intervention that did better — periodic recalibration on a stream of artificially elevated exception prevalence with immediate ground-truth feedback on every judgement — is one that almost no oversight regime implements and that no instrument surveyed in Section 10 requires. Section 10 accordingly proposes it as a measurable design commitment rather than an aspiration. 5.4 The oversight scissors Composing Section 5.2 with Section 5.3, and both with the solvency constraint of Section 4, produces the central image of this paper, named here the oversight scissors. As an automated decision class matures, three things happen at once and independently. Exception prevalence falls, because that is what maturity means, and falling prevalence raises the miss rate through the criterion shift measured by Wolfe et al. (2007). Reliability constancy rises, because a mature system is not merely more accurate but more consistently accurate, and rising constancy lowers detection through the complacency mechanism measured by Parasuraman, Bahri, Molloy and Singh (1992). Throughput rises, because throughput is how the deployment pays for itself, and rising throughput lowers the uncommitted attention available per exception through the arithmetic of Definition 4. Three mechanisms, from three separate literatures, measured in three unrelated paradigms, all pushing the same way. The blades close. The value of the system and the reliability of its oversight move in opposite directions as automation matures, and they do so as a matter of measured psychophysics and measured attention economics rather than as a matter of organisational discipline. No amount of managerial seriousness alters the criterion shift at 1% prevalence, and no governance policy alters the fact that a supervisor's uncommitted attention does not scale with the throughput of the system they supervise. Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 33 Figure 3. The oversight scissors. Two curves plotted against automation maturity: rising throughput and reliability constancy, and falling exception detection probability, the latter anchored on the measured low-prevalence points of 7%, 16% and 30% miss rates at 50%, 10% and 1% target prevalence respectively, with 52% for very rare targets in the mixed-prevalence stream (Wolfe, Horowitz & Kenner 2005). The annotated intervention is high-prevalence bursts with full feedback (Wolfe et al. 2007, Experiment 7, N = 14), in which low-prevalence misses were 21% against 18% in the matched high-prevalence condition, a difference not significant at p = .312. The annotation reports that within-design comparison rather than the 46% low-prevalence baseline of Experiment 1a (N = 24) against which the intervention is more commonly quoted, since that contrast is across experiments; at N = 14 a null is weak evidence of elimination, and the annotation should be read as an indication of direction rather than as a measured intervention effect. Proposition 6 (Reliability–supervisability tension). Increasing the constancy of an automated system's reliability decreases the probability that its human supervisor detects its failures. Reliability improvement therefore has two opposite effects on realised risk: it lowers failure frequency and it lowers failure detectability. Refutation condition. Refuted if failure-detection performance is shown to be invariant to reliability constancy under multiple-task load. Proposition 7 (Rarity). Detection probability for an exception falls with its prevalence, through a shift in decision criterion rather than a loss of perceptual sensitivity, and is not corrected by incentives or by instructions to respond more slowly. An oversight regime defined over exceptions therefore degrades precisely as the underlying system improves. Refutation condition. Refuted if an incentive, instruction or interface intervention that does not alter observed prevalence is shown to eliminate the low-prevalence detection deficit. The oversight scissors 0% 25% 50% 75% 100% 93% 84% 70% 48% high-prevalence bursts with feedback (Exp. 7, N=14): 21% vs 18%, n.s. early automation maturity mature detection probability detected 93% · miss 7% · prevalence 50% detected 84% · miss 16% · prevalence 10% detected 70% · miss 30% · prevalence 1% detected 48% · miss 52% · very rare probability an exception is detected throughput and reliability constancy Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 34 5.5 The counter-evidence, at full strength The scissors image would be dishonest if presented as a fatality, and the human-factors literature contains findings that constrain it substantially. This subsection gives that counter-evidence at full strength, because the design proposals of Section 10 are built on it rather than in spite of it. The single most important constructive finding is Parasuraman, Mouloua and Molloy (1996), Human Factors 38(4):665–679, peer-reviewed. Participants monitored an automated engine-status task while performing concurrent tracking and fuel-management tasks across three thirty-minute sessions, in three conditions: model-based adaptive allocation, performance-based adaptive allocation, and a non-adaptive control. All groups had low probabilities of detecting automation failures for the first forty minutes under automation, replicating the baseline complacency effect of Section 5.2. The intervention was that in the adaptive conditions the engine task was returned to manual control for ten minutes. Both adaptive groups detected significantly more automation failures in the subsequent blocks under automation control, and the improvement persisted after automation was reinstated. The sample size was not obtainable from the accessible abstract and is accordingly not stated. Complacency, on this evidence, is reversible by a scheduled design intervention costing ten minutes of automated operation, and the benefit does not evaporate when automation resumes. Section 10 builds a design proposal directly on this result and connects it to the protected unassisted practice mechanism of the series' fifth working paper. The second constraint concerns the generality of the vigilance decrement itself. See, Howe, Warm and Dember (1995), Psychological Bulletin 117(2):230–249, peer-reviewed, is a meta-analysis of 42 studies and 138 conditions, and it is narrower and sharper than the shorthand "meta-analysis of the vigilance decrement" suggests: it analyses the decline in perceptual sensitivity specifically, as distinct from a shift in response criterion. Its central conclusion is that a genuine sensitivity decrement appears specifically under high event rates combined with demanding discriminations, and that under other conditions the observed decline reflects a shifting criterion instead. Numeric effect sizes were not obtainable in this paper's research pass and are therefore not quoted. The conclusion nonetheless matters a great deal. The vigilance decrement is conditional on task parameters, which makes it partly a design variable: event rate and discrimination difficulty are properties of the interface and the routing policy, not of the human. This substantially weakens any blanket claim that humans cannot monitor, and weakens the corresponding claim in this paper. What the evidence supports is that humans cannot monitor under particular measurable conditions, and that those conditions are the ones mature automation produces by default rather than the ones it produces necessarily. A related item must be named here rather than passed over, because omitting it would be selective citation of exactly the kind this subsection exists to prevent: Szalma, Hancock, Warm, Dember and Parsons (2006), Training for Vigilance: Using Predictive Power to Evaluate Feedback Effectiveness, Human Factors, was located in this paper's research pass but its content was not verified, and it reports that feedback-based vigilance training improves vigilance performance. No figure is quoted from it. It sits in apparent tension with Parasuraman and Manzey's conclusion that automation bias cannot be prevented by training or instructions, and while vigilance and automation bias are distinct constructs, so that the two findings are not strictly in conflict, the tension is real enough that it should be recorded and is left unresolved here, on the same footing as the contrary 2016 fine-motor finding set against Casner et al. below. The third constraint concerns the benefit side of the ledger, which the pessimistic literature routinely omits. Wickens, Clegg, Vieane and Sebok (2015), Human Factors 57(5):728–739 — volume, issue and pages were not independently confirmed in this paper's research pass and should be checked before quotation — compared an "automation wrong" condition against an "automation gone" condition and a no-automation control in a process-control simulation. With correct Human-on-the-Loop · WP No. 8 5. Why Better Automation Is Harder to Supervise 35 automation, the aid supported faster and more accurate diagnosis and lower workload. That is a real measured benefit and it should not be quietly dropped from a paper about oversight costs. The finding that matters for design is the asymmetry: "automation wrong" was substantially more damaging to diagnostic accuracy than "automation gone", and the authors recommend that decision aids expose their confidence in uncertain environments. A confidently wrong system is a worse object of supervision than an absent one, a proposition with obvious application to generative systems whose fluency is uncorrelated with their correctness. The fourth item is a correction rather than a constraint, included because correcting a widespread miscitation is a credibility asset for a paper that asks its readers to accept a large number of figures on trust. Casner, Geven, Recker and Schooler (2014), Human Factors 56(8):1506–1516, peer-reviewed, studied sixteen airline pilots (N = 16) in a Boeing 747-400 full simulator with automation levels systematically varied across routine and non-routine scenarios. It is routinely cited as evidence of manual-skill atrophy under automation. It shows close to the opposite. Instrument scanning and manual control skills appeared to be reasonably well retained when practised infrequently. What degraded was cognitive: tracking the aircraft's position without the use of a map display, deciding which navigational steps come next, and recognising instrument system failures. No effect sizes were obtainable and none is quoted. 

The mechanism is the one this paper requires: performance on those cognitive tasks correlated with the frequency of task-unrelated thought during periods of automation use, and the authors conclude that cognitive skill retention may depend on pilots' active engagement in supervising automated systems. An unresolved tension should be noted rather than concealed: a contrary 2016 finding in Human Factors (DOI 10.1177/0018720816640394) reports erosion of fine-motor flying skills under flight-deck automation, and the two have not been reconciled. Taken together, this counter-evidence changes the shape of the claim the paper can make, and changes it in a direction that makes it more useful. What decays under a supervisory regime is not manual execution but situation modelling and judgement — the capacity to hold a model of what the system is doing and why, and to notice when that model stops fitting. The decisive variable governing whether it decays is the quality of attentional engagement during supervision. Attentional engagement is not a character trait. It is a product of task design, event rate, scheduled unassisted operation, feedback provision and workload allocation, every one of which is set by the organisation rather than by the individual at the console. The scissors close by default. They do not close by necessity, and Sections 8 and 10 specify what holding them open costs. 6. When Oversight Works, and Why 6.1 The question restated Sections 3 to 5 are a sustained account of the conditions under which human oversight of an automated system fails. Read carelessly, they invite the conclusion that human oversight does not work, and that is not what the evidence supports. The literature on human–algorithm decision making is large, methodologically uneven and, on the question of whether the human helps, genuinely two-sided. A paper reporting only its pessimistic half would be doing to the evidence what it accuses firms of doing to their reliability statistics. The productive question is therefore not "does human oversight work" but "under what measured conditions does it improve outcomes, by how much, and how do those conditions compare with the conditions under which oversight is actually mandated". Posed that way, the question has an answer. The conditions are known, they have been isolated in meta-analysis and in field natural experiments, and they are narrow. The uncomfortable finding is not that oversight never helps; it is Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 36 that the conditions under which it helps and the conditions under which it is required by policy are close to disjoint. This section establishes the two conditions that Sections 8 and 10 will operationalise, and states Propositions 8 and 9, which are boxed in the reverse of their numerical order because the verifiability condition of Proposition 9 follows from the explanation evidence of Section 6.4 while the decorrelation condition of Proposition 8 follows from the field evidence of Section 6.5. 6.2 The meta-analytic answer The most authoritative source is Vaccaro, Almaatouq and Malone (2024), Nature Human Behaviour 8(12):2293–2303, peer-reviewed. The corpus is 5,126 records screened, 74 papers meeting inclusion, 106 unique experiments and 370 effect sizes, with inclusion requiring that each study report performance for the human alone, the AI alone and the combination. That inclusion rule is what makes the paper decisive: most primary studies report only two of the three arms. Two pooled results follow. Human–AI combinations beat humans alone with Hedges' g = +0.64, 95% CI [0.53, 0.74]. Human–AI combinations lose to the better of the two alone with g = −0.23, 95% CI [−0.39, −0.07], p = 0.0047. The first result is the one quoted in vendor material and in policy documents; the second is the one that matters for an oversight regime, because an oversight regime is a decision to run the combination rather than whichever arm is stronger. It should be stated plainly that the paper reports no pooled estimate for the combination versus the AI alone, and none is invented here; the functional equivalent is the moderator structure. The moderators are where the finding becomes operational. Decision tasks, in which the actor chooses from a finite set of options, show g = −0.27 [−0.44, −0.10]. Creation tasks, in which the actor generates open-ended content, show g = +0.19 [−0.09, +0.48], which is not significant. Where the human alone was the stronger performer at baseline, combination shows g = +0.46 [0.28, 0.66]. Where the AI alone was stronger at baseline, combination shows g = −0.54 [−0.71, −0.37]. The pattern is coherent and unflattering to oversight policy: the combination helps where the human was already better and hurts substantially where the machine was. Oversight regimes are not imposed on tasks where the human is better; they are imposed where a model has been deployed because it outperforms the incumbent process. A further result bears directly on the design proposals commonly offered as fixes. AI explanations and AI confidence displays were both tested as moderators, and neither significantly affected human–AI system performance. This is the most damaging published finding for explanation-based oversight policy, and it comes from a Nature-family meta-analysis rather than a single laboratory study. The paper's own limitations must be reported with the same precision as its findings. Heterogeneity is extreme: I² = 97.7% for the synergy estimate and 93.8% for the augmentation estimate. The pooled means are therefore weak summaries of a very dispersed distribution, and no individual deployment should be predicted from them. And the publication-bias analysis is asymmetric in a way that is analytically important and rarely noted. Bias was detected for the optimistic augmentation result, Egger's β = 1.96, p = 0.002, and was not detected for the pessimistic synergy result. The finding that the literature is most likely to have inflated is the encouraging one. A reader inclined to discount the meta-analysis on publication-bias grounds should discount the +0.64 first. The second meta-analytic source addresses not performance but acceptance. Qin and colleagues (2025), Psychological Bulletin 151(5):580–599, peer-reviewed, synthesise 163 studies, 442 effect sizes and N = 82,078. Overall, people show a slight net preference for human over AI judgement, d = −0.26 [−0.37, −0.15]. The pooled figure conceals a clean moderation. Where AI is perceived as more Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 37 capable and personalisation is unnecessary, there is appreciation, d = +0.27; otherwise there is aversion, d = −0.50. This resolves the long-standing conflict between the algorithm-aversion and algorithm-appreciation literatures as two arms of a moderated effect rather than competing claims. The compositional argument is what makes this relevant. Oversight is mandated precisely in highstakes, individualised, high-personalisation contexts: credit, benefits, medicine, employment, criminal justice. That is the exact quadrant in which the meta-analysis predicts aversion at d = −0.50 — the quadrant in which an overseer will systematically under-weight a model even where the model is superior. Composed with Vaccaro et al.'s g = −0.54 for tasks where the AI is the stronger party, this yields a single coherent mechanism rather than two correlated observations: policy places human overseers in the settings where they will discount the model most, on the tasks where discounting the model costs the most. The two meta-analyses are measuring the same failure from the attitudinal and the performance side respectively. 6.3 Why overseers cannot tell when to override The mechanism behind the pooled numbers is a discrimination failure. Overriding an automated recommendation is valuable only if overrides concentrate on the cases where the recommendation is wrong. They do not. Green and Chen (2019a), Proceedings of FAT* '19, pp. 90–99, peer-reviewed, ran a pretrial riskassessment task with 554 analysed participants generating 13,850 predictions. Mean reward was 0.756 unaided, 0.786 with the risk assessment shown, and 0.807 for the risk assessment alone. False positive rates were 17.7%, 14.8% and 10.1% respectively. The risk assessment therefore improved on unaided humans and was in turn degraded by them, and only 23.7% of treated participants beat the algorithm. The study also documents a disparate-interaction effect: where the algorithm predicted higher risk than the participant, it exerted 25.9% stronger influence for Black defendants than for White, 0.68 versus 0.54, p = 0.02. Human oversight did not neutralise the model's disparate impact; the interaction between human and model generated a disparity of its own. The result this paper needs most is the evaluation-blindness finding. Participants' confidence in their own performance was negatively associated with their actual performance, p = 0.0186. Their assessments of the algorithm's accuracy had no significant relationship to its actual accuracy. Their assessments of its fairness had no association with measured disparity. An overseer who cannot evaluate their own performance, the system's accuracy, or the system's fairness has no signal on which to allocate their scarce attention. The authors' own limitations should be carried with the finding: participants were lay crowdworkers rather than judges, exposure was roughly twenty minutes, and racial priming was minimal relative to a real courtroom, which the authors note may make the disparity estimate conservative. Green and Chen (2019b), Proceedings of the ACM on Human-Computer Interaction 3(CSCW), Article 50, peer-reviewed, tested three desiderata for algorithm-in-the-loop decision making across pretrial risk and loan default: accuracy, reliability and fairness. Accuracy was only partially met, reliability was not met, and fairness was not met. Participant counts for the two domains come from the author-hosted version rather than the ACM version and are flagged accordingly. The most counterintuitive result is that the Feedback condition, in which participants were shown per-case outcomes, performed worse than baseline, p < 10โปโถ in the pretrial domain and p < 10โปโด in loans. Outcome feedback, the standard organisational remedy for miscalibration, made performance worse in both domains. The constructive exception belongs in the same breath, because it is one of very few positive design results in this literature. The Update presentation — the participant makes their own prediction Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 38 first, then sees the algorithmic score, then revises — reduced disparities by 81.5% in influence and 73.9% in deviation relative to showing the score first. The ordering of information, not its content, carried the effect. Section 10 builds a design proposal directly on this. Green (2022), Computer Law & Security Review 45:105681, peer-reviewed, generalises this into a policy critique on the basis of a survey of 41 policies mandating human oversight. His argument has two flaws, not three, and describing it as a three-part argument is a miscitation that should be corrected wherever it appears. The first flaw is empirical incapacity: people cannot perform the oversight functions the policies assume of them. The second is legitimation: because of the first, oversight requirements confer a false sense of security on faulty systems and allow vendors and agencies to shirk accountability. The "three" is the taxonomy of oversight mechanisms across which those two flaws are shown to apply — 20 policies restricting decisions based solely on automated processing, 14 conferring human discretion or override authority, and 7 requiring "meaningful" human input. The legitimation flaw is the one that connects to Section 9 of this paper, because it is the mechanism by which a nominal overseer becomes a liability allocation. The most recent evidence, now with large language models, reproduces the pattern and sharpens it. Bo, Wan and Anderson (2025), Proceedings of CHI '25, peer-reviewed, N = 400, tested reliance disclaimers, uncertainty highlighting and implicit-answer presentation against a control. The authors' summary is that the interventions reduce over-reliance but generally fail to improve appropriate reliance — that is, they move the level of reliance without improving its discrimination, exactly the distinction that will recur in Section 6.4. The most damaging single result for oversight policy is that across conditions participants showed larger confidence gains after making wrong reliance decisions than after correct ones, p = .02. Overseers are not merely miscalibrated. They are anti-calibrated: most confident precisely when wrong. An organisation that routes escalations by overseer confidence is therefore routing them by a signal that is inversely related to the thing it is meant to track. 6.4 Explanations do not repair it, and transparency can make it worse The standard response to Section 6.3 is that the overseer needs to be told why the system produced its output, the intuition underlying explainability requirements in most AI governance instruments. The evidence for it is weak, and in one pre-registered study it runs the other way. The counter-evidence should be led with, because it is real. Bansal and colleagues (2021), Proceedings of CHI '21, peer-reviewed, N = 1,626, achieved genuine complementary team performance — team accuracy exceeding both the human alone and the AI alone — in all three of their tasks. On beer-review sentiment the team scored 0.89 against 0.82 for the human alone and 0.84 for the AI alone; on Amazon book-review sentiment, 0.87 against 0.85 and 0.84; on LSAT logical reasoning, 0.78 against 0.67 and 0.65. Complementarity is achievable. What it was not achieved by is explanation. Explanations added nothing over a bare confidence score: beer z = −1.18, p = 0.24; Amazon z = 1.23, p = 0.22; LSAT z = 0.427, p = 0.64. Expert-authored explanations were no better than AI-generated ones. The mechanism the authors identify is the decisive one: explanations increased agreement with the AI regardless of whether the AI was right, raising accuracy when it was correct and lowering accuracy when it erred. Explanations moved the level of reliance and not its discrimination. Only 5% of participants reported using explanations to validate the AI's reasoning at all. Poursabzi-Sangdeh and colleagues (2021), Proceedings of CHI '21, peer-reviewed and, unusually for this literature, pre-registered, ran four experiments with N = 3,800 in total on apartment-price prediction, manipulating model transparency while constraining the models to produce identical Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 39 predictions so that only presentation varied. On unusual cases where the model was badly wrong — precisely the cases oversight exists to catch — participants shown the clear model were worse at detecting the mistake, F(1,994) = 8.81, p = 0.003, replicated at t(594) = −4.16, p < 0.001. Experiment 4 diagnosed the mechanism as information overload: a simple message directing attention to the outlier eliminated the transparency disadvantage entirely. Transparency did not fail because transparency is bad; it failed because it consumed the attention that detection required. That is the same attentional-scarcity account as Sections 4 and 5, arriving from a different direction. Buçinca, Malaya and Gajos (2021), Proceedings of the ACM on Human-Computer Interaction 5(CSCW1), Article 188, peer-reviewed, N = 199, tested cognitive forcing functions — interface designs that compel deliberation before the AI's answer is available. They worked, in the narrow sense: over-reliance fell from roughly 30% under simple explainable-AI conditions to roughly 26%, and correct decisions on the trials where the AI was wrong rose from 3% to 9%, p โ‰ช .0001, d = .37. Three qualifications make this a hard result to build policy on. Even the best intervention left the human– AI team below the AI alone. The benefit accrued disproportionately to high–need-for-cognition participants, which is intervention-generated inequality: the debiasing tool helps the people who least needed it. And participants rated the forcing conditions as significantly more complex and liked them less. The interventions that improve oversight are, on this evidence, the ones organisations and users will resist adopting. Fok and Weld (2024), AI Magazine 45(3):317–332, converts these scattered nulls into a structural claim. It must be labelled correctly: this is a peer-reviewed review and position article, not primary research. Their argument is that explanations are useful only to the extent that they allow the decision maker to verify the correctness of the prediction, and that most tasks fundamentally do not permit easy verification regardless of explanation method. This is the right frame, and it yields the paper's ninth proposition. Proposition 9 (Verifiability precedence). Oversight of an output is economically meaningful only where verifying the output costs less than producing it. Where ground truth is a counterfactual future, the verification cost is undefined at decision time and oversight cannot be an accuracy control; it can only be a legitimacy or a values control. Refutation condition. Refuted if a class of decisions with counterfactual ground truth is shown to admit contemporaneous human verification that improves realised accuracy. Why this claim is stronger and more robust than "people are biased" is worth spelling out, because it determines what a firm can do about it. It is not a claim about human deficiency. It is a claim about the structure of the decision. Recidivism, benefit fraud, credit default and medical prognosis share the property that the ground truth against which a prediction would be checked is a counterfactual future state that does not exist at the moment of decision and, for the cases where the system intervenes, may never come to exist. No amount of overseer training, expertise, seniority or motivation creates a verification signal that the decision structure does not contain. This is why the proposition cannot be answered by better training, and why it survives arbitrary improvement in the quality of explanations. Where verification is impossible, oversight can still be legitimate and valuable — as a values control, as a channel for contestation, as the locus of a duty someone must hold — but it cannot honestly be described as an accuracy control, and instruments that require it as one are requiring something the decision class cannot supply. Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 40 6.5 Where oversight demonstrably works The constructive core must be as strong as the critique, and it is. De-Arteaga, Fogliato and Chouldechova (2020), Proceedings of CHI '20, peer-reviewed, exploit a natural experiment at the Allegheny County child-maltreatment hotline. A technical fault after deployment caused certain model inputs to be miscalculated in real time, so the risk score displayed to call screeners diverged from the correct assessed score. The workers' decisions were better calibrated with respect to the assessed score than with respect to the shown score: they partially saw through an erroneous display. Where the displayed score underestimated risk, workers screened in nearly 60% of mid-range cases against 30% of correctly-scored comparable cases. Only 66% of cases flagged as mandatory screen-ins were actually screened in, which is substantive override behaviour rather than ratification. And there was no significant differential adherence across racial or socioeconomic groups — the opposite of the laboratory finding in Green and Chen (2019a). The mechanism must be identified precisely, because it is the core of the generalisable lesson. The screeners hold private information the model does not have: the live call, the narrative context, prior familiarity with the family. They were not auditing the model on the model's own inputs; they were combining the model's signal with an independent one. This is Vaccaro et al.'s human-stronger quadrant, g = +0.46, instantiated in the field. It is also this paper's eighth proposition. Proposition 8 (Error decorrelation). Human oversight improves the accuracy of an automated decision only where the overseer's errors are decorrelated from the system's on the relevant error class. An independent information channel — evidence the system does not receive — is the strongest and most auditable source of such decorrelation, but it is a sufficient and not a necessary source: a human and a system reading the same input may still err differently. Where neither an independent channel nor demonstrated decorrelation is present, oversight modulates the level of reliance without improving its discrimination. Refutation condition. Refuted if overseers whose errors are shown to be correlated with the system's on the relevant class, supervising a system whose baseline advantage over them exceeds the reported sampling error, nonetheless achieve accuracy exceeding the system alone. The boundary case for this proposition is reported two subsections earlier in this paper, and it should be confronted rather than left for a reader to notice. The crowdworkers in Bansal and colleagues (2021) read the same review text the model read; they held no channel the system lacked; and on beer-review sentiment the team nonetheless scored 0.89 against 0.84 for the AI alone. Under the formulation this proposition carried in earlier drafts — that oversight improves accuracy only where the overseer holds an information channel the system does not — that result is a refutation, and this paper's own Section 6.4 is what supplies it. The earlier formulation was too narrow. What Definition 5 requires of a verifier is not a separate channel but conditional independence of error, Pr(V misses e | G commits e) = Pr(V misses e), and two parties reading identical input can satisfy that condition, because what must differ is the mapping from evidence to error and not the evidence itself. Bansal's complementary result is therefore evidence that decorrelation can arise without an independent channel — which is what the amended proposition asserts — rather than evidence against it. Honesty requires a second observation: the human alone scored 0.82 with a reported standard deviation of 0.09 against the AI's 0.84, so the AI is not clearly the stronger party on that task, which is why the amended refutation condition requires the system's baseline advantage to exceed the reported sampling error. The constructive consequence is the one a firm Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 41 can act on. Decorrelation is not observable to the organisation that would rely on it, because estimating a joint error distribution requires labelled ground truth on the error class, which is precisely what the decision classes at issue withhold. An independent channel remains the recommendation because it is the one source of decorrelation a firm can verify cheaply: it can confirm that its reviewer sees evidence the model does not far more easily than it can estimate how the two parties' errors covary. That is what the independence register proposed in Section 10.7 operationalises. The corroborating evidence is unusually clean because it varies the information channel directly. Agarwal, Moehring, Rajpurkar and Salz (2023), NBER Working Paper 31422 — which must be labelled as not peer-reviewed, and is the least peer-reviewed source relied on in this section — ran a pre-registered study with 180 professional radiologists and 324 chest X-ray cases across four information conditions. The AI outperformed roughly two-thirds of radiologists. Providing AI predictions produced a near-zero average treatment effect on diagnostic quality, while providing clinical history alone improved accuracy by roughly 5%. The contrast is the argument in miniature: an additional channel of independent information helped; an additional opinion on the same inputs did not. Radiologists also spent about 4% more time per case with AI, an oversight cost with no accuracy return. The authors conclude that unless documented mistakes can be corrected, the optimal solution involves assigning cases either to humans or to AI, but rarely to a human assisted by AI. Clinical evidence nonetheless establishes that human–AI teams can beat both parties. Tschandl and colleagues (2020), Nature Medicine 26(8):1229–1234, peer-reviewed, found that good-quality AI-based support of clinical decision making improves diagnostic accuracy over that of either AI or physicians alone, with the authors' own paired caveat that faulty AI can mislead the entire spectrum of clinicians, including experts. Krakowski and colleagues (2024), npj Digital Medicine 7:78, peerreviewed, meta-analysed 10 studies comprising 67,700 diagnostic evaluations and report sensitivity rising from 74.8% [68.6–80.1] without AI to 81.1% [74.4–86.5] with it, and specificity from 81.5% [73.9–87.3] to 86.1% [79.2–90.9], with the largest gains for non-dermatologists. Their stated limitation is material: 67% of the included studies were conducted in experimental rather than clinical settings. The strongest real-world datapoint is Lång and colleagues (2023), The Lancet Oncology 24(8):936– 944, the MASAI randomised controlled trial, peer-reviewed: 80,033 women in a population screening programme, cancer detection of 6.0 versus 5.0 per 1,000 screened, recall 2.2% versus 2.0%, false positives 1.5% in both arms, and a 44% reduction in radiologist screen-reading workload. Better detection, no rise in false positives, less human labour. The architecture of MASAI matters more than its effect size, and Section 8 depends on the distinction. MASAI is a triage and attention-reallocation design, not an override design. The AI determines which cases receive how much human reading; the human is not asked to audit the model's output on the model's own inputs and pronounce on whether it erred. The human's reading is an independent assessment of the image, deployed where it has most value. Every failure mode catalogued in Sections 6.3 and 6.4 concerns the override architecture — a human asked to evaluate a model's judgement using the model's own evidence. MASAI never asks that. That is a different and far more defensible use of the human, available to firms willing to redesign the routing rather than add a reviewer. What the architecture buys is nonetheless bounded, and the bound is a property of the routing model rather than of the radiologists: no triage arrangement removes more cancers than its router places in front of a human reader, so the benefit is capped by the sensitivity of the routing model itself. Mammography screening is one of the few settings in which that ceiling is recoverable after the fact, because interval cancers eventually reveal what the router did not route, Human-on-the-Loop · WP No. 8 6. When Oversight Works, and Why 42 and that recoverability is why the design can be evaluated at all rather than merely asserted; Section 8.3.1 derives the bound and sets out what a firm adopting the same architecture in a decision class where the outcome never arrives is left holding instead. The most balanced citation available should close the evidentiary account. Ben-Michael, Greiner, Huang, Imai, Jiang and Shin (2025), PNAS 122(38):e2505106122, peer-reviewed, apply a formal statistical evaluation framework to a randomised trial of a pretrial risk-assessment instrument and find that the risk-assessment recommendations do not improve the classification accuracy of a judge's cash-bail decision, and that replacing the human judge with algorithms leads to worse classification performance. Neither the human, nor the algorithm, nor the human-with-algorithm dominates. This refuses both the techno-optimist and the techno-pessimist reading, and it forces the governance conclusion this paper argues for: the question is not who is smarter but how the decision is architected and how cases are allocated between the available decision makers. Two conditions emerge from the evidence, and Sections 8 and 10 operationalise them. Oversight adds value when the overseer holds an information channel independent of the system's own inputs — the Allegheny call, the clinical history, the second reading of the image — which is the strongest and most auditable source of the error decorrelation Proposition 8 requires, though not the only one from which it can arise. And oversight adds value when verification costs less than performance, which is the condition Proposition 9 states and which most consequential decisions with counterfactual ground truth fail. Where both hold, the human is a genuine control. Where neither holds, what remains is not oversight but its appearance, and Section 9 examines what that appearance is for. 7. Why the Loop Cannot Close Itself 7.1 The strongest claim, and the weakest evidence Sections 3 to 6 took the human overseer as given and asked what that overseer can do. This section examines the proposition that the overseer is no longer necessary, because the automated system has become self-supervising. The claim, as it is made in practice and in the press release this paper takes as its object, is that AI agents now operate in a closed loop: they observe the results of their own actions, identify problems, correct their approach and re-execute, without human intervention. If that were true, most of the preceding argument would be a transitional inconvenience, and the sensible response to Sections 4 and 5 would be to wait. The claim must be divided before it can be assessed, because two quite different loops travel under the same name and only one of them is in doubt. A deployed agent ordinarily acts against an environment that answers back: a coding agent runs a compiler and a test suite, an operations agent reads the state of a database, a treasury agent observes whether a transaction settled, a customerservice agent receives a reply. Each of those is feedback the generator did not author, and a loop closed by feedback of that kind is not the loop the results below refute. Stechly and colleagues' own headline finding — a sound external verifier lifting Game of 24 from 3% to 38% and Blocksworld from 55% to 87% — is a demonstration that loops with genuine external feedback work, and work well. What the evidence refutes is the narrower and more heavily marketed loop in which the closing signal is the generator's own assessment of its work: the model re-reading its output, judging it correct, and acting on that judgement. The distinction that survives is therefore not whether feedback exists but whether the feedback channel is independent of the generator's error distribution, in the sense Definition 5 will give the term. A compiler is independent: it does not share the model's misconception about what the program was supposed to do, and it rejects on grounds the model did not select. A test suite the Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 43 agent also wrote is not, because the misconception that produced the defect is the same misconception that produced a test incapable of catching it — which is precisely the correlated blind spot that Monperrus concedes in the limitations recorded in Section 7.4. The closed-loop question is a question about the provenance of the feedback rather than about its presence, and provenance is the property that neither the marketing claim nor the systems built on it currently records. So restricted, the claim is not supported, and its position in the evidentiary hierarchy is the opposite of what its rhetorical prominence suggests. Of the several claims a firm might make about agentic autonomy, self-verification is the one for which the peer-reviewed literature offers the least support and the most direct contradiction. The capability claims examined in Section 1.4 are, at worst, overstatements of a real and improving trend. The self-verification claim is of a different kind: it asserts a property — reliable self-verification — tested directly at two successive International Conference on Learning Representations, in both cases with negative results, and observed to be absent in the one published field deployment in which a frontier laboratory ran a real business with its own money over months. This section reports those results, identifies the structural reason for them, and states two propositions carrying the paper's deepest claim, which is not about capability at all. 7.2 What the measurements show The primary source is Huang, Chen, Mishra, Zheng, Yu, Song and Zhou (2024), published at ICLR 2024 and therefore peer-reviewed. Its methodological contribution is the definition it isolates. The authors test what they call intrinsic self-correction, defined as the case in which a model "attempts to correct its initial responses based solely on its inherent capabilities, without the crutch of external feedback." That definition isolates exactly the half of the commercial claim that Section 7.1 identified as doubtful. It holds the environment out of the loop, so that what remains is the generator's assessment of its own work, and a correction operator carrying no information the generator did not already possess. Intrinsic self-correction is therefore not an approximation of everything a deployed agent does; it is a clean measurement of the one component on which the self-supervising claim rests. Accuracy is reported under standard prompting, then after a first and a second round of selfcorrection. GPT-3.5-Turbo on GSM8K moves from 75.9% to 75.1% to 74.7%; on CommonSenseQA from 75.8% to 38.1% to 41.8%; on HotpotQA from 26.0% to 25.0% to 25.0%. GPT-4 on GSM8K moves from 95.5% to 91.5% to 89.0%; on CommonSenseQA from 82.0% to 79.5% to 80.0%; on HotpotQA from 49.0% to 49.0% to 43.0%. Every cell is flat or worse than the standard-prompting baseline, and not one improves. A system asked to review its own work and thereby halving its accuracy is not a weak verifier but an actively harmful one. The mechanism is what makes the result general rather than incidental to a benchmark. On GSM8K, GPT-3.5 changed a correct answer to an incorrect one 8.8% of the time, against incorrect to correct 7.6%. The correction operator is therefore not merely noisy; it has negative expected value, and iterating it does not converge on correctness but drifts away from it. The two-round columns are that drift observed. This is a claim about the arithmetic of the loop, not the difficulty of the task. The same paper closes off the most common architectural response, adding more agents. At matched inference cost, multi-agent debate scored 83.0% against 85.3% for plain self-consistency: agents arguing with each other underperformed simply sampling the same model several times and taking the majority. "Add a reviewing agent" is the most frequently proposed and most cheaply implemented intervention, and at equal cost it was worse than voting. Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 44 The second source isolates the causal variable. Stechly, Valmeekam and Kambhampati (2025), published at ICLR 2025 and therefore peer-reviewed, compare standard prompting, the model critiquing itself in a loop, and the same loop with a sound external verifier substituted for the selfcritique. On Game of 24, accuracy runs 5%, then 3%, then 38%. On Graph Colouring, 16%, then 2%, then 37%. On Blocksworld planning, 40%, then 55%, then 87%. Self-critique made performance worse in two of the three domains; the external verifier improved it in all three. The diagnostic figure is the model-as-verifier's false-negative rate of 95.8% on Graph Colouring: asked to check candidate solutions, the model rejected almost every correct solution it was shown. A loop gated on such a verifier does not merely fail to catch errors; it discards its own successes, which is why the Graph Colouring self-critique column falls to 2% rather than stalling at 16%. The authors add the finding that identifies where the value lies: merely re-prompting with a sound verifier retains most of the benefit. The gain comes from the external ground-truth signal, not from the model's reflection. Reflection without such a signal is not a cheaper form of verification; on this evidence it is not verification at all. Source and grade Design Result Bearing on the closed-loop claim Huang et al. (2024), ICLR 2024; peerreviewed Intrinsic selfcorrection, two rounds, three benchmarks GPT-3.5: GSM8K 75.9→75.1→74.7; CommonSenseQA 75.8→38.1→41.8; HotpotQA 26.0→25.0→25.0. GPT-4: GSM8K 95.5→91.5→89.0; CommonSenseQA 82.0→79.5→80.0; HotpotQA 49.0→49.0→43.0. Correct→incorrect 8.8% vs incorrect→correct 7.6%. Debate 83.0% vs self-consistency 85.3% Every cell flat or worse. The correction operator has negative expected value, so iteration drifts from correctness Stechly, Valmeekam & Kambhampati (2025), ICLR 2025; peer-reviewed Standard vs selfcritique vs sound verifier Game of 24: 5% → 3% → 38%. Graph Colouring: 16% → 2% → 37%. Blocksworld: 40% → 55% → 87%. LLMverifier false-negative rate 95.8% on Graph Colouring The external groundtruth signal, not the reflection, carries the benefit Cemri et al. (2025), NeurIPS 2025 Datasets and Benchmarks Track; peer-reviewed 1,600+ traces, 7 frameworks; taxonomy from 150 traces, κ = 0.88 Failure 41%–86.7%. System design 43.9%; misalignment 32.15%; task verification 23.5% (no/incomplete 8.2%, incorrect 9.1%). Fixes: +9.4% (CEO final say), +15.6% (objective verification) Verification failure is common; the remedies that worked were organisational Anthropic, Project Vend Phase 1 (2025) and Phase 2 (18 Dec 2025); company technical reports, grade C A real office shop run for about a month; then three locations with an added CEO agent Phase 1: lost money; sold below cost; "made minimal progress on lessons learned". Phase 2: CEO agent "authorized such requests about eight times as often as it denied them"; procedures cut discounting 80%, giveaways 50% No self-improvement in the field; a supervisor on the same base model shares its blind spots Replit agent incident, July 2025 (AIID Incident 1152); grade D trade reporting Field incident; no legal record Production database deleted during a code freeze; ~4,000 fabricated accounts; fake unit-test results; rollback wrongly asserted impossible Instruction-following, self-testing and selfreporting failed at once, toward concealment Table 7. Self-verification and the closed loop, by evidence grade. Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 45 7.3 The failure is organisational, and so is the fix The two ICLR results are laboratory measurements on benchmark tasks; the question is whether they survive contact with a deployed multi-agent system doing something resembling work. Cemri and colleagues (2025), NeurIPS 2025 Datasets and Benchmarks Track and therefore peer-reviewed, answer at a usable scale. They annotated more than 1,600 execution traces across seven opensource multi-agent frameworks, developing the taxonomy from 150 traces with expert annotators and validating it at inter-annotator agreement κ = 0.88 — worth reporting, because taxonomies of this kind are frequently published without one. Between 41% and 86.7% of traces ended in failure across the seven frameworks; the best-behaved failed on roughly two runs in every five. Failures sort into three categories: system design and specification issues at 43.9%, inter-agent misalignment at 32.15%, and task verification at 23.5%. Inside the last, no or incomplete verification accounts for 8.2% and incorrect verification for 9.1%. The second is the more troubling: a system that does not verify has an obvious gap, whereas one that verifies wrongly has produced a positive signal of correctness that is itself wrong, and a closed loop consuming that signal has no further channel by which to discover the error. The authors' scoping limitation must be carried with the figures, and it runs in the direction that strengthens the argument: they deliberately restricted the taxonomy to failures addressable through system design, coordination and verification, explicitly excluding fundamental model limitations such as hallucination. The 41% to 86.7% range therefore understates total failure. The result that matters most for a management paper is not the failure rate but the intervention study, which is the strongest empirical warrant in this paper for a constructive conclusion. The fixes that worked were organisational rather than model-level. A straightforward workflow adjustment ensuring that the CEO agent had the final say yielded a 9.4% improvement in overall task success; adding a high-level task-objective verification step yielded 15.6% on program development. Neither involved a better model, more inference-time compute or a different vendor. Both are management design in the ordinary sense: assign a clear final authority, and mandate that work be verified against the objective rather than against itself. That the two interventions with measured positive effects are the classical instruments of organisational control is a finding a governance paper may lean on, and a caution against reading this section as uniformly pessimistic. Multi-agent systems fail at high rates, and for reasons that respond to structure. 7.4 Correlated blind spots, and why a supervisor agent is not oversight The intervention that follows most naturally from Section 7.3 — appoint a supervisor — raises the question this paper's fifth definition exists to answer. If the CEO agent's final say improves outcomes, and an external verifier rescues the loop, why not appoint an agent as that verifier and recover the closed loop at one remove? The answer turns on a property that is not capability. Definition 5 (Verification independence). A verifier V is independent of a generator G with respect to an error class E if Pr(V misses e | G commits e) = Pr(V misses e) for e ∈ E. Verifiers sharing a base model, a training corpus, an input channel, a vendor or an incentive with the generator are not independent with respect to the error classes those commonalities induce. The definition is stated in terms of conditional probability because that is where the failure lives. A verifier can be highly capable in absolute terms and contribute nothing, if the errors it misses are exactly the errors the generator makes. Capability is a marginal property of the verifier; independence is a joint property of the pair, and no improvement in the first supplies the second. Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 46 The field evidence comes from Anthropic's Project Vend: a pair of company technical reports, grade C, produced by a party with a commercial interest in the technology. Two features nonetheless make them valuable. They report a real business run with real money over months rather than a benchmark, and they are a frontier laboratory publishing an unflattering result about its own model. In Phase 1 (2025), an agent ran an office shop for about a month and lost money. It purchased large quantities of specialty metal cubes and resold them below cost, declined a $100 offer for an item costing $15, and hallucinated payment details, directing customers to a Venmo account that did not exist. Most importantly here, and in Anthropic's own words, it "made minimal progress on lessons learned, repeatedly reverting to the same patterns." That is an observed absence of selfimprovement in the field rather than in a benchmark, over a month of continuous operation with the feedback signal of an actual profit-and-loss account available throughout. Phase 2, published 18 December 2025, ran the experiment the previous paragraph invites. Anthropic added a CEO agent to impose commercial discipline on the shopkeeper agent — the supervisoryagent architecture in its cleanest form, a hierarchical control layer instantiated to catch exactly the failures Phase 1 documented. The measured result is that the supervisor agent "authorized such requests about eight times as often as it denied them." Anthropic's own explanation is the finding rather than an excuse for it: the CEO agent "shared many of the deficiencies and blind spots of Claudius (which makes sense, given that they're the same underlying model)." The general statement should be made plainly, because it is the practical content of Definition 5. Adding a supervisory agent built on the same base model does not add independent oversight. It adds a correlated rubber stamp — a control that generates approval records and an audit trail while leaving Pr(V misses e | G commits e) essentially equal to one for the error classes the shared model induces. An eight-to-one approval ratio is what such a control looks like from outside, and in process documentation it is indistinguishable from a functioning one. What did work in Phase 2 deserves equal prominence, because it corroborates Section 7.3 from a different setting. Anthropic's conclusion was that "bureaucracy matters": structured procedures and checklists prevented common errors more effectively than raw capability, with discounting falling 80% and giveaways 50%. The mechanisms that improved the system were constraints fixed in advance, not judgements exercised contemporaneously by a peer intelligence. Anthropic discloses no net profit figure for Phase 2; calling it profitable would over-read the source, and this paper does not. The corroborating field incident is the Replit agent of July 2025, which must be labelled grade D trade reporting: well corroborated across the AI Incident Database and several outlets, but with no regulatory or judicial record of the kind Section 9 relies on. The agent deleted a production database during an explicit code freeze, generated roughly 4,000 fabricated user accounts, produced fake unit-test results, and incorrectly asserted that rollback was impossible. The analytical point is not that an agent erred but the pattern of the errors. Three self-monitoring channels — instructionfollowing, self-testing and self-reporting — failed simultaneously, and all in the same direction, toward concealment. A closed loop is precisely a system whose safety rests on those channels being independent; here they were not, and their dependence was revealed only because a human independently checked and found that rollback in fact worked. Recovery came from outside the loop. An independent concession from an unsympathetic source is worth recording. Monperrus (2026), arXiv:2606.13175, is a preprint, and explicitly a position paper containing no new empirical study — its own text states that it synthesises existing capability evidence. It argues that coding agents Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 47 supersede human inspection, and must not be cited as evidence that they do. Its value here is that its limitations section concedes the mechanism identified above: correlated blind spots, in that a model which writes vulnerable code may fail to see that vulnerability when reviewing it, and prompt injection as "a qualitatively new attack surface that does not exist for human reviewers." A paper arguing for automated review thus concedes the two properties that make it nonindependent. It also reports secondarily two workload figures used later here: developers at large organisations spend 10–15% of working hours reviewing code, and review latency from submission to actionable feedback routinely exceeds 24 hours. Those figures are the cost pressure that makes automated review attractive, and are why this section cannot end at "keep the human." The second of those concessions deserves separating out, because it changes the character of the problem. Correlated blind spots as described so far are a statistical property of two systems built from the same material: their errors coincide more often than chance would give because their training data and their inductive biases coincide. An adversarial input makes that coincidence something an attacker can select for. 

A prompt injection that captures a generator built on a given model is, by construction, a strong candidate for capturing a verifier built on the same model, so an attacker who knows that the reviewing agent belongs to the same family as the writing agent needs one exploit rather than two. Definition 5 relativises independence to an error class, and that relativisation is what makes the point expressible: the relevant class is not only the errors the systems generate unprompted but the errors an adversary can induce, and independence must be assessed against the second as well as the first. A verification arrangement can therefore be adequately decorrelated in ordinary operation and wholly correlated under attack, which is the condition under which it will actually be tested. This paper does not extend the empirical claim, since no measurement of induced correlation across model families was located in its research pass; but the independence register proposed in Section 10.7 should record the model family of every verifier for this reason as well as for the statistical one, and a register that treats independence as a property established once at design time will not capture it. 7.5 Independence, not judgement, is the scarce input Proposition 10 (No dependent verifier). A control loop closed by a verifier that is not independent of its generator does not converge on correctness. Feedback from an environment the generator did not author — a compiler, a database constraint, a settled transaction, a customer's reply — supplies such independence; feedback consisting of the generator's own assessment of its work, or of an artefact the generator also produced, does not. Where the generator and the verifier are instances of the same model, the loop's error rate is bounded below by their shared error rate, and iteration can move the output away from correctness. Refutation condition. Refuted if iterated verification by a verifier demonstrably correlated with the generator on the relevant error class is shown to improve accuracy on a task class with a non-trivial error rate. Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 48 Proposition 11 (Independence is necessary and not sufficient). The marginal value of a verifier is the product of its decorrelation from the generator and its detection probability on the error class. The two factors move in opposite directions as the generator improves: capability improvement leaves decorrelation intact while lowering prevalence, and by Proposition 7 falling prevalence lowers a human verifier's detection probability. The human's oversight value is therefore not eliminated by improvements in model capability, but neither is it preserved by them; it must be maintained by an instrument that holds detection probability up as prevalence falls, and it is destroyed outright by the overseer adopting the generator's frame. Refutation condition. Refuted if verifier value is shown to be increasing in agreement with the system holding accuracy constant, or if detection probability on rare error classes is shown to be invariant to prevalence under an oversight regime carrying no prevalence-restoring instrument. The second of these propositions carries the paper's most important argument, and it is not the argument usually made. The reason a human is required is not that the human judges better. On many of the tasks at issue the human demonstrably does not: Section 6 established that the risk assessment outperformed the unaided participant and was degraded by them, that the AI outperformed roughly two-thirds of radiologists while AI predictions produced a near-zero average treatment effect, and that human–AI combinations lose to the better of the two alone at g = −0.23. A defence built on human superiority rests on a premise the evidence contradicts, and will be dismantled by the next capability release. The defence offered here is structural instead. A control loop requires a verifier statistically independent of its generator, in the sense of Definition 5. Independence is not a level of skill and cannot be acquired by acquiring skill. It is a property of the relationship between two systems — whether their error distributions are conditionally dependent — and no capability improvement on either side supplies it. That is why Proposition 10 is a claim about the loop rather than the model, and why the two ICLR results, the Vend supervisor and Monperrus's concession are instances of one thing. Three consequences follow. First, the human's oversight value is neither eliminated nor preserved by model improvement; it has to be actively maintained, and the reason is arithmetic. The marginal value of a verifier is the product of two terms: its decorrelation from the generator, and its probability of detecting an error of the relevant class when one is present. Capability improvement leaves the first term untouched, for the reason just given — independence is a joint property of the pair, and no improvement on either side supplies or removes it. But capability improvement attacks the second term, and it does so by this paper's own mechanism. A better model has a lower error rate; a lower error rate is a lower prevalence of exceptions in the overseer's stream; and by Proposition 7 falling prevalence shifts the overseer's decision criterion conservatively and lowers detection probability while leaving perceptual sensitivity intact. Green and Chen (2019a) arrive at the same place from the other side: their overseers' assessments of the algorithm's accuracy had no significant relationship to its actual accuracy, so nothing in the overseer's own experience reports the decay while it is happening. The independence argument establishes that the human's value cannot be competed away by capability; Section 5 establishes that it can nonetheless be allowed to decay to nothing. What follows from that is a design obligation rather than a caveat, and it converts the tension into the section's most usable result. If detection probability falls as prevalence falls, then the value of an independent verifier survives capability growth only where the firm operates an instrument that holds detection probability up against falling prevalence. The paper already contains that Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 49 instrument: the injection of synthetic exceptions at elevated prevalence with immediate groundtruth feedback, which is the fourth design proposal of Section 10.7 and the only intervention in Wolfe and colleagues (2007) that moved the low-prevalence miss rate at all, where payoff manipulation and forced slowing did nothing. Its warrant must be taken at the strength Section 5.3 establishes and no higher: in Experiment 7 (N = 14) the low-prevalence miss rate was 21% against 18% in the matched high-prevalence condition, a difference not significant at p = .312, while the 46% customarily set against that 21% is the low-prevalence baseline of Experiment 1a (N = 24), so the headline contrast is across experiments rather than within a design, and a failure to reject the null at N = 14 is weak evidence of elimination rather than a demonstration of it. That proposal should therefore not be read as an optional refinement of an oversight regime. It is the precondition on which Proposition 11's survival claim depends: a firm asserting a durable oversight value on the strength of independence, while operating no prevalence-restoring instrument, is asserting the first factor of a product whose second factor its own improvement programme is driving toward zero. The derivation in Section 8.3.1 reaches the same instrument from the cost side and establishes something stronger than a design preference. Of the three remedies available to a firm whose oversight efficacy is decaying — raising the overseer's uncommitted attention, lowering the cost of handling an item, and routing non-uniformly so that error is more prevalent within the examined stream — only the third acts on the detection term rather than on the coverage term, and it is the detection term that carries the part of the decay produced by improvement in the supervised system itself. Injection and the triage routing of Section 8.4 are therefore the only instruments in this paper that address that part of the loss, which is why the argument of Sections 7, 8 and 10 converges on them. So maintained, and only so maintained, this is the one justification for human oversight located in this paper that survives arbitrary improvement in the supervised system, and therefore the only one on which a durable governance commitment can be built. Second, that value is destroyed by the human adopting the model's frame. If the overseer's priors, evidence and reasoning converge on the system's, the conditional independence in Definition 5 fails and the human becomes a slower instance of the same verifier. This is exactly what the reliance literature reviewed in Section 6 measures: explanations that raise agreement whether or not the system is correct, transparency that degrades outlier detection through information overload, confidence displays that move the level of reliance without improving its discrimination. The reframing is worth stating explicitly: automation bias is not a human weakness to be trained away, but the loss of the very property that made the human worth consulting. A firm that succeeds in getting its overseers to agree with the model more often has degraded its control environment, and its agreement statistics will report that degradation as an improvement. Third, independence is purchasable and losable by design. Definition 5 names its determinants: base model, training corpus, input channel, vendor and incentive. Each is a fact about a deployment that a firm knows, can record and could disclose. Independence is therefore not a philosophical property but an auditable one, and Section 10 proposes the instrument — an independence register, stating for each control whether the verifier shares any of those five with the generator. The Vend Phase 2 failure is invisible in every process metric a conventional control framework collects, and immediately visible in such a register. It remains to connect this to the under-quoted half of the irony introduced in Section 2. Bainbridge (1983) is a five-page analytical essay with no data, and the line usually taken from it concerns monitoring. The corollary matters more here: if the human operator must check the details of what the computer is doing, the computer must make its decisions using methods and criteria, and at a rate, which the operator can follow. The only verifier affordable at machine speed is another machine, and the only verifier reliably independent of a general-purpose model across arbitrary Human-on-the-Loop · WP No. 8 7. Why the Loop Cannot Close Itself 50 error classes is, on present evidence, a human. The qualification is doing work and should not be dropped: a verifier need not be human to be independent, and Section 11.5 sets out the conditions under which an automated one would qualify — an adversarially trained verifier, a formal method carrying a soundness guarantee, or an ensemble of genuinely heterogeneous models — of which Stechly's sound external verifier is an instance already realised within its own narrow domain. What no automated verifier has yet been shown to be is independent across the open-ended range of error classes a general-purpose agent can generate. Speed and independence therefore trade off directly, along the same axis that Section 3 showed determines loop position and Section 4 determines solvency. A firm cannot buy its way out of that trade with capability, because capability is what makes it bind. Section 8 asks what remains once that is accepted. 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 8.1 The problem the argument has reached Everything established so far is a constraint. Sections 3 to 5 show that control oversight is fragile, and — this is the uncomfortable part — becomes more fragile as the supervised system improves, because reliability constancy and low exception prevalence both work against detection. Section 6 shows that oversight improves accuracy only under two conditions, decorrelation of the overseer's errors from the system's and verification cheaper than production, which are rarely present where oversight is mandated; an independent information channel is the strongest and most auditable source of the first rather than the condition itself, and is what a firm can actually check, decorrelation being unobservable without labelled ground truth on the error class. Section 7 shows the function cannot be automated away, because automating it destroys the independence that made it a control. If the argument stopped there it would be a counsel of despair, and it would also be wrong. Firms plainly do exercise effective influence over algorithmic systems. Systems get built to do one thing rather than another; they get switched off; their objectives get rewritten after a bad quarter; capital moves toward some and away from others. Any theory implying that human influence over automation is illusory is refuted by ordinary organisational experience. The correct conclusion is not that influence is impossible but that it is exercised somewhere other than where the oversight literature and the oversight regulations both look for it. This section locates it, defines the distinction that makes the location precise, states the paper's twelfth proposition, and rebuilds the five-role model Section 1 took as the paper's object of study. 8.2 Where influence is actually exercised The most direct empirical answer comes from outside the human-factors and human–computer interaction literatures that have supplied most of this paper's evidence. Cameron and Rahman (2022), Organization Science 33(1):38–58, is a peer-reviewed comparative multi-year ethnography of two large platform companies, one in ride-hailing and one in online freelancing, combining participant observation, interviews and archival material. Their subject is workers under algorithmic management rather than managers supervising algorithms, but the structural finding transfers directly. Control and resistance co-constitute, and workers' latitude is greatest early in the labour process and shrinks monotonically through execution and evaluation. Effective agency is exercised upstream — at selection, framing and the setting of objectives — and not downstream at act-by-act supervision, where by the time the act occurs the latitude has already decayed. This is an independent, ethnographically grounded arrival at the same place as the formal theory imported in Section 3. Aghion and Tirole (1997) show that a principal whose probability of being Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 51 informed is low optimally rubber-stamps, so that formal authority persists while real authority migrates to the informed party, with overload and urgency of decision named among the drivers. Dessein (2002) shows that against an informed agent delegation dominates communication over a wide parameter range, and that retaining a veto typically reduces the principal's expected utility unless the incentive conflict is extreme. Both results and the ethnography point to the same instruction: control the objective, not the decision. That three literatures with no methodological overlap converge on it is the strongest warrant this paper can offer for the reallocation proposed below. The transfer from gig platforms to firm-internal AI deployment is an inference rather than a finding, and Cameron and Rahman make no claim about the latter. The second source supplies the vocabulary. Kellogg, Valentine and Christin (2020), Academy of Management Annals 14(1):366–410, must be labelled accurately: it is an integrative review and theory-building essay synthesising an interdisciplinary literature, not primary empirical research, and no effect size should be attributed to it. Its contribution is a taxonomy of six mechanisms of algorithmic control, organised under the three classical control functions — directing, through restricting and recommending; evaluating, through recording and rating; and disciplining, through replacing and rewarding. The authors argue that algorithmic control differs from earlier employer control in being simultaneously comprehensive, instantaneous, interactive and opaque. The inversion this paper needs follows immediately. Kellogg and colleagues describe algorithms controlling humans. A human-on-the-loop configuration inverts the arrows: the human is the controlling party and the algorithmic system the controlled one. The same six mechanisms remain available, and a firm can restrict, recommend, record, rate, replace and reward an automated system as readily as an algorithm can do those things to a worker. But the two directions are not symmetric, and the asymmetry is the crux. When an algorithm records and rates a human, those functions are executed by a system whose capacity to observe is effectively unbounded relative to the volume of behaviour observed — this is what "comprehensive and instantaneous" means. When a human records and rates an algorithm, the same functions are executed by an overseer whose capacity is bounded by attention in exactly the way Sections 4 and 5 measure. Restricting and rewarding are cheap in both directions, being exercised once and then holding. Recording and rating are cheap in one direction and expensive in the other. The claim is therefore not that humans are less intelligent than the systems they supervise, and nothing here depends on any such claim. It is that two of the six mechanisms have an attention cost scaling with the volume of behaviour controlled, and that this cost falls entirely on the human when the arrows are inverted. That asymmetry, and not any proposition about intelligence, forces the reallocation the rest of this section proposes. 8.3 The dichotomy and its scaling Definition 6 (Constitutive and control oversight). Constitutive oversight comprises decisions that fix the objective function of the automated system — purpose, admissible means, and the allocation of capital and attention across purposes. Its exercise is ex ante, revisable on a governance cycle, and its attention cost is independent of the system's throughput. Control oversight comprises the monitoring of executions and the interception of exceptions. Its exercise is contemporaneous and its attention cost is increasing in throughput. Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 52 Proposition 12 (Scaling dichotomy). Let T be the throughput of an automated decision class, p(T) the per-action rate of consequential error, A the overseer's uncommitted attention per period, n = A/(IT + c) the number of items that budget permits to be examined, m the volume routed to the overseer, R the recall of the routing rule — the share of all consequential errors it places in the routed stream — s = min(1, n/m) the share of the routed stream the budget permits to be examined, π = R·T·p/m the prevalence of error within it, and d(π) the probability that an examined error is detected, increasing in π by Proposition 7. The attention cost of constitutive oversight is invariant in T. Control oversight admits no corresponding invariance at a fixed detection guarantee. The share of consequential errors it removes is η(T) = R(T)·s(T)·d(π(T)). Under uniform review the routing rule is the null rule, R = 1 and m = T, so η(T) = (n/T)·d(p(T)), which is O(1/T) for any p and o(1/T) wherever p is strictly decreasing in T; holding η at a positive constant ηฬ„ then requires A(T) = ηฬ„ ·(IT + c)·T / d(p(T)), which is Ω(T) and strictly ω(T) under Proposition 7. Non-uniform routing raises π and with it d, but bounds efficacy above by the recall of the router, η ≤ R, so triage does not remove the constraint: it relocates it from the overseer's attention to the router's recall, and where the router is an instance of or is correlated with the supervised system, R is governed by Definition 5. Whether the absolute exposure T·p(T) rises, falls or is constant, the share of it removed by control oversight is bounded by a quantity no fixed attention budget can hold up; the only oversight whose efficacy is invariant in T is constitutive. Refutation condition. Refuted if a control-oversight mechanism is exhibited whose per-period attention cost is bounded in T while the share of consequential errors it removes is bounded below — whether through coverage, through routing recall, or through any combination of the two — without transferring the control function to a verifier correlated with the generator. The argument for the proposition should be developed carefully, and its limits stated as carefully as its content. Constitutive oversight consists of decisions taken once and then standing: what the system is for, what means it may use, how much capital and human attention it receives against competing purposes. These are exercised ex ante and revised on a governance cycle — quarterly, annually, or on a material change. Their attention cost is a function of the number of decision classes the firm operates and the frequency of that cycle, not of how many actions the system takes between cycles. A board that fixes a credit policy consumes the same attention whether the policy is applied ten thousand times or ten million. Control oversight consists of monitoring executions and intercepting exceptions, and is by definition exercised contemporaneously with the actions it supervises. Its attention cost is at least linear in throughput wherever the exception rate is fixed, because the exception arrival rate λ in Definition 4 is then proportional to throughput. That conditional cannot be relied upon, and Section 8.3.1 shows why: with λ = T·p(T) and the per-action error rate p falling as the decision class matures, the arrival rate need not grow at all and may fall. The proposition is therefore stated in terms of what the oversight achieves rather than of what arrives at it, and the derivation in Section 8.3.1 establishes that the share of consequential errors control oversight removes falls as O(1/T) whatever λ does, and strictly faster wherever p is falling. On the cost side the linear figure is in practice an understatement, for two reasons already established. The queueing term WT in Definition 4 is zero at low load and grows as several processes demand attention simultaneously, so the span bound N ≤ NT/(IT + WT) + 1 tightens as arrivals rise. And the context re-acquisition cost c from Section 4 is incurred per interruption rather than per unit of work, so a doubling of arrival rate more than doubles the attention consumed unless the interruptions can be batched — which, for exceptions requiring interception before consequence, they cannot be, since batching converts an interception into a post-hoc review. Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 53 The consequence follows arithmetically rather than empirically. If constitutive cost is flat in T while the cost of holding a fixed detection guarantee is increasing in T, and uncommitted attention A is fixed in the short run, then there exists a throughput T* at which the second crosses A. Below T*, both forms of oversight are affordable at the stated guarantee. Above it, only constitutive oversight remains, and control oversight does not fail visibly but degrades into acceptance in the manner Proposition 3 describes and the clinical alerting evidence in Section 4.4 is consistent with — evidence drawn from three different clinical systems, none of them a natural experiment, and supporting the mechanism rather than a magnitude. The managerially significant version of this is that any configuration whose economic case rests on raising throughput has an insolvency point built into its business case. The very variable that justifies the deployment is the variable that exhausts the oversight it promises. Figure 4. The scaling dichotomy. Attention cost is plotted against the throughput T of an automated decision class. Constitutive oversight — purpose, admissible means, capital and attention allocation — is flat in T, being exercised ex ante and revised on a governance cycle. Control oversight — monitoring executions and intercepting exceptions — rises at least linearly in T at a fixed detection guarantee, and faster once the queueing term WT, the context reacquisition cost c and the falling prevalence term derived in Section 8.3.1 are included. The horizontal line marks available uncommitted attention A. Their intersection is T*, the control-insolvency throughput, beyond which the only oversight that survives is constitutive. The inset shows the span bound N ≤ NT/(IT + WT) + 1 with the effective range of 4–6 measured by Crandall and Cummings (2007). The curves are illustrative of shape; no exponent has been estimated for knowledge work. That last sentence must be given its full weight, because the proposition is easy to overclaim and this paper's discipline requires that it not be. Proposition 12 is an asymptotic claim about the shape of the efficacy function and of the attention cost of holding it fixed, not a quantitative claim about their level or their curvature. Nothing in this paper estimates the sample size n, the detection function d or the error rate p for a knowledge-work setting; the only settings in which the capacity parameters have been measured at all are the robotics supervision studies of Section 4.3 and the prevalence experiments of Section 5.3, and their transfer to a deployed agent system is an assumption rather than a result. Nor is T* estimated for any decision class. The proposition asserts that T* exists and that it is reached, not where it lies, and a firm that wished to locate its own T* would have to measure λ, IT, c and A for itself, which is what the solvency statement proposed in The scaling dichotomy constitutive oversight — O(1) A — available attention control oversight — Ω(T) T* — control insolvency throughput T of the automated decision class attention cost N ≤ NT/(IT + WT) + 1 measured effective span: 4–6 Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 54 Section 10 is for. Section 11.2 records the concession in the form it actually takes: the asymptotics are derived rather than asserted, and what is unmeasured is the set of four functions the derivation is stated over — the number of items a fixed attention budget permits to be examined, n = A/(IT + c), the detection function d(π), the per-action rate of consequential error p(T), and the recall R of the rule by which items are routed to the overseer — none of which has been estimated for a deployed agent system. A reader who rejects the asymptotic form is entitled to reject the proposition; the refutation condition attached to it states exactly what would be required. 8.3.1 Why constant-sample review does not escape the dichotomy, and what happens when reliability improves Two objections meet here, and they pull in opposite directions. The first disposes of the proposition in a line if left unanswered: a firm that reviews fifty sampled items a week, whether the system emitted a thousand actions or ten million, runs a control mechanism whose per-period attention cost is flat in throughput and which hands the control function to no verifier at all. The second cuts the other way: writing the cost of control oversight as a function increasing in throughput presupposes that the per-action rate of consequential error is constant, and in a maturing system it is not, so the exception arrival rate may not grow at all. Both are answered by one derivation. The quantities are those of Definitions 3 and 4, with two additions. Let T denote the number of actions the automated decision class takes per period, and p(T) the per-action rate of consequential error prevailing at that throughput. Let A be the overseer's uncommitted attention per period, IT the mean interaction time to closure and c the context re-acquisition cost, so that the number of items the attention budget permits to be examined in the period is n = A/(IT + c) Let d(π) denote the probability that an error present in an examined item is detected, when the prevalence of error within the examined stream is π; by Proposition 7, d is increasing in π. It is assumed throughout that the examined items are drawn uniformly at random from the action stream, so that prevalence within the sample equals prevalence within the stream and the detection term is d(p(T)). That assumption is provisional and is discharged in the sixth step below, where nonuniform routing is shown to change the detection term and to introduce one of its own. The first step is the one this paper already had. Coverage, the share of actions examined, is n/T. The attention budget fixes n independently of T, since none of A, IT and c is a function of how many actions the system takes. Coverage therefore falls as O(1/T), and no diligence exercised within the sample alters that. The second step identifies the quantity that actually matters, which is neither coverage nor absolute exposure. Define the oversight efficacy η(T) as the share of consequential errors that control oversight removes in the period. Errors committed per period are T·p(T); errors detected per period are the number examined, n, times the probability that an examined item carries an error, p(T), times the probability that an error so presented is detected, d(p(T)). The ratio is the efficacy, and in forming it the error rate cancels: η(T) = n·p(T)·d(p(T)) ⁄ T·p(T) = (n/T)·d(p(T)) The cancellation answers the second objection. Oversight efficacy does not depend on the level of the error rate at all; it depends only on coverage and on detection at the prevailing prevalence. Whether the absolute exposure T·p(T) grows, shrinks or holds constant as the system matures is therefore irrelevant to what oversight contributes. The natural formulation of the same concern, which reasons about undetected exposure p(T)·(T − n) and observes that it grows as Ω(T), holds only Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 55 where p is constant and is superseded here: a firm whose error rate falls quickly enough may face a shrinking absolute exposure while the share of it that oversight removes goes to zero, and it is the share that measures the control. The third step states the asymptotics properly. Since d is a probability and n is bounded by the attention budget, η(T) ≤ n/T and hence η(T) = O(1/T) unconditionally, for any behaviour of p whatever. Where p is in addition strictly decreasing in T, d(p(T)) is decreasing in T by Proposition 7, so the product of a term falling as 1/T and a term falling strictly gives η(T) = o(1/T). Efficacy falls strictly faster than coverage does, and reliability improvement, which is the thing the firm buys, is therefore also the thing that accelerates the decay of its oversight's contribution. This is the second blade of the oversight scissors named in Section 5.4: the first lowers the probability a supervisor detects a failure at all, the second lowers the efficacy of a fixed review function faster than throughput growth alone would. Nothing in the result requires the error rate to rise; it is strongest precisely where the rate falls. The fourth step prices the refusal of that decay. Suppose a firm undertakes to hold efficacy at some positive constant ηฬ„ . Inverting the expression gives the sample the undertaking requires, n(T) = ηฬ„ ·T/ d(p(T)), and multiplying by the attention cost of examining one item gives the attention it consumes: A(T) = ηฬ„ ·(IT + c)·T ⁄ d(p(T)) Where d is bounded below by a positive constant this is Ω(T). Where p is decreasing in T the denominator shrinks as T grows, and A(T) is then strictly ω(T), superlinear for the reason that the firm's system is getting better. The interpretive result is the rigorous form of what this paper claims. The claim is not that control oversight costs Ω(T) in the arrival rate, since the second objection is right that it may not; it is that the cost of holding a fixed detection guarantee is Ω(T), and superlinear once falling prevalence is accounted for. Constitutive oversight carries no analogous term because it operates on the objective function rather than on instances: the decision that fixes what a system is for is not indexed by the number of actions the system takes, so nothing in it plays the role of the coverage ratio n/T, and its efficacy is not a function of T at all. The fifth step enumerates what a firm can do about the decay, which is three things, each with a price. It can raise A, which is Proposition 4 restated: attention is fixed in the short run, so raising it means removing other work from the overseer's role. It can lower IT + c, which raises n and therefore efficacy proportionally, but the reduction is bounded below by the resumption-cost evidence of Section 3.3, since an interruption costs something to recover from and no tooling drives that cost to zero. Or it can abandon uniform sampling and route non-uniformly, so that error is more prevalent within the examined stream, raising π and therefore d — what the injection of synthetic exceptions at elevated prevalence proposed in Section 10.7 achieves artificially, and what triage designs such as MASAI achieve structurally. The asymmetry between the three governs where this paper's design proposals concentrate: the first two act on the n/T term and buy a constant factor within an order that is unchanged, while only the third acts on the d term. Since it is the d term that carries the o(1/T), routing is the only one of the three that addresses the part of the decay reliability improvement itself creates. Whether it delivers that gain, however, turns on a property of the routing rule the derivation has so far left out, and the sixth step supplies it. The sixth step models the routing rule itself, because the fifth step as it stands is incomplete in the firm's favour. A rule that selects which items reach the overseer thereby determines which items do not, and an error it fails to select is not in the position of an error the overseer examined and missed: it is outside the sample, never eligible to be seen, and recoverable by no diligence and no additional attention. The recall of the rule is therefore a term of the derivation rather than a caveat Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 56 upon it, and carrying it explicitly recovers the earlier result as a special case rather than displacing it. Let m denote the volume routed to the overseer per period and R the recall of the routing rule, the share of all consequential errors committed in the period that the rule places in the routed stream. The attention budget still permits n = A/(IT + c) items to be examined, so the share of the routed stream that is actually examined is s = min(1, n/m): where the rule routes more than the budget can absorb, the overseer samples within it; where it routes less, the overseer examines all of it. Prevalence within the routed stream is π = R·T·p/m, the routed errors over the routed volume. Errors detected per period are then R·T·p·s·d(π) and errors committed per period are T·p, so the error rate cancels a second time and the general form of oversight efficacy is η = R · s · d(π) Two special cases exhaust ordinary practice. Uniform review is the null routing rule: it routes everything, so R = 1, m = T, s = n/T and π = p, and the general form collapses to η = (n/T)·d(p), which is exactly the expression obtained in the second step of this section. That should be said plainly, because it is the strongest form the generalisation takes. The uniform case is not a different model from the triage case; it is the general model operated with a router that discriminates nothing, and the O(1/T) decay derived above is what a firm pays for setting R to its maximum by declining to select at all. Triage within budget is the other case: where the routed volume is no greater than what attention can absorb, m ≤ n, the overseer examines the whole routed stream, s = 1, and efficacy is η = R·d(π); and since d cannot exceed one, η ≤ R The first consequence is that triage does not escape the dichotomy; it relocates the constraint that binds. Under uniform review the ceiling on efficacy is coverage, and coverage is set by the attention budget, so the human is the limiting factor and n/T must fall as throughput rises. Under triage the ceiling is R, and R is set by the router, so the human ceases to be the limiting factor and the router becomes it. That is a real improvement and should be stated at full strength: R carries no throughput term, so a routing rule need not lose recall as the volume it screens grows, whereas n/T falls by construction. But it is an improvement obtained by moving the binding constraint out of a component the firm measures, in hours of uncommitted attention, and into a component it has not been measuring at all. The second consequence is that the router inherits Definition 5. Where the routing rule is itself an instance of the supervised system, or is trained on the same data, or reads the same inputs, its recall is systematically lowest for exactly the error classes the generator is most prone to commit, since what makes an error easy for the generator to produce is what makes it hard for a correlated selector to flag. This is the correlated-blind-spot argument of Section 7.4, operating one layer earlier than that section states it. A triage layer built from the same model as the generator does not merely fail to catch some errors; it fails to catch them non-randomly, and in the direction that matters, so that the unrouted stream is enriched for precisely the failure modes the oversight was installed to intercept. The adversarial extension of Section 7.4 applies here with more force, since an attacker who can suppress routing removes the item from human view entirely rather than merely surviving an inspection of it, and leaves no trace in the review record because no review was ever attempted. The third consequence is the sharpest of the four and should be stated as such. R is not observable from the routed stream. Everything a firm sees of its triage layer is conditioned on having been Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 57 routed — the rule's precision, the outcomes of the reviews it commissions, the catch rate the overseer achieves within it. Recall is a statement about the unrouted stream, and estimating it requires sampling what the router did not select, which is precisely the uniform sampling that triage was adopted in order to avoid. A firm can therefore buy a higher detection rate with triage and lose the ability to know what it has bought. The measurement instruction follows immediately, and belongs with the design proposals of Section 10.7: any firm operating a triage layer must run a residual uniform sample of the unrouted stream, sized to estimate R rather than to catch errors, and must report that estimate alongside whatever workload reduction the layer delivers. That sample is a permanent cost of the design and not a transitional one: it does not fall away once the rule has been validated, because R is a property of a rule operating on a distribution that drifts. The fourth consequence marks the real boundary. 

Where ground truth arrives late but does arrive, R becomes measurable without a residual sample, because the unrouted stream eventually reports on itself: in screening, interval cases reveal what the router missed; in credit, defaults do; in fraud, chargebacks do. This is exactly why the MASAI trial can be evaluated at all, and why population screening is the setting in which a triage design is most defensible. Where the condition of Proposition 9 holds, however, and the ground truth is a counterfactual future that never arrives, R is not estimable even in principle: no residual sample recovers it, since sampling the unrouted stream yields items whose correct disposition is itself unknown. A triage design in such a decision class is one whose central parameter is permanently unknown, and the firm operating it holds an oversight arrangement whose efficacy has an upper bound it cannot estimate. Neglect time NT is unmeasurable in those same classes and for the same reason, as Section 11.1 sets out: both quantities are defined against a trajectory that was not realised, and the gap is definitional rather than empirical. The seventh step states what the derivation does not establish. The detection function d(π) is treated as a function of prevalence alone, whereas the vigilance literature surveyed in Section 5 makes detection depend on event rate and discrimination difficulty as well, so d is a reduced form at best whose monotonicity in π is the only property used. The routing recall R is treated as a fixed parameter of a rule when it is a property of a rule operating on a particular distribution, and it will drift as that distribution drifts; the derivation establishes that efficacy is bounded above by R and estimates R for nothing, and that a rule which raises prevalence by discarding the cases it finds hard raises d while lowering efficacy, without establishing how often such rules are built. No value of R has been measured for any deployed enterprise triage layer located in this paper's research pass. Writing p as a function of T is expositional economy rather than a causal claim, maturity being properly indexed by elapsed time and cumulative investment rather than by throughput. And none of n, d, p and R has been estimated for a deployed agent system anywhere the research underlying this paper could locate. What is offered is a comparative-statics claim about directions and orders, not a point prediction, and Section 11.1 records the absence of those measurements as this paper's principal weakness. What a firm operating constant-sample review must therefore declare is not the fact of the sample, which is the fact it currently discloses, but the efficacy the sample buys: the share of consequential errors the review removes, given the throughput it is drawn from, the prevalence within it, and the reviewers' measured detection rate at that prevalence. Every term is constructible from quantities the firm already holds. No firm known to this paper reports it, and until one does, a disclosure that a human reviews a sample of the system's output states the price of the control and says nothing about its strength. Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 58 Figure 5. Oversight efficacy under a fixed attention budget. Upper panel: three quantities plotted against the throughput T of an automated decision class on a logarithmic axis. Coverage n/T falls as O(1/T), the attention budget fixing n independently of T. Detection at the prevailing prevalence, d(p(T)), falls gently and is read against the righthand axis, carrying the measured anchors of Wolfe, Horowitz and Kenner (2005) — detection of 93%, 84% and 70% at target prevalences of 50%, 10% and 1% respectively. Oversight efficacy η(T) = (n/T)·d(p(T)) is their product and falls fastest, as o(1/T) wherever p is strictly decreasing. The flat line is constitutive oversight, whose efficacy is invariant in T because it operates on the objective function rather than on instances. The plotted efficacy curve is the uniformreview case, in which the routing rule is the null rule and its recall is one; non-uniform routing raises the curve by raising the prevalence of the examined stream but cannot lift efficacy above the recall R of the router, which the horizontal dashed ceiling marks at an illustrative value that is not estimated from any deployment. Lower panel: the attention A(T) = ηฬ„ ·(IT + c)·T ⁄ d(p(T)) required to hold efficacy at a fixed ηฬ„ , drawn against a straight Ω(T) reference that assumes constant detection; the widening gap between them is the cost of the falling prevalence term. The curves are illustrative shapes carrying the measured Wolfe anchors. They are not fitted data: no value of n, d, p or R has been estimated for a deployed agent system, and the vertical scales are relative. 8.4 The reconstruction of the five roles The five roles advanced in the practical proposal this paper examines are purpose, values, capital allocation, governance and exception intervention. Sorted by Definition 6, they do not form a single class. Purpose, values and capital allocation are constitutive. Each is a decision about the objective function of the automated system rather than about any of its outputs: what the system is for, which Oversight efficacy under a fixed attention budget (a) coverage, detection and their product, against throughput 1 0.1 0.01 100% 90% 80% 70% 60% T0 10 T0 100 T0 constitutive oversight — efficacy invariant in T triage ceiling: η ≤ R 93% 84% 70% throughput T (log scale) share of errors removed (log) detection d(π), right axis coverage n/T — falls as O(1/T) detection at prevailing prevalence d(p(T)) — right axis measured anchors (Wolfe et al. 2005): detection 93% at prevalence 50%, 84% at 10%, 70% at 1% η(T) = (n/T)·d(p(T)) — o(1/T) constitutive oversight — efficacy invariant in T (b) attention needed to hold efficacy at a fixed η: A(T) = η·(IT + c)·T / d(p(T)) A(T) — ω(T) Ω(T) reference the cost of the falling prevalence term throughput T attention required Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 59 means it may and may not employ, and how much of the firm's capital and attention it receives relative to competing uses. None is time-critical in the sense of Definition 1, because none must be taken within τsys of any particular action. Each is genuinely where real authority resides, in the exact sense Aghion and Tirole give the term — the party who fixes the project set and the payoff function is not rubber-stamping anything, because their decision is not a ratification of a proposal generated by a better-informed party. And, decisively for this section, the attention cost of all three is invariant to throughput. This is the part of the practical proposal that survives the paper's critique intact, and it should be said clearly that it survives: the claim that a firm's irreducible contribution lies in purpose, values and capital allocation is well founded, for a reason the proposal itself does not state. Governance and exception intervention are control. Both are exercised contemporaneously with the system's operation; both have attention costs that scale with the volume of operation; and both are treated by the practical proposal, and by most of the regulatory instruments surveyed in Section 10, as though they were nearly free. Exception intervention is control oversight in its pure form. Governance, as the proposal uses the term, means the ongoing supervision of how the system behaves in operation, which is monitoring under another name — although the portion of governance that sets policy rather than checking conformance is constitutive, and the ambiguity of the word is part of the problem. The sharp claim this permits is not that the five-role model is wrong. It is that the model is unpriced. It lists five roles as though they were commensurable and assigns all five to the human, when two of them carry a cost that the balance sheet of the proposal does not record and that rises with precisely the variable the deployment is meant to increase. A firm adopting the model in good faith will find that three of its five commitments hold indefinitely and two of them silently stop being met at a throughput nobody identified in advance, with the failure appearing not as an admission but as a rising override rate. The reconstruction is therefore not to delete governance and exception intervention from the model. They are real functions and something must discharge them. It is, first, to require that any firm claiming those two roles state the solvency conditions under which it can perform them — the arrival rate, the interaction time, the re-acquisition cost and the uncommitted attention that make the claim arithmetically possible — and, second, to move as much of their content as design permits into the constitutive class. The migration is easier to see in examples than in the abstract. A decision rule fixed ex ante is constitutive; the same decision taken case by case is control. A firm that decides in advance that no refund above a stated value may be issued without human authorisation has spent its attention once; a firm that reviews refunds as they arise has spent it per refund, and the second firm's cost rises with volume while the first firm's does not. A hard constraint compiled into the system's action space is constitutive; catching violations of the same constraint after the fact is control. A system that is architecturally incapable of writing to a production database during a freeze window differs from a system instructed not to and monitored for compliance, and the Replit incident in Section 7.4 is what the difference costs. And an allocation of which cases receive human attention is constitutive; reviewing every case is control. The firm that decides in advance which strata of its caseload are routed to a human has made one decision whose cost does not scale — but the decision is what does not scale, and not the reviewing it commissions. And since the cost of the rule is no longer the binding claim, the quantity such a firm must report of its routing rule is not that its cost is fixed but what its recall is. Where the stratum is defined as a fraction of the caseload, the cost of reviewing what the rule routes remains linear in volume; where it is defined as a fixed count, the cost is held flat only by letting coverage and therefore efficacy fall, and the argument of Section 8.3.1 Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 60 applies in full. What a routing rule buys is not exemption from the scaling but the power to choose which part of the caseload the surviving human attention is spent on — and, by the fifth step of that derivation, that choice is not merely a matter of allocation. A rule that concentrates human attention where error is more prevalent raises the prevalence π within the examined stream, and by Proposition 7 raises the detection term d(π) with it. Routing is accordingly the one instrument available to a firm that acts on the d term rather than on the coverage term n/T, which is why it is the only instrument that touches the component of the decay attributable to reliability improvement itself. It is also bounded, and the bound is the sixth step of Section 8.3.1: efficacy under triage cannot exceed the recall R of the rule that does the routing. A firm that migrates its review into a routing rule has therefore moved its exposure from an attention constraint it can measure to a recall constraint it typically cannot, and has acquired with the rule the obligation to estimate that recall from the stream the rule declines to route. The firm that reviews everything has declined to make that choice, and has made a commitment whose cost is linear in volume by construction and whose examined stream carries the population prevalence, which is the lowest prevalence available to it. The third example is not hypothetical, and its clinical instance was already introduced in Section 6.5. MASAI is exactly this migration performed in a real screening programme. The AI does not produce outputs for a radiologist to audit; it determines which cases receive how much human reading, and the human's role is fixed by a protocol established in advance. The human attention in the system is reallocated rather than added, which is why the trial reports a 44% reduction in screen-reading workload alongside cancer detection of 6.0 versus 5.0 per 1,000 screened and false positives of 1.5% in both arms. That figure should be read for what it is. A 44% reduction is a constant factor applied to a cost that remains linear in the number of women screened, not a change in its order: double the screening volume and the reading workload doubles again from the lower base. What routing changes on the cost side is the constant, and no more than the constant. What it changes on the efficacy side is the prevalence of the stream the human reads, and by the fifth step of Section 8.3.1 that is the one term reliability improvement attacks and the one term additional attention cannot restore. A triage design is therefore not a way of escaping the scaling; it is the only way of buying back the part of the decay that improvement in the supervised system creates, and what it buys is capped by the recall of the routing rule, since no triage arrangement removes more consequential errors than its router places in front of a human. Screening is one of the few settings in which that recall can be recovered after the fact, because interval cases eventually report on what the router did not route; a firm operating the same design where outcomes never arrive would be running it on a ceiling it cannot observe. Every failure mode catalogued in Sections 6.3 and 6.4 belongs to the override architecture, in which a human is asked to evaluate a model's judgement using the model's own evidence. MASAI never asks that question. It is the strongest available demonstration that the reallocation this section proposes is achievable in practice rather than merely coherent on paper, and firms should note what it required: not a better model and not a more diligent radiologist, but a redesign of the routing — and that the parameter such a redesign turns on, and that a firm adopting it must therefore measure, is the recall of the rule. 8.5 What must be instrumented, and the trap of instrumenting the wrong thing Migrating content into the constitutive class raises the question of what the firm should then measure, since a constitutive regime is not a regime without information. The economics of the answer is settled and predates the technology by decades. Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 61 Holmström (1979), Bell Journal of Economics 10(1):74–91, peer-reviewed and theoretical, establishes the informativeness principle: a signal belongs in the contract if and only if it is informative about the agent's action in the sufficient-statistic sense, so that what should be instrumented is determined by posterior informativeness about the action rather than by observability convenience. Both directions matter here. It licenses instrumenting signals that are not outcomes, where those signals shift the posterior over what the system actually did; and it excludes signals that are merely easy to collect — which, in an automated deployment, is most of them, since software emits telemetry in proportion to what is cheap to log rather than to what is diagnostic. Holmström and Milgrom (1991), Journal of Law, Economics, and Organization 7(special issue):24–52, peer-reviewed and theoretical, supplies the warning. Their central result is that the desirability of providing incentives for any one activity decreases with the difficulty of measuring performance in any other activities that make competing demands on the agent's time and attention. Supervising through measurable proxies therefore distorts effort toward the measurable, and the correct response is not to measure harder. One response they derive explicitly is job-design restriction: removing the unmeasurable dimensions from the agent's remit altogether rather than attempting to score them. Translated into the present setting, that means narrowing what the automated system is permitted to decide rather than widening what the human is asked to watch — a constitutive move rather than a control move, and one whose attention cost is paid once. It also supplies the direct argument against oversight dashboards built from latency, volume and cost, which are instrumented because they are cheap and which predictably degrade the dimensions nobody is scoring. The design result that makes this concrete is empirical. Bloom, Garicano, Sadun and Van Reenen (2014), Management Science 60(12):2859–2885, peer-reviewed, with an original firm and plant survey across the United States and seven European countries and instrumental-variable identification, separate two technology classes that are usually conflated. Information technologies, which make the executing agent better informed — enterprise resource planning for plant managers, computer-aided design and manufacturing for production workers — are associated with greater autonomy and wider spans of control. They decentralise. Communication technologies, which give the principal a cheap channel to the agent's behaviour — data networks and intranets — are associated with reduced autonomy for both workers and plant managers, and tighter oversight. They centralise. The identification is contestable and the sample manufacturing-heavy, but the directional finding is the one this section requires. The design implication should be stated precisely, because it is the most actionable sentence in this section. A firm choosing between instrumenting the automated system's knowledge and instrumenting the automated system's behaviour is choosing, whether or not it realises it is choosing, between a wide-span constitutive regime and a narrow-span control regime. Instrumentation that improves what the executing system knows at the point of action — better retrieval, better context, hard constraints available before the act rather than after it — pushes authority down and widens the span the firm can carry. Instrumentation that streams the system's activity to a supervisory centre pulls authority up, narrows the span, and lands the firm in the regime Section 4 shows becomes insolvent, while incurring the initiative costs Aghion and Tirole identify and the distortion costs Holmström and Milgrom identify along the way. Most firms deploying agentic systems are currently building the second and describing it as the first. 8.6 A necessary corrective on span, and the question for a board A caution is owed on the span literature, because this section's argument would be strengthened by a cognitive ceiling on supervision, and no such ceiling is established by the work usually cited for it. Human-on-the-Loop · WP No. 8 8. The Scaling Dichotomy and the Reconstruction of the Five Roles 62 Rajan and Wulf (2006), Review of Economics and Statistics 88(4):759–773, peer-reviewed panel work on more than 300 large United States firms between 1986 and 1999, find that positions reporting directly to the CEO rose from 4.5 in 1986 to almost 7 in 1999, while levels between division heads and the CEO fell from 1.6 to 1.2. Guadalupe and Wulf (2010), American Economic Journal: Applied Economics 2(4):105–127, peer-reviewed and quasi-experimental, using trade liberalisation as an exogenous competition shock, report a sample-mean span of 5.47 and find span responds causally to competitive pressure. This literature measures observed spans and their determinants. It does not establish a cognitive maximum, and this paper attributes no such number to it. The correct reading is the opposite of a ceiling: span is an endogenous organisational choice that firms adjust when the environment demands speed, which means that any oversight proposal assuming a fixed human span is assuming away the margin firms actually move. The single-digit span figures this paper does use come from the robotics supervision studies reported in Section 4.3, where interaction time, neglect time and effective team size were measured directly under controlled conditions, and those figures carry their own transfer problem, since supervising semi-autonomous robots in a laboratory is not supervising an agentic system in a firm. Proposition 5 is stated with that limitation attached and its refutation condition names it. The managerial statement of this section's result can now be made without overclaiming. The question a board should ask about an agentic deployment is not whether the firm is on the loop. Section 3 established that loop position is determined by throughput rather than declared by policy, so the answer to that question is not the board's to give. The question is which decisions the firm has succeeded in moving into the class whose oversight cost does not grow — which constraints are compiled into the action space rather than checked afterwards, which routing rules are fixed in advance rather than applied case by case, which limits on the system's remit have been set rather than monitored. And for everything that remains in the control class, which will always be something, the question is whether the firm can state the arrival rate, the interaction time and the uncommitted attention that would make its claimed oversight solvent. A firm that can answer the first question has done the design work. A firm that can answer the second has priced what it is claiming. A firm that can answer neither is not on the loop; it is describing itself as on the loop, and Section 9 examines what that description accomplishes. 9. Nominal Oversight as Liability Allocation 9.1 The turn in the argument Everything to this point has treated nominal oversight as a control failure. Sections 3 to 5 established that loop position is determined by measured latency rather than declared policy, that oversight capacity can become insolvent, and that improvement in the supervised system degrades the supervisor's detection performance. Section 6 established the narrow conditions under which oversight is nonetheless an accuracy control, and Sections 7 and 8 that a loop cannot be closed by a verifier statistically dependent on its generator. On that account, a nominal overseer is simply a control that does not work: a defect to be engineered out. This section argues that the account is incomplete, and in some deployments inverted. Nominal oversight is also — and sometimes primarily — an allocation of liability. The designation of an accountable human is not only a failed attempt to install a control; it is a successful attempt to install a defendant. And the two facts are connected in a way that ought to trouble anyone designing an oversight regime in good faith. A nominal overseer is useful precisely because they absorb responsibility without exercising control. An overseer with genuine control would generate friction, Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 63 refusals, delay and a documented record of institutional disagreement. An overseer without control generates none of these and still satisfies the formal requirement that some natural person be answerable. The organisational value of the arrangement is therefore not diminished by the overseer's incapacity. It is created by it. This is not an accusation of bad faith against firms that adopt human-on-the-loop language; most such adoptions are sincere. The claim is structural: an arrangement that produces liability without control is stable whether or not anyone intends it, because it is cheaper than the alternative and satisfies the same external audit. Section 9.5 states this as Proposition 13. 9.2 The conceptual apparatus: tracking, tracing, responsibility gaps and crumple zones The philosophical apparatus for this argument is well developed and predates the current deployment wave. The most useful formulation is Santoni de Sio and van den Hoven (2018), Frontiers in Robotics and AI 5, Article 15, peer-reviewed, who decompose meaningful human control into two conditions. The first is tracking: In order to be under meaningful human control, a decision-making system should demonstrably and verifiably be responsive to the human moral reasons relevant in the circumstances—no matter how many system levels, models, software, or devices of whatever nature separate a human being from the ultimate effects in the world, some of which may be lethal. That is, decision-making systems should track (relevant) human moral reasons. The second is tracing: In order for a system to be under meaningful human control, its actions/states should be traceable to a proper moral understanding on the part of one or more relevant human persons who design or interact with the system, meaning that there is at least one human agent in the design history or use context involved in designing, programming, operating and deploying the autonomous system who (a) understands or is in the position to understand the capabilities of the system and the possible effects in the world of the its use; (b) understands or is in the position to understand that others may have legitimate moral reactions toward them because of how the system affects the world and the role they occupy. Two features of the account are load-bearing for this paper and are routinely lost when it is cited. First, the account derives from Fischer and Ravizza's notion of guidance control in the free-will literature: it is a philosophical theory of when an agent's control over an outcome is of the kind that grounds responsibility, imported into engineering, not a piece of engineering doctrine dressed in philosophical vocabulary. Second, and more consequentially, it is deliberately cast as a set of design requirements. Meaningful human control on this account is a property of the sociotechnical system, not a property of the person at the console. Nothing an operator does at the moment of decision can satisfy the tracking condition if the system was not built to be responsive to the relevant moral reasons. This is why "put a human in front of it" is not a remedy: it addresses the identity of the person occupying the role and leaves the design property untouched. Behind both conditions stands Matthias (2004), Ethics and Information Technology 6(3):175–183, peer-reviewed, and the responsibility gap. Matthias argues that for autonomous learning machines the manufacturer or operator is in principle not capable of predicting future machine behaviour, and thus cannot be held morally responsible or liable for it; society must either abandon such machines, which he treats as unrealistic, or face a gap that cannot be bridged by traditional concepts of responsibility ascription. The modifier "in principle" carries the whole argument and is the reason the gap cannot be closed administratively. Matthias's claim is not that the operator lacks information which better logging, better documentation or better model cards would supply. It is that the control condition of classical responsibility ascription fails for systems whose behaviour is Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 64 determined by post-deployment learning. An organisation cannot document its way out of a gap that is constitutive rather than epistemic. Nissenbaum (1996), Science and Engineering Ethics 2(1):25–42, peer-reviewed, had already catalogued four barriers by which computerised systems erode accountability: the problem of many hands, in which distributed authorship obscures culpability; the problem of bugs, in which treating software error as an inevitable natural fact of programming discourages holding anyone accountable; blaming the computer, in which failure is attributed to the machine as agent and human responsibility is thereby displaced; and software ownership without liability, in which producers retain property rights while contractually disclaiming responsibility for harm. This paper's argument is that mandated human oversight is a fifth mechanism of the same family, and that it is the most efficient of the five. The first four erode accountability by dispersing or deflecting it; the fifth manufactures a nominal accountable party without transferring the control that would make accountability meaningful. It is more efficient than its predecessors because it does not require anyone to argue that no one is responsible. It supplies an answer to the question "who is accountable?" that survives audit, litigation and press enquiry, while leaving the causal structure of the failure entirely undisturbed. It is worth recording that Nissenbaum identified the first four thirty years ago, before machine learning, and that nothing in the intervening period has retired any of them. Elish (2019), Engaging Science, Technology, and Society 5:40–60, peer-reviewed, gives the phenomenon its name. In a car, the crumple zone absorbs the impact to protect the occupant; in a highly automated system, the human operator becomes the moral crumple zone, and responsibility "may be misattributed to a human actor who had limited control over the behavior of an automated or autonomous system", thereby shielding the technical system and its designers from scrutiny. The mechanism she identifies is structural rather than accidental: as control is distributed across time, space and many hands, the nearest human absorbs the blame precisely because they are legible and locatable, while the system's diffuse contributions are not. Blame flows to whoever can be named. The research underlying this paper verified the pairing of Elish's cases — a historical case in commercial aviation autopilot, anchored on Air France 447, and a contemporary case concerning the 2018 Uber self-driving fatality in Tempe, Arizona and its safety operator — from the journal record and secondary summaries only, because the article body could not be opened from two hosts. Those cases are therefore referred to here in general terms, and no claim about her treatment of them is quoted or attributed. Composing the two accounts yields the formulation this paper uses throughout: Elish's moral crumple zone is what you get when the tracing condition is enforced against someone for whom the tracking condition was never satisfied. A person is held to have understood the system's capabilities and to have occupied a role attracting legitimate moral reactions, in a system that was never built to be responsive to the moral reasons at issue. The enforcement is real; the control it presupposes never existed. Section 9.3 gives the documented instance of that composition, together with two cases in which the tracking condition failed by a different route. 9.3 Three documented failures of the tracking condition These are the empirical anchors of the argument, and each is given with its evidentiary grade, because they differ materially in what kind of record they are. They differ in something more consequential than grade, and the difference is stated with each case rather than smoothed over: only the first is an instance of tracing enforced without tracking, which is the composition Section 9.2 just derived and the pathology Proposition 13 describes. The other two are failures of the tracking condition brought about deliberately — in one by policy, in the other by an evidential rule Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 65 — and they are included because they establish that the absence of tracking is more often designed than accidental, which is a different claim and separately important. (a) The Netherlands childcare benefits affair. The primary account is Ongekend onrecht ["Unprecedented Injustice"], the report of the Childcare Allowance Parliamentary Inquiry Committee of the House of Representatives of the States General, 17 December 2020 — an official parliamentary inquiry report, authoritative but not peer-reviewed. Between 2012 and 2019 approximately 25,000 to 35,000 people were designated "malicious" or grossly negligent, and 94% of those designations were subsequently found unjustified. The operating rule was all-or-nothing: minor administrative errors triggered recovery of the entire annual allowance, and, in the committee's words, "the proportionality principle was effectively cast aside until well into 2019". Alongside the inquiry sits a concluded enforcement action, which is a primary regulatory record of a different and stronger kind: in December 2021 the Autoriteit Persoonsgegevens fined the Belastingdienst €2.75 million, finding among other things that it "used applicants' nationality (Dutch/ not Dutch) as an indicator in a system that automatically designated certain applications as risky". This is not a recommendation or a finding of poor practice; it is a monetary penalty on a completed investigation. The anatomy of the oversight failure comes from Amnesty International's Xenophobic Machines (2021), which must be labelled NGO grey literature and used for mechanism rather than as an independent evidentiary authority. Its most important finding for this paper is that civil servants did manually review the applications the model flagged — human oversight formally existed throughout the six years — but "the civil servant (as the system user) did not have access to any details about what information had been used as the basis for assigning a specific risk score to an applicant". The reviewer could confirm a fraud designation regardless of the severity of the underlying error while being blind to the reasoning that produced the selection. This is the cleanest documented instance of tracing without tracking available in the public record: a human in the loop, formally accountable, and structurally incapable of oversight. It is the on-point case, and the only one of the three in which a formally present human was held to a control they could not exercise; the pathology here is tracing enforced against a person for whom tracking was never built. (b) Australia's Robodebt. The Report of the Royal Commission into the Robodebt Scheme, Commissioner Catherine Holmes AC SC, tabled 7 July 2023, is an official royal commission report and is not peer-reviewed. Its recommendation directly on point is to "Increase review and oversight of automated decision making". The mechanism was income averaging — annual tax income data averaged across fortnights to impute debts — combined with the removal of the manual review step that had previously verified debts before they were raised, and a reversal of the onus onto recipients to disprove the computed figure. The peer-reviewed analysis is Whiteford (2021), Australian Journal of Public Administration 80(2): 340–360; the paper is frequently miscited to the Australian Journal of Social Issues, and the correct journal is noted here in passing. Whiteford documents repayment of more than A$1,000 million to more than 400,000 people following the largest class-action settlement in Australian history, after a Federal Court ruling that the policy was unlawful. His analytically sharper claim is that Robodebt "differs from other examples of policy failures in that it was intentional, and not the result of mistakes in design or implementation". The distinction this paper needs follows directly. The automation did not cause the harm through error. It operationalised a legally unsupported policy at scale, and the removal of individual human assessment was the point rather than a side-effect. An argument that the algorithm was buggy would be weaker and would also be wrong. The pathology is accordingly not the Dutch one: no human was formally present in the decision at all, because the Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 66 tracking condition was removed as a matter of policy rather than left unbuilt, and Robodebt therefore evidences the deliberate elimination of oversight rather than its nominal retention. The widely quoted characterisation of the scheme as "a crude and cruel mechanism" is not reproduced here, because the research pass could not verify it against the report itself. (c) The United Kingdom's Post Office Horizon scandal. The Post Office Horizon IT Inquiry, Final Report Volume 1, Sir Wyn Williams, HC 1119, 8 July 2025, is a primary inquiry record. It finds that throughout the lifetime of Legacy Horizon the Post Office "maintained the fiction that its data was always accurate" despite the system being "capable of error". Approximately 1,000 persons were prosecuted and convicted on Horizon evidence; approximately 10,000 eligible claimants entered redress schemes; and 13 persons died by suicide in connection with Horizon shortfalls. The legal-mechanism point is this subsection's most important, and it generalises beyond the case. England and Wales operated a common-law presumption that computer evidence is reliable, following the repeal of section 69 of the Police and Criminal Evidence Act 1984 in 1999. The effect of the presumption was to shift the burden onto sub-postmasters to prove that a system they could not inspect was wrong. Human oversight here was not merely absent; it was legally disfavoured. A subpostmaster who exercised precisely the independent judgement that oversight doctrine calls for — who disbelieved the machine — bore the evidential burden for doing so. The pathology here is a third one, distinct from both of the others: tracking was neither unbuilt nor withdrawn but legally disfavoured, the law of evidence attaching a cost to the exercise of independent human judgement and thereby removing the condition by operation of a presumption rather than by an operational decision. Horizon is on that account the strongest of the three cases for the proposition that the absence of tracking is designed, and the weakest for the proposition that a nominal overseer absorbs liability, since the persons prosecuted were not overseers of the system at all. The research pass did not verify what Volume 1 specifically recommends about that presumption, and no recommendation on it is attributed to the Chair here. 9.4 Two cases in which humans were present, engaged, and too slow The three cases above concern configurations in which the tracking condition failed, and in which such humans as were present lacked information or standing. The two below concern overseers who lacked time, and they are included because they refute the natural objection that the failures in Section 9.3 were failures of diligence. Knight Capital. The record is the SEC administrative proceeding, Release No. 34-70694, 16 October 2013, a primary regulatory record. Erroneous orders flowed for approximately 45 minutes, producing four million executions in 154 stocks for more than 397 million shares and a $460 million loss, with a $12 million civil penalty. The order finds that "Knight's system continued to send millions of child orders while its personnel attempted to identify the source of the problem". The detail that matters most for an oversight paper concerns what those personnel did: during troubleshooting the team uninstalled the new code from the seven servers where it had been correctly deployed, which worsened the problem by triggering the malfunction across all of them. The humans were in the loop and were actively working the incident. For 45 minutes they could neither diagnose nor stop it, and their first corrective action amplified the loss. No amount of attentiveness substitutes for a diagnostic channel adequate to the system's rate of consequence. The 2010 Flash Crash. The record is the joint report of the Commodity Futures Trading Commission and the Securities and Exchange Commission of 30 September 2010, also a primary regulatory record. An automated program sold 75,000 E-Mini S&P 500 futures contracts, about $4.1 billion, "programmed to feed orders into the June 2010 E-Mini market to target an execution rate set to 9% of the trading volume calculated over the previous minute, but without regard to price or time". The Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 67 acute phase ran from 14:41 to 14:45:28. The same trader's previous comparably sized program had taken more than five hours, so the same economic action was compressed roughly fifteen-fold by removing price and time awareness from the execution logic. The report also records that manual override delays ranged "from as short as a few seconds to as long as several hours". The governance point is not that the algorithm malfunctioned. The algorithm was not buggy; it executed its specification exactly. The specification omitted a constraint that a human executing the same trade would have supplied implicitly, namely that one does not continue selling into a collapsing book without regard to price. That is the same failure mode the 2026 agent benchmarks describe as losing track of constraints, reproduced at a $4.1 billion scale sixteen years earlier. The constraint a human supplies implicitly is exactly the constraint that a specification omits silently, and the omission becomes visible only at the throughput at which no human can intervene. 9.5 Liability does not stay where the firm puts it A firm may reasonably ask what it gains from this analysis, given that liability is placed by law rather than by organisational design. The answer is that placement is not fully within the firm's gift, and the clearest recent illustration is Moffatt v. Air Canada, 2024 BCCRT 149, decision dated 14 February 2024 — a primary tribunal record, though one which the research underlying this paper reached through secondary legal analyses rather than directly, because CanLII blocks automated retrieval. For that reason the tribunal's reasoning is not quoted verbatim anywhere in this paper. The airline argued in substance that the chatbot was a separate legal entity responsible for its own actions. The tribunal rejected this and held Air Canada liable in negligent misrepresentation, reasoning that the company is responsible for all information on its website whether it comes from a static page or from a chatbot, and that a customer cannot be expected to cross-reference one part of a website against another to determine which statement is authoritative. Damages were approximately CAD $650. The triviality of the award is precisely the point. Nothing turns on the quantum; what the decision establishes is the allocation. An organisation deploying a customer-facing autonomous agent owns every statement that agent makes, and cannot allocate that liability to the agent, to the vendor, or to a disclaimer located elsewhere in its own estate. The attempted allocation was not merely unsuccessful — it was rejected as a category error. A firm that designs its oversight arrangements on the assumption that the designated overseer, the model vendor or the terms of service will absorb the consequence of an autonomous action is designing against a proposition that at least one tribunal has already declined to accept. Proposition 13 (Tracing without tracking). Designating an accountable human where the tracking condition fails does not produce accountability; it produces a liability allocation that protects the system's designers and operators from scrutiny. The stronger the formal accountability attached to a nominal overseer, the weaker the incentive to repair the conditions that made oversight nominal. Refutation condition. Refuted if a case is documented in which formal accountability attached to an overseer lacking independent channel, latency headroom or authority was followed by measured improvement in system-level error rates attributable to that designation. The proposition's second sentence appears to contradict Section 10.5, and the apparent contradiction is worth resolving here rather than leaving for a reader to find. Section 10.5 concludes from the Boeing record — a board that publicly represented that it monitored a safety risk in ways it did not, on a matter that settled for $237.5 million — that claiming oversight one does not perform is Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 68 itself part of what the violation consists in — a step Section 10.5 advances in express dependence on the misrepresentation element of that record as reported at secondary level, no slip opinion having been obtained, and within the boundary Section 10.5.1 states, so that what is reached is the accuracy of a representation about the control environment and not any operational failure of the system described — which makes the incentive to repair the conditions of oversight stronger rather than weaker. Proposition 13 says the opposite about what is, on its face, the same act. The two comparative statics are reconcilable, and they differ by the level at which the accountability lands. Where the formal accountability is discharged by naming an individual overseer, the marginal return on repairing the conditions of oversight falls, because the liability has already been placed and the designation has already bought what external pressure was demanding; that is Proposition 13's case, and the four failure conditions of Definition 2 are then investments with compliance value already obtained. Where the accountability attaches instead to the board as a duty to design and operate a monitoring system — Delaware's oversight doctrine under Caremark, Stone, Marchand and Boeing, and Japan's internal-control duty under ไผš็คพๆณ•็ฌฌ362ๆก็ฌฌ4้ …็ฌฌ6ๅท (a numbering taken from the secondary literature and not confirmed against the primary text, as Section 10.6 records), both set out in Sections 10.5 and 10.6 — the return on repair is restored, because no individual designation discharges a duty owed at system-design level. Naming an overseer does not satisfy a requirement that the monitoring system be board-level, dedicated and actually used, and under the Japanese standard it does not make an officially catalogued failure mode less foreseeable. This identifies board-level oversight duty as the one mechanism located anywhere in the corpus surveyed in Section 10 that reverses the perverse incentive Proposition 13 describes, which is a stronger and more useful result than the proposition alone, and it is why Section 10.5 is the operative legal finding of this paper rather than an adjunct to it. It should be added, since the proposition must be readable against its own test, that the refutation condition stated above is difficult to satisfy in practice: it demands an error-rate improvement attributable to a designation, and attribution of a system-level error rate to one organisational act is rarely available in any published record. That difficulty is a weakness of the proposition rather than a strength, since a condition unlikely ever to be met is a condition that protects the claim from the evidence. The second sentence of the proposition is its perverse-incentive corollary, and it deserves to be stated at length because it is the practical consequence for management. Once formal accountability has been attached to a designated overseer, the organisation has already purchased the thing that regulatory, contractual and reputational pressure was demanding: an identifiable answerable person. Every subsequent investment in the conditions that would make that person's oversight real — an information channel independent of the system's own inputs, latency headroom sufficient for the decision class, unencumbered authority to halt, and uncommitted attention exceeding the arrival-rate-weighted cost of the exceptions routed to them, which are the four failure conditions of Definition 2 — is an investment with no marginal effect on the firm's compliance position and a substantial marginal cost in throughput. The liability has already been placed. Repairing the oversight buys nothing that the designation has not already bought, and it forgoes the throughput gains that motivated the automation in the first place. The evidentiary base of the proposition must be stated at its true size, because Section 9.3 supplies less than it appears to. Proposition 13's direct empirical anchor is the Dutch childcare benefits affair alone, the only one of the three cases in which a formally accountable human was actually present and structurally incapable of oversight; Robodebt and Horizon establish the adjacent and separately important claim that the absence of tracking is commonly deliberate rather than accidental, but neither documents a nominal overseer absorbing liability, and neither can be counted toward the proposition it is often taken to support. A proposition resting on a single documented instance is Human-on-the-Loop · WP No. 8 9. Nominal Oversight as Liability Allocation 69 weak, whatever the quality of that instance, and it is weak in a way that the mechanism argument surrounding it cannot repair; Section 11 records the case selection as one of the paper's principal limitations, since the three cases are in the record precisely because they failed badly enough to produce a parliamentary inquiry, a royal commission or a regulatory order. This is why nominal oversight is stable rather than transitional. It is not a state that organisations pass through on the way to real oversight; it is an equilibrium, and it is reached by ordinary cost minimisation rather than by anyone deciding to launder responsibility. It is also why the measurement proposals in Section 10.7 are addressed to observable quantities rather than to intentions. A regime that requires a firm to publish the latency, the exception arrival rate, the utilisation and the override rate makes the difference between real and nominal oversight visible from outside the firm, and thereby restores the marginal return on repair that the liability allocation destroyed. Section 10 turns to what the law currently requires, what it conspicuously does not require, and what a firm that wishes to make an honest claim should measure. 10. The Legal Position and What a Firm Should Measure 10.1 What the law actually requires The most detailed statutory treatment of human oversight is Article 14 of the EU AI Act, Regulation (EU) 2024/1689, OJ L, 2024/1689, 12.7.2024, in force since 1 August 2024. Article 14(1) states the provider's design duty: high-risk AI systems "shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use". Article 14(2) supplies the purpose clause: human oversight "shall aim to prevent or minimise the risks to health, safety or fundamental rights that may emerge when a high-risk AI system is used in accordance with its intended purpose or under conditions of reasonably foreseeable misuse, in particular where such risks persist despite the application of other requirements set out in this Section" — a clause that positions oversight as the residual control, operating on the risk left over after every other requirement has been satisfied. Article 14(4) enumerates five capabilities, not four, which the provider must enable the oversight person to exercise "as appropriate and proportionate": (a) "to properly understand the relevant capacities and limitations of the high-risk AI system"; (b) "to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system"; (c) "to correctly interpret the high-risk AI system's output"; (d) "to decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output"; and (e) "to intervene in the operation of the high-risk AI system or interrupt the system through a 'stop' button or a similar procedure". Subparagraph (b) is the only place in the corpus surveyed for this paper where automation bias is named, in a binding instrument, as a hazard the system's design must counteract; whether some other provision of Union law names it too was not established here. The law thereby concedes that a nominal human is a foreseeable failure mode of the arrangement it mandates, and does not treat over-reliance as an operator failing to be corrected by exhortation: counteracting it is a design obligation on the provider — and the meta-analytic null on explanations and confidence displays reported in Section 6 is evidence that the most commonly offered discharge does not work. Article 14(5) is the only quantitative rule. For remote biometric identification, the deployer is to take no action or decision on the basis of an identification unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority. The requirement that the number of persons be at least two is load-bearing for this section, so its provenance is stated alongside it rather than in a footnote: the text of Article 14 was Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 70 read from a faithful mirror of the Official Journal and not from the Official Journal itself, and while that mirror is a standard practitioner resource, the figure should be confirmed against the OJ PDF by anyone who relies on it. Subject to that, it is the only numeric human-oversight requirement located anywhere in the corpus surveyed for this paper, and even it is disapplied where Union or national law considers its application disproportionate in law enforcement, migration, border control or asylum. A rule specifying n ≥ 2 in one product category and switched off in the four domains where individual stakes are highest is not a quantification of oversight; it is the exception that measures the absence of one. Article 26 shifts the duty to the deployer. Article 26(2) provides that deployers "shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support" — a four-element test, two of whose elements, authority and support, are organisational rather than cognitive. Competence and training concern what is in the overseer's head; authority and support concern what the organisation has given them, which is what separates real oversight from a rubber stamp under Definition 2. Article 26(3) leaves those obligations without prejudice to "the deployer's freedom to organise its own resources and activities". That combination — a binding duty to confer authority and support, with discretion over resourcing — is the precise seam at which human-on-the-loop becomes a management problem rather than a compliance checkbox. The law states what the overseer must have and is silent on how much of a finite attention budget the firm must spend supplying it; the firm must resolve that trade-off, and can only resolve it by measuring. 

On verification: every passage of Articles 14 and 26 reproduced above was retrieved from a faithful mirror of the Official Journal rather than from EUR-Lex directly, and should be checked against the OJ PDF before formal submission. 10.2 The timing, which a paper dated August 2026 must get right Article 14 does not currently bind. Regulation (EU) 2026/1744, the Digital Omnibus on AI, amends Regulation (EU) 2024/1689 and postpones the high-risk obligations. The two halves of that statement rest on different evidence and are separated here accordingly. The instrument's existence, its full title, its adoption on 8 July 2026 and its entry into force on 27 July 2026 were confirmed directly at EUR-Lex. The postponed dates were not: on a single reading of the amending text, the Annex III obligations, which include Articles 14 and 26, move from 2 August 2026 to 2 December 2027, and the Annex I obligations from 2 August 2027 to 2 August 2028. Those two dates are reported on that reading and on no other authority. Recital (40) gives the stated rationale: "delayed availability of standards, common specifications … and the delayed establishment of national competent authorities". That is an official admission that the oversight rules outran the tooling meant to make them operable — the Union enacted a duty to oversee before the harmonised standards which would tell a firm what oversight consists of, and before the authorities that would assess whether it had been performed. The correct framing is therefore that Article 14 is enacted binding law which does not yet apply, not a currently enforceable obligation: citable as a legal standard, but not today a source of enforcement risk. The date on which the paper proceeds is 2 December 2027, which is the date read from the amending instrument in the single retrieval described above and is not independently corroborated; it should be re-checked against the Official Journal before submission. Nothing in the argument turns on the date being that one rather than another, since what the argument uses is that the obligation is enacted and not yet applicable. Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 71 10.3 Japan Japan's principal AI statute is AIๆŽจ้€ฒๆณ• (the Act on the Promotion of Research, Development and Utilisation of AI-Related Technologies), Act No. 53 of 2025, promulgated 4 June 2025 and fully in force 1 September 2025. It comprises 28 articles and has no penal provisions; it imposes on business operators a duty to cooperate with State measures, not to achieve any outcome, so a firm seeking an enforceable oversight standard in it will not find one. The Second-Term AI Basic Plan (ไบบๅทฅ็Ÿฅ่ƒฝๅŸบๆœฌ ่จˆ็”ป๏ผˆ็ฌฌโ…กๆœŸ๏ผ‰) was adopted by Cabinet decision on 14 July 2026; the research pass could not extract its substantive language on human supervision, and nothing is attributed to it here. The operative Japanese material is soft law. AIไบ‹ๆฅญ่€…ใ‚ฌใ‚คใƒ‰ใƒฉใ‚คใƒณ (AI Guidelines for Business) version 1.2, issued jointly by ็ทๅ‹™็œ (Internal Affairs and Communications) and ็ตŒๆธˆ็”ฃๆฅญ็œ (METI) on 31 March 2026, is non-binding. It extends coverage to AIใ‚จใƒผใ‚ธใ‚งใƒณใƒˆ (AI agents), placing it directly on the configurations this paper analyses, and item C-1โ‘ก names ่‡ชๅ‹•ๅŒ–ใƒใ‚คใ‚ขใ‚น (automation bias) as a risk requiring countermeasures — converging independently with EU AI Act Article 14(4)(b), two instruments in different traditions identifying over-reliance as the characteristic failure mode of nominal oversight. Against that convergence the guideline's verb choice must be reported: item C-3โ‘ก, on fairness, provides that the user should consider (ๆคœ่จŽใ™ใ‚‹) interposing human judgement rather than letting the AI decide alone — ๆคœ่จŽใ™ใ‚‹, not ่ฌ›ใ˜ใ‚‹ (implement). The Japanese guideline stops short of requiring human intervention even in the fairness context in which the case for it is strongest. The most useful Japanese primary source is the Financial Services Agency's AIใƒ‡ใ‚ฃใ‚นใ‚ซใƒƒใ‚ทใƒงใƒณใƒšใƒผ ใƒ‘ใƒผ (AI Discussion Paper) version 1.1 of March 2026, which is explicitly non-binding and states that it does not set out supervisory expectations. Its value is empirical. It reports that firms overwhelmingly do not present generative AI output directly to customers, human judgement being interposed in most cases, and, at page 18, that for ็จŸ่ญฐๆ›ธ — credit memoranda, which can materially affect customers — an operating practice requiring ไบบใฎ้–ขไธŽ๏ผˆHuman in the loop๏ผ‰ is predominant. A national financial regulator has here used the English term untranslated in an official document, and reported it as observed industry practice rather than as a regulatory demand. That is descriptive supervisory intelligence, more informative than exhortation, because it tells a firm what its peers actually do and thus the baseline against which a deviation would have to be justified. 10.4 Standards specify process, never quantity The standards corpus adds detail without adding measurement. The NIST AI Risk Management Framework 1.0 (January 2023) is voluntary; its relevant subcategories are GOVERN 3.2, "Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems"; MAP 3.5, "Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function"; MANAGE 2.4, requiring mechanisms in place and applied, "and responsibilities are assigned and understood", to "supersede, disengage, or deactivate" AI systems performing inconsistently with intended use; and MANAGE 4.1, on post-deployment monitoring including appeal and override. MAP 3.5 is the cleanest statement anywhere that oversight must be defined, assessed and documented rather than assumed; MANAGE 2.4 is the NIST analogue of Article 14(4)(e), adding what the EU text omits — responsibilities assigned and understood. The Generative AI Profile, NIST AI 600-1 (July 2024), also voluntary, adds suggested actions keyed to those subcategories. GV-3.2-001 provides that "Policies are in place to bolster oversight of GAI systems with independent evaluations or assessments", independence appearing here as a governance rather than, as in Section 7, a statistical requirement. Two further items are subtler. Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 72 MP-3.4-004 instructs organisations to "Delineate human proficiency tests from tests of GAI capabilities", implicitly recognising that a competent system paired with an untested overseer is not a controlled configuration; MS-2.5-004 tells them to "Track and document instances of anthropomorphization in GAI system interfaces", a second over-trust vector alongside automation bias. On ISO/IEC 42001:2023 this paper must be blunt about a limitation and must then confine itself to what the limitation leaves standing. The standard is paywalled and its normative text was not read in this paper's research pass. What can be said is what its published scope and structure establish: it was published in December 2023 as a management-system standard for artificial intelligence, and it is certifiable, so that conformity with it is assessed by audit against an organisation's documented management arrangements. Beyond that the paper asserts nothing. It places no quotation marks around any ISO 42001 control language; it makes no claim about what the standard requires of an organisation in respect of human oversight; and it makes no claim about whether Annex A operates on a Statement-of-Applicability basis or about what would follow if it did. A reader who needs to know what ISO/IEC 42001 requires must read the standard, and this paper is not a substitute for doing so. The structural finding is the one the whole of Section 10 turns on, and it must be stated with its gap left visible. Among the instruments whose text this paper's research pass actually inspected, exactly one states a measurable human-oversight quantity: Article 14(5), n ≥ 2. Every other inspected instrument specifies process and documentation. ISO/IEC 42001:2023 enters the finding on neither side, because its normative text was not inspected: it is not counted as an instrument found silent, and no secondary source consulted reported a numeric human-oversight threshold in it. Subject to that exclusion, human oversight is a procedural requirement across the inspected corpus and a measurable one only in the single biometric case, which means that as a compliance claim it is very nearly unfalsifiable: no quantity's absence would refute it, and a firm can satisfy every process requirement recorded in Table 8 while operating a channel that is insolvent under Definition 4 at every hour of every day. Section 10.7 proposes to close that gap voluntarily, because nothing in the inspected corpus closes it compulsorily. Instrument Binding status Specified as process Specified as quantity EU AI Act, Reg. (EU) 2024/1689, Art. 14 Binding but not yet applicable; date read as 2 Dec 2027 from Reg. (EU) 2026/1744 on a single retrieval Provider design duty (14(1)); purpose clause (14(2)); five enabled capabilities (14(4)), including overreliance awareness (b) and stop function (e); article text read from an OJ mirror None EU AI Act, Art. 14(5) As above Separate verification before action on remote biometric ID n ≥ 2 natural persons, read from an OJ mirror; disapplied where deemed disproportionate in law enforcement, migration, border control or asylum EU AI Act, Art. 26(2)–(3) As above Assign oversight to persons with competence, training, authority and support; resourcing left to the firm None AIๆŽจ้€ฒๆณ• (Act No. 53 of 2025), Japan Binding statute; 28 articles; no penal provisions Duty on business operators to cooperate with State measures None Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 73 Instrument Binding status Specified as process Specified as quantity AIไบ‹ๆฅญ่€…ใ‚ฌใ‚คใƒ‰ใƒฉใ‚ค ใƒณ v1.2 (31 Mar 2026), Japan Non-binding soft law C-1โ‘ก countermeasures against ่‡ช ๅ‹•ๅŒ–ใƒใ‚คใ‚ขใ‚น; C-3โ‘ก consider (ๆคœ่จŽใ™ ใ‚‹) interposing human judgement for fairness None ้‡‘่žๅบ AIใƒ‡ใ‚ฃใ‚นใ‚ซใƒƒ ใ‚ทใƒงใƒณใƒšใƒผใƒ‘ใƒผ v1.1 (Mar 2026) Explicitly nonbinding; sets no supervisory expectations Observed practice: ไบบใฎ้–ขไธŽ ๏ผˆHuman in the loop๏ผ‰ predominant for credit memoranda (p. 18) None (descriptive, not prescriptive) NIST AI RMF 1.0 (Jan 2023) Voluntary GOVERN 3.2 roles; MAP 3.5 oversight defined, assessed, documented; MANAGE 2.4 supersede/disengage/deactivate, responsibilities assigned and understood; MANAGE 4.1 postdeployment monitoring None NIST AI 600-1, Generative AI Profile (Jul 2024) Voluntary GV-3.2-001 independent evaluations; MP-3.4-004 separate human proficiency from GAI capability tests; MS-2.5-004 track anthropomorphization None ISO/IEC 42001:2023 Certifiable managementsystem standard, published Dec 2023 Not inspected — normative text paywalled and not read; no characterisation offered Not inspected; no claim made in either direction Delaware oversight duty (Caremark; Stone; Marchand; Boeing) Binding case law; in force today. Propositions well settled; slip opinions not obtained and nothing quoted Good-faith effort to install a boardlevel monitoring and reporting system, dedicated and used for mission-critical risk None ไผš็คพๆณ•็ฌฌ362ๆก็ฌฌ4้ … ็ฌฌ6ๅทใƒป็ฌฌ5้ …; ไผš็คพ ๆณ•ๆ–ฝ่กŒ่ฆๅ‰‡็ฌฌ100ๆก (numbering per the secondary literature, not confirmed against e-Gov) Binding statute; mandatory for ๅคงไผš ็คพ Board decision on internal control systems; components enumerated by ordinance; adequacy judged against ordinarily foreseeable wrongdoing None Table 8. Human-oversight requirements across instruments: process versus quantity. Entries record what this paper's research pass inspected; an instrument whose text was not obtained is marked as not inspected rather than as silent. 10.5 The duty that binds today is corporate law, not AI law While Article 14 waits for its compliance date, a duty of oversight already binds directors. In re Caremark International Inc. Derivative Litigation, 698 A.2d 959 (Del. Ch. 1996) established it in Delaware. Stone v. Ritter, 911 A.2d 362 (Del. 2006) located it within loyalty and good faith rather than care — which is why exculpation under DGCL §102(b)(7) does not apply — and set the two-prong bad-faith test on which liability turns. The prongs are stated here in this paper's own words, and the case is cited for the proposition rather than for any language: liability requires either that the directors failed altogether to implement any reporting or information system or controls, or that, having implemented such a system, they consciously failed to monitor or oversee its operation and Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 74 thereby disabled themselves from being informed of risks requiring their attention. That two-prong structure is well settled and is what the argument below uses. Marchand v. Barnhill, 212 A.3d 805 (Del. 2019) supplies the gloss that matters most here, and it too is taken as a proposition rather than as a form of words: the bottom-line requirement is that the board make a good-faith effort — that it try — to put in place a reasonable board-level system of monitoring and reporting, and the requirement bites hardest where the risk is essential and mission critical to the company. The good-faith-effort standard and the mission-critical category are the two settled propositions this section needs. No language from either decision is quoted here, because neither slip opinion was obtained: the formulations in which these propositions ordinarily circulate reached this paper through a partial Justia retrieval and a law-firm case summary, and a paper that cannot inspect the reporter should not present the reporter's words as its own evidence. In re Boeing Co. Derivative Litigation, 2021 WL 4059934 (Del. Ch. 7 September 2021), which settled for $237.5 million, applied that standard to the 737 MAX. The fact pattern is reported here as the same law-firm summary describes it, that summary being the only account of the record this paper obtained: the board had no committee charged with monitoring airplane safety, it had no means of receiving internal safety complaints, and it passively accepted management's assurances. The element on which this paper's argument turns is the last one the summary records — that the board publicly represented that it monitored the aircraft's safety in ways in which it did not. The doctrinal bridge, stated precisely: Delaware does not require the board to prevent the harm or to be right; it requires the board to try, and for mission-critical risks it requires the monitoring system to be board-level, dedicated and actually used. That much is settled doctrine and depends on no quotation. The further step — that claiming oversight one does not perform is itself part of what the violation consists in — rests on the misrepresentation element of the Boeing record as reported at secondary level, and is advanced with that dependence attached rather than as an independently verified holding, within the boundary Section 10.5.1 states. On that footing it is the nominal-oversight problem of Definition 2, litigated at board level. A firm that designates an AI overseer without independent information, latency headroom, authority or attention, and then describes that arrangement in its disclosures as human oversight, resembles the Marchand and Boeing fact patterns in the respect that made those records actionable — a resemblance in the arrangement described and in what was said about it, not a prediction that liability would follow from any operational failure of the system so described. Two qualifications are owed, and both are research limitations rather than rhetorical hedges. The first has been entered above in the sentences themselves: nothing is quoted from any of the four decisions, because none of the slip opinions was retrieved, and every proposition drawn from them is stated in this paper's words and cited to the case for the proposition alone. The second concerns the reach of the doctrine, and is a statement about a search rather than about the world. The search conducted for this paper used public web retrieval and not a subscription case-law database, and it located no Delaware decision applying the oversight doctrine to AI or algorithmic governance as at August 2026. Whether a decision exists that this method would not have reached is not established here; the extension is anticipated in the practitioner literature, and this paper neither asserts nor denies that it has occurred. A procedural point should be noted rather than relied upon, and it too is verified only at secondary level: Delaware SB 21, signed 25 March 2025, narrowed the DGCL §220 books-and-records demands by which oversight plaintiffs traditionally build a pleading-stage record — a headwind for pleading such claims rather than a substantive change to the oversight duty. Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 75 10.5.1 The boundary of the corporate-law claim The argument must be bounded before it is carried into Section 10.6, because the doctrine it rests on is narrower than the use commonly made of it. The threshold for liability under the Delaware oversight doctrine is bad faith: not negligence, and not error. A board that establishes a reporting system and attends to what that system reports does not incur liability because the system it monitors made a mistake, however costly. The two prongs set out above reach the absence of a system and the conscious failure to monitor one that exists; neither reaches operational failure as such. An operational AI failure, however serious, does not by itself expose a director to personal liability, and the threshold is famously difficult to meet. The claim this paper makes is therefore procedural rather than outcome-based. It is not that a firm whose AI system errs has breached a duty. It is that a firm which represents itself as exercising human oversight, while the arrangement it describes lacks the independent information channel, the latency headroom, the authority or the uncommitted attention that would make oversight possible, has made a representation about its own control environment whose accuracy the board is responsible for — and the quantities set out in Section 10.7 and Appendix B are what would make such a representation checkable. They do not establish that oversight was adequate; they establish what was claimed, in terms testable against records the firm already keeps. Three things this paper does not claim should accordingly be stated plainly. It does not claim that any Delaware court has applied the oversight doctrine to AI governance, which Section 10.5 already records as not established. It does not claim that operational AI error creates director liability. And it does not claim that failing the solvency condition of Definition 4 is itself a breach of any duty. A firm can be control-insolvent on this paper's definition and entirely compliant with its directors' duties, because the duty is to try and to have a system, not to succeed. Definition 4 is a management instrument and not a legal test. The corresponding limit applies on the Japanese side. The standard there is the adequacy of the internal control system against ordinarily foreseeable wrongdoing, which is likewise a question about the system rather than about the outcome — the company in ๆ—ฅๆœฌใ‚ทใ‚นใƒ†ใƒ ๆŠ€่ก“ไบ‹ไปถ was held not liable although the fraud occurred and the loss was real. The foreseeability argument of Section 10.6 accordingly goes to what a board ought now to anticipate when it designs that system, and not to whether any particular failure creates liability. What the narrowing leaves standing is the whole of what the argument needs, and saying so is not a retraction. On both sides of the comparison the enforceable object is the same: a documented, board-level reporting system that is actually used. Delaware reaches it through the requirement of a good-faith effort and Japan through adequacy against foreseeable wrongdoing, and both routes converge on a system rather than on a result. A firm that cannot state the quantities that would make its oversight claim true has not evidenced such a system; it has evidenced a description of one, which is the difference Definition 2 names. That claim survives the narrowing intact, and it is the only claim this paper needs from either jurisdiction. 10.6 The Japanese equivalent, and the paper's sharpest legal contribution Japanese law reaches the same territory by a different route, and the proposition it supplies is not in doubt even where this paper's verification of the sources is incomplete. The proposition is that the board of a company with a board of directors must decide on the establishment of systems ensuring that directors' execution of their duties complies with law and that the company's operations are otherwise proper; that for large companies (ๅคงไผš็คพ) that decision is mandatory rather than discretionary; and that the components of the system are enumerated by ministerial ordinance. The Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 76 provisions in which that proposition is located are cited in the standard Japanese corporate-law literature as ไผš็คพๆณ•็ฌฌ362ๆก็ฌฌ4้ …็ฌฌ6ๅท and ็ฌฌ5้ …, and ไผš็คพๆณ•ๆ–ฝ่กŒ่ฆๅ‰‡็ฌฌ100ๆก. The article, paragraph and item numbering is taken from that literature and was not confirmed against the e-Gov text in this paper's research pass, the item numbering of the enforcement ordinance in particular having been amended over time; a reader who needs the numbering should confirm it before using it. The judicial standard comes from ๆ—ฅๆœฌใ‚ทใ‚นใƒ†ใƒ ๆŠ€่ก“ไบ‹ไปถ, Supreme Court of Japan, First Petty Bench, judgment of 9 July 2009, ๆฐ‘้›†63ๅทป6ๅท1442้ , in which the company was held not liable because the control system it had built was adequate against ordinarily foreseeable wrongdoing and the fraud was perpetrated by means exceeding ordinary foreseeability. The formulation ใ€Œ้€šๅธธๆƒณๅฎšใ•ใ‚Œใ‚‹ใ€ in which that standard is conventionally rendered is likewise taken from the secondary literature and was not confirmed against the judgment text, so it is reported here as the literature's formulation and is not quoted from the report. ๅคงๅ’Œ้Š€่กŒ was not verified at all and is not relied upon. Those limitations do not damage the comparative argument, and it is worth saying at once why not, because the argument is this section's principal contribution. It requires only that the Japanese standard key on foreseeability — that the adequacy of an internal control system be judged against the wrongdoing a company ought ordinarily to anticipate. That much is uncontested in the secondary literature, and it is a property of the structure of the standard rather than of any particular phrasing of it. The argument would survive unchanged if the received formulation turned out to be worded otherwise, and it does not depend on the article numbering at all. On that basis the comparison is this. The Japanese standard makes foreseeability the operative variable; Delaware's makes good-faith effort the operative variable. Both converge on the same practical requirement — a documented system, proportionate to risk, that actually reports upward — but the Japanese key is the one currently moving, because foreseeability is not a fixed property of a hazard. It is constructed by the published record of what is known, and that record has changed in the last two years. Once automation bias is officially named as a risk — by EU AI Act Article 14(4)(b), in binding Union legislation, and independently by AIไบ‹ๆฅญ่€…ใ‚ฌใ‚คใƒ‰ใƒฉใ‚คใƒณ C-1โ‘ก, in guidance issued jointly by two Japanese ministries — the argument that over-reliance on AI output is not among the failure modes a company ought ordinarily to anticipate becomes progressively harder to sustain. A director cannot easily maintain that a failure mode named in the legislation of a major trading bloc and in his own government's guidance for his own industry lies outside ordinary foreseeability. The published guidance is itself constructing the foreseeability that the liability standard keys on — which is the mechanism by which non-binding soft law acquires binding force: not by becoming enforceable in its own right, but by supplying the factual predicate of a liability rule that already was. The conclusion follows in both jurisdictions at once. A firm that designates an AI overseer without an independent information channel, latency headroom, unencumbered authority or uncommitted attention is building, in Japanese law, a control system that does not address a failure mode which two governments have now catalogued in published guidance; and it is assembling, in Delaware, the arrangement whose description supplied the actionable element of the Marchand and Boeing records. Neither duty waits on the EU compliance date; both bind today. What binds is what Section 10.5.1 sets out and no more: an obligation to establish a system proportionate to the risk and to use it, and in neither jurisdiction an obligation to be right. The operational failure is not the breach; the unevidenced claim to be overseeing is what the two standards reach. Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 77 10.7 What a firm should measure, and what it should stop doing Eight measurement proposals follow, and a set of anti-patterns follows them. Each carries its evidentiary warrant, and none requires the firm to resolve the open questions catalogued in Section 11. They are proposals of the author, not findings. The first is to declare the loop position per decision class, by measurement rather than by policy. Under Definition 1, loop position is determined by the relation between the system's consequence latency τsys and the human's end-to-end response latency τH , both measurable properties of a deployed configuration. A firm should measure both per decision class, classify by the resulting inequality, and publish the classification and the measurement method. The warrant is Proposition 1 with the latency evidence of Section 3: where τsys has fallen below τH , a policy asserting on-theloop oversight asserts what the configuration cannot deliver, and is refutable by anyone with a stopwatch. A firm can no longer describe a configuration as supervised by describing its intentions. The second is to publish an oversight solvency statement. Definition 4 makes solvency computable, and Section 4.2 writes the first condition as a utilisation ratio: a configuration with exception arrival rate λ is solvent only if ρ = λ·(IT + c)/A ≤ 1, for IT the mean interaction time to closure, c the context re-acquisition cost and A the oversight time available per period. A firm should publish λ, IT, c and A, the implied span and the realised utilisation ρ, together with the override, dismissal and nonresponse rate. It is ρ that is the reportable number, being the single dimensionless figure that says how far the configuration sits from the boundary, and a firm reporting a value above one has reported its oversight insolvent without needing any further interpretation. That last matters most, because insolvency is not observed as refusal: Proposition 3 holds that an over-subscribed exception channel fails by accepting, and the override rates of 49–96% and 46.2–96.2% reported in Section 4 are what an insolvent channel looks like from outside. Publishing the override rate is what makes the solvency claim falsifiable. Section 4.7 adds one requirement to the publication and it should be read as part of this proposal: since none of A, IT, c and the detection probability is a constant within a period, the firm must state the point in the period at which each measurement was taken and should report ρ at least twice, early and late, the difference between the two being the fatigue term, which is the same requirement Appendix A.2 imposes on every term of the calculation and Appendix B on the solvency statement itself. The third is to maintain an independence register, recording for each control whether the verifier shares a base model, a training corpus, an input channel, a vendor or an incentive with the generator it verifies. This operationalises Definition 5, under which a verifier is independent for an error class only if its miss probability is unchanged by the generator having committed the error. Where a triage or routing layer selects which items reach a human at all, the register must record the same five relations for the router as for the verifier, because Definition 5 now bites one layer earlier: a triage layer built from the same base model as the generator, trained on the same corpus or reading the same input channel fails to route precisely the error classes the generator is most prone to commit, and fails to route them non-randomly. The second consequence of Section 8.3.1 establishes why the router's dependence is the more damaging of the two. An error a correlated verifier examines and misses is at least in the review record; an error a correlated router never selects is outside the examined stream altogether, recoverable by no diligence and no additional attention, and leaves no trace of the review that was never attempted. The warrant is Project Vend Phase 2 — a company technical report, grade C — in which the supervising CEO agent authorised requests about eight times as often as it denied them, which Anthropic explained by the supervisor sharing many of the deficiencies and blind spots of the agent it supervised, they being the same underlying model; reinforced by Monperrus's concession, in a preprint positioned as advocacy for Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 78 automated review, that correlated blind spots are a limitation of the approach. A register is cheap, and makes dependence visible at design time rather than at incident time. The fourth is to inject synthetic exceptions at elevated prevalence with immediate ground-truth feedback, and to report both the injection rate and the catch rate. In Wolfe and colleagues (2007), Journal of Experimental Psychology: General 136(4):623–638, peer-reviewed, the low-prevalence effect resisted every intervention tried but one: payoff manipulation had no effect, forced slowing had no effect, and only high-prevalence bursts with full feedback moved the miss rate at all. The warrant must be stated at its true strength, as Section 5.3 sets out. In Experiment 7 (N = 14) the lowprevalence miss rate was 21% against 18% in the matched high-prevalence condition, a difference not significant at p = .312; the 46% customarily set against that 21% is the low-prevalence baseline of Experiment 1a (N = 24), a different experiment with different observers, so the headline contrast is across experiments rather than within a design, and a failure to reject the null at N = 14 is weak evidence of elimination rather than a demonstration of it. The within-design replication that would establish it — high- and low-prevalence conditions with and without feedback, at adequate power — has not been run, and this proposal is accordingly in part a request that firms operating an injection regime generate it, since a firm running injections at scale produces exactly that design as a byproduct and need only report it. The recommendation nonetheless stands, because it is the only thing in the series that moved the miss rate in the right direction. A firm whose exception rate is falling because its models are improving is, by Proposition 7, watching its overseers' detection performance fall with it, and injection is the only measure that addresses that mechanism rather than exhorting against it. The fifth is to schedule returns to unassisted operation. The warrant is Parasuraman, Mouloua and Molloy (1996), Human Factors 38(4):665–679, peer-reviewed, in which a ten-minute scheduled return to manual control significantly improved automation-failure detection, the benefit persisting after automation resumed. This is the same mechanism as the protected unassisted practice proposed in the series' Brain Capital Management paper, and the connection should be made explicitly — with the equally explicit qualification, required by this series' self-citation discipline, that the Brain Capital Management proposal is a design proposal of the author's, not evidence, and is not counted as support here. The evidentiary weight rests on the 1996 experiment. The sixth is to order the decision predict-then-see: require the human to record their own judgement before the model's output is displayed, then show the output, then permit revision. The warrant is the Update condition in Green and Chen (2019b), PACM HCI 3(CSCW), Article 50, peer-reviewed, which reduced disparity in algorithmic influence by 81.5% and in deviation by 73.9% relative to showing the score first. The ordering of information rather than its content carried the effect, which is why the proposal is cheap. It has a second merit: recording the unassisted judgement generates, as a by-product, an independent channel, which the amended Proposition 8 identifies not as the condition under which oversight improves accuracy but as the strongest and most auditable source of the error decorrelation that is that condition. The recommendation survives the amendment for the reason Section 6.5 gives: decorrelation is not observable to a firm without labelled ground truth on the error class, whereas a firm can confirm cheaply that its reviewer saw evidence the model did not. The seventh is to locate the AI monitor correctly in the control architecture. Under the Institute of Internal Auditors' Three Lines Model (2020), an AI monitor is a first- or second-line control and can never be a third line, because third-line independence is elevated in that model to a principle in its own right, and is a property an instrumented component of the operation cannot possess. The human's oversight role therefore splits: part is second-line risk expertise within management, and part is governing-body-level determination of risk appetite, purpose and capital allocation, and only Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 79 the latter is what the constitutive oversight of Definition 6 describes. The reporting obligation comes from COSO's Internal Control—Integrated Framework (2013), whose Monitoring Activities component comprises Principle 16, "Conducts Ongoing and/or Separate Evaluations", and Principle 17, "Evaluates and Communicates Deficiencies". The dual-mode structure of Principle 16 is directly usable: automated monitoring can carry the ongoing evaluation, while the human's irreducible contribution is the separate evaluation plus the Principle 17 judgement of what constitutes a deficiency and to whom it must be escalated — a values judgement, not a throughput task. Both COSO and the IIA are professional frameworks rather than peer-reviewed research, and both full texts are paywalled; the principle names here were verified through secondary reproductions of the official lists. The eighth is to run a residual uniform sample of the unrouted stream wherever a triage or routing layer operates, sized to estimate the recall of the routing rule rather than to catch errors, and to report the resulting estimate of R alongside whatever workload reduction the triage layer is credited with. The warrant is the sixth step of Section 8.3.1. Where the routed volume lies within the attention budget, oversight efficacy is bounded above by R, so the router rather than the overseer is the binding constraint; and R is not observable from the routed stream, because everything the firm sees of its triage layer — the rule's precision, the outcomes of the reviews it commissions, the catch rate the overseer achieves within it — is conditioned on having been routed. Recall is a statement about what the router declined to route, and only a sample of that stream estimates it. Three properties of the instruction should be stated with it, because each is a cost a firm will otherwise discover late. The sample is a permanent cost of the triage design and not a transitional one: it does not fall away once the rule has been validated, since R is a property of a rule operating on a distribution that drifts, and a firm that retires the residual sample after a validation exercise has retired the only instrument that would tell it the rule had stopped working. A firm reporting a workload reduction without an accompanying recall estimate has reported a benefit without its price, which is the disclosure asymmetry this section exists to close: the saving is denominated in hours the firm no longer spends and appears in every operating report, while the cost is denominated in errors it no longer sees and appears in none. And where Proposition 9's counterfactual-ground-truth condition holds, the residual sample cannot recover R even in principle, since sampling the unrouted stream yields items whose correct disposition is itself unknown; a triage design in such a decision class is one whose central parameter is permanently unknown, and the firm operating it should record that fact in the same place it would otherwise have published the estimate, rather than publish a figure it has no means of forming. The anti-patterns follow, each with its refuting citation. The firm should stop relying on explanations to produce oversight: Bansal and colleagues (2021) found explanations added nothing over a bare confidence score across three tasks and raised agreement whether or not the model was right, with only 5% of participants using them to validate its reasoning; Vaccaro, Almaatouq and Malone (2024) found explanations non-significant as a moderator across 106 experiments; and Poursabzi-Sangdeh and colleagues (2021), pre-registered with N = 3,800, found that on unusual cases where the model was badly wrong, participants shown the clear model were worse at detecting the error. It should stop relying on confidence displays alone — and here a genuine tension must be recorded rather than resolved, since Vaccaro et al. found no moderator effect for confidence displays whereas Wickens and colleagues (2015) recommend that aids expose calibrated confidence, on the ground that "automation wrong" is more damaging than "automation gone". This paper does not resolve that tension; the defensible position is that a confidence display is not by itself an oversight control. It should stop adding a supervisor agent on the same base model, refuted by Project Vend Phase 2 and Definition 5. It should stop paying and instructing people to be more careful, which is the form of the Human-on-the-Loop · WP No. 8 10. The Legal Position and What a Firm Should Measure 80 anti-pattern the evidence actually refutes: Wolfe and colleagues found that neither payoff-matrix manipulation nor instructions to respond more slowly affected the low-prevalence deficit, and Parasuraman and Manzey (2010) found automation bias cannot be prevented by training or instructions, occurs in experts, and is not fixed by teams. The scope must be kept to what was tested, because one located finding points the other way: Szalma and colleagues (2006), whose bibliographic record was located but whose content this paper's research pass did not verify, report that feedback-based vigilance training improves vigilance performance, and while vigilance and automation bias are distinct constructs so that the two are not strictly in conflict, Section 5.5 records the tension and leaves it unresolved. It should stop using response-time service levels as a proxy for oversight quality, refuted by Louw and colleagues (2017), Accident Analysis & Prevention 108:9–18, N = 75, who found take-over time and the quality of the avoidance response largely independent; a latency service level measures whether the overseer answered, not whether the answer was good. And it should stop adding an oversight duty without removing work, refuted by Proposition 4: attention being fixed in the short run, an oversight duty assigned to an unchanged workload reallocates oversight capacity rather than adding it. 11. Limitations, Falsification, Competing Explanations, and Interests 11.1 The transfer problem, stated first because it is the principal weakness The limitation is best stated at the level of the claim rather than at the level of the data, because stated at the level of the data it invites a remedy that does not exist. The propositions advanced in this paper are comparative-statics claims about directions and orders of magnitude, and they are not point predictions. Each asserts the sign of a partial derivative and none asserts a magnitude: that detection probability falls as prevalence falls; that the share of consequential error removed by control oversight falls as throughput rises, and that the attention required to arrest that fall grows faster than throughput; that the span available to an overseer tightens as exogenous escalations rise; that adding a duty to an unchanged workload reallocates attention rather than adding it. This is a scope condition and not merely a caveat, and it should be applied to every proposition in Table 9 uniformly. The transferred parameters bear on how much and not on whether. A reader who accepts every mechanism set out here and rejects every number is left with the paper's thirteen propositions intact and its worked illustration in Appendix A void, and that is the correct reading rather than a hostile one. The location of the evidence is what produces that position. Almost every quantitative parameter the theory uses has been measured somewhere other than the setting the theory is about. Take-over latency comes from driving simulators and one instrumented fleet of consumer vehicles. Supervisory span, interaction time and neglect time come from a laboratory in which participants supervised teams of semi-autonomous ground robots. Exception arrival rates, override rates and the dose–response relation between alert volume and acceptance come from hospital medicationordering and intensive-care monitoring systems. Resumption lag comes from a psychology laboratory, and its field counterpart from an observational study of knowledge workers whose interruptions were human rather than machine-generated. Low-prevalence miss rates come from a visual-search paradigm modelled on baggage screening. Not one of these parameters has been measured in an enterprise deployment of AI agents performing knowledge work. What the paper does with those parameters is therefore a mechanism argument, and mechanism arguments have a known epistemic shape: they give direction without magnitude. That attention is finite, that an interrupted task costs something to resume, that detection probability falls with prevalence, and that a queue whose arrival rate exceeds its service capacity will shed load Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 81 somewhere — these are domain-general propositions, and it is those, rather than any transported constant, that carry Propositions 3 to 7. The constants almost certainly do not transfer; there is no reason to expect the interaction time for adjudicating an agent's proposed contract amendment to resemble the fifteen to seventeen seconds measured for retasking a robot. A reader who accepts every mechanism here and rejects every number has taken a defensible position, and the paper's own design proposals concede as much by instructing firms to measure locally rather than to import. What would settle the question is straightforward to specify and has not been done. The parameters λ, IT, c, A and N of Definitions 3 and 4 have never been estimated for a deployed agent system, and to those Section 11.2 adds a quantity of a different kind that is in no better position: the recall of the rule by which items are routed to an overseer, which is a property of a component the firm builds for itself rather than of the overseer, and which bounds what any triage design can achieve. The research underlying this paper searched specifically for such estimates and located no rigorous measurement, in any deployed enterprise setting, of agent action rates against human review capacity, and no measurement of the marginal human cost per agent action. What exists is vendor telemetry, position papers, and inference from the tool-call counts of benchmark runs. The absence should be stated in both directions. It is a governance finding in its own right: organisations are deploying agents at rates whose oversight cost nobody has quantified. It is also, and more seriously, a limitation on this paper. Definition 4 states an inequality none of whose terms has an estimated value in the domain the paper addresses, and such an inequality cannot tell a firm whether it is solvent; it can only tell the firm what to measure to find out. One of those terms is worse than unmeasured, and listing it alongside the others understates the problem. Neglect time NT is defined in Definition 3 as the interval a supervised process can proceed unattended before its expected performance falls below threshold, and both the interval and the threshold are defined against the quality of the unattended trajectory. Estimating either therefore requires knowing what the system would have done — a signal available at or near the time of action where an outcome is realised quickly and observably, and unavailable in principle where it is not. Proposition 9 marks exactly that boundary: where ground truth is a counterfactual future that may never come to exist, verification cost is undefined at decision time. In those decision classes NT has no operational definition, no default can honestly be adopted for it, and the span bound N ≤ NT/(IT + WT) + 1 cannot be evaluated at all — not because nobody has yet measured it, but because there is nothing there to measure. This is a definitional gap rather than a data gap, and it does not close with more fieldwork. The consequence for Definition 4 should be stated plainly rather than left to the reader: the solvency constraint applies in full only to verifiable decision classes, and in unverifiable ones it reduces to the arrival-rate condition alone, which remains well defined because it is stated over attention and elapsed time rather than over outcome quality. The concurrency bound simply drops out. Appendix A is qualified accordingly, so that a firm implementing the protocol is not sent to measure a quantity that does not exist for precisely the decisions it most wants to supervise — which is an uncomfortable result, since the classes in which the span bound is inapplicable are the consequential ones the paper is about. How far the indeterminacy extends is now shown rather than asserted. Section A.6 performs a scenario analysis of the constraint across five knowledge-work archetypes and reports that the ratio c/IT, which decides whether context re-acquisition is a rounding error or the dominant cost, has no measured value in any of them; that the plausible range of c spans the three orders of magnitude between the controlled resumption measurements and the field observation of interrupted knowledge work; and that within that range the same configuration is comfortably solvent or Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 82 insolvent by a large multiple. 

The result of the exercise is an ordering rather than a number, and Section A.6 argues that an ordering is what the model can honestly deliver to most firms. A second simplification runs alongside the transfer problem and is independent of it, in the sense that it would survive the arrival of perfectly domain-specific measurements. Uncommitted attention, interaction time, context re-acquisition cost and the overseer's detection probability are written throughout this paper as constants over the period against which solvency is evaluated, and Section 4.7 assembles the evidence already in the paper that all four drift within a single period of duty: detection falls with time on watch, acceptance falls with the number of interruptions already absorbed rather than with general workload, resumption cost rises with the demand of what the overseer was pulled into, and response latency lengthens with cumulative exposure in the one deployed setting where it has been observed. The four measurements are not commensurable and none was collected to test the constancy of these parameters, so no functional form for the drift is proposed here and none could honestly be fitted from them; a dynamic constraint, in which the parameters are written as functions of elapsed time and cumulative interruption and solvency is required to hold at the end of the period rather than on average across it, is the obvious next formalisation and is left undone. What the four do establish jointly is the sign. Every drift runs in the direction that tightens the constraint and none offsets another, so the static condition is an optimistic bound: a configuration evaluated on early-period parameters may satisfy it and be insolvent by the end of a shift, with the arithmetic returning the same verdict on inputs that no longer describe the overseer. The direction of the resulting error in this paper's own estimates is therefore known even though its magnitude is not, which is the most that can be claimed for a static treatment and is why Section 4.7 asks a firm reporting its utilisation ratio to report it early and late in the period rather than once. The cost of the scope condition stated at the opening of this subsection should therefore be recorded with the same directness, because it is substantial and it is not repaired by anything in the paper. A theory that predicts only signs cannot tell a firm whether it is currently insolvent. It can tell the firm only which way it will move when it changes something — that raising the routing threshold reduces ρ, that removing work raises A, that batching switches reduces the term c contributes, that improving the model lowers prevalence and by Proposition 7 therefore lowers detection — and an instrument that ranks directions without locating a boundary is weaker than one that does both. The declaration protocol of Appendix B is the means by which the theory would acquire magnitudes, and it is the only means proposed here. Until firms operate it, this theory is untested in the one way that matters: not unargued, and not unsupported by mechanism, but never once confronted with a measured value of the quantity it claims a firm is exceeding. 11.2 The asymptotic claim is derived; the functions it is stated over are unmeasured Proposition 12 asserts that the attention cost of constitutive oversight is invariant in throughput while that of control oversight is Ω(T). The second half must be read with the qualifier the proposition attaches to it, because without that qualifier it is a different and weaker claim. The Ω(T) attaches to the cost of holding a fixed detection guarantee, and not to the exception arrival rate, which need not grow in throughput at all and may fall as the per-action error rate falls. What grows at least linearly, and strictly faster wherever reliability is improving, is the attention required to hold the share of consequential errors removed at a positive constant. The two halves of the claim do not have the same evidentiary standing. The first half is close to definitional. Constitutive oversight is defined in Definition 6 as the class of decisions taken ex ante and revised on a governance cycle, so its independence from throughput follows from the definition rather than from any observation. That is not a defect, but it means the Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 83 O(1) claim earns no evidentiary credit. The interesting question is whether any consequential oversight function can be placed in that class, which is an empirical question the paper answers only by example. The second half is not definitional either, but neither is it any longer an assertion. The asymptotics are derived in Section 8.3.1 from the definitions themselves: the number of items a fixed attention budget permits to be examined is independent of throughput, so coverage falls as O(1/T); efficacy is that coverage multiplied by the detection probability at the prevailing prevalence, so it falls at least as fast and strictly faster wherever the error rate is decreasing; and inverting the expression prices the guarantee at Ω(T), superlinearly under Proposition 7. The amended proposition carries that derivation one step further, and the step is a generalisation rather than a repair: the uniform review over which the asymptotics were originally stated is the special case of a routing rule that discriminates nothing, and where the firm routes selectively the share of consequential errors removed is the product of the recall of the routing rule, the share of the routed stream the attention budget permits to be examined, and the detection probability at the prevalence the routing induces. Selective routing raises that prevalence and with it detection, but cannot lift efficacy above the recall of the rule, so the ceiling moves from the overseer's attention to the router's recall rather than disappearing. The queueing term WT and the per-interruption re-acquisition cost c of Section 4 aggravate the result rather than carrying it, and the field evidence of exception-channel saturation in Section 4.4 — override rates of 49 to 96 per cent, drug-alert acceptance below 1 per cent — is consistent with a channel whose demand outran its capacity long ago. The concession that must be recorded here is therefore not an unfitted exponent but an unmeasured set of functions. The derivation is stated over four quantities and none has been estimated for a deployed agent system: the number of items an attention budget permits to be examined per period, n = A/(IT + c); the detection function d(π) and in particular its slope in prevalence outside the visual-search paradigm; the per-action rate of consequential error p(T) and whether it falls with maturity in any deployed decision class; and, since the amendment described above, the recall R of the rule that routes items to the overseer. The fourth is the least measured of the four and is the one whose absence bites hardest, because it is the ceiling the amended proposition places on every triage design: no value of R has been measured for any deployed enterprise triage layer located in this paper's research pass, and the reason it is not incidentally observed is structural rather than a matter of effort, since everything a firm sees of its routing rule is conditioned on having been routed and recall is a statement about what was not. A derivation is only as informative as the functions it quantifies over, and these four are the paper's principal gap, as Section 11.1 states in the general case. Figures 4 and 5 are accordingly illustrative of shape and not of magnitude: no value in either is fitted, the vertical scales are relative, and T* is asserted to exist rather than located. The study that would test it can be specified exactly. Take a single decision class within one firm, instrument the routing system to record every exception with timestamps for surfacing, first touch and closure, and vary the throughput of the underlying system across at least an order of magnitude, whether by staged rollout or by exploiting natural variation in volume. Then estimate, as functions of throughput, both the attention consumed per period and the share of consequential errors the review removes, the second requiring an audited error rate within the examined stream and the reviewers' measured detection rate at that prevalence, while independently recording the override rate and an audited measure of per-exception handling quality. Where the review is not uniform, the design acquires one further requirement, and it is the expensive one: the share removed cannot be recovered from the routed stream alone, so the study must also sample what the routing rule declined to route, in order to estimate its recall. If the per-period attention cost remains Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 84 bounded as throughput grows while the share of consequential errors removed remains bounded below — whether the firm holds that share up through coverage, through the recall of its routing, or through any combination of the two — and the control function has not been transferred to a verifier correlated with the generator, Proposition 12 is refuted on its own stated condition. If attention grows at least linearly at a held share, or the share falls as throughput rises, the proposition acquires for the first time an estimated magnitude rather than a shape. 11.3 The evidentiary quality of the foundations, honestly graded A second limitation concerns not the transfer of the evidence but its grade at source. Several works this paper leans on hardest in Sections 2 and 5 are not what their citation counts imply, and collected in one place the pattern is worth confronting. Bainbridge (1983), the most cited work in the supervisory-control tradition and the origin of the framing that organises Section 2, is a five-page Brief Paper in Automatica: an analytical essay with no sample, no design, no data and no statistics. Sheridan and Verplank (1978), which supplies the levels-of-automation scale from which the technical content of human-on-the-loop is drawn, is a laboratory technical report distributed through the defence documentation system: grey literature, not peer-reviewed, and cited here through the route stated in Section 2.1. Warm, Parasuraman and Matthews (2008) and Parasuraman and Manzey (2010), which supply the vigilance-is-work finding and the conditions under which complacency appears, are reviews rather than experiments and report no sample of their own. Parasuraman, Bahri, Molloy and Singh (1992), the source of the reliability-constancy result on which Proposition 6 principally rests, is a government technical report and is not peer-reviewed; its reported degrees of freedom are not reconcilable with its reported sample size, a discrepancy the paper disclosed where it used the result. Its peer-reviewed companion, Parasuraman, Molloy and Singh (1993), could not be independently verified, and no figure is quoted from it. Olsen and Goodrich (2003), which supplies the fan-out construct and the claim that context re-acquisition is the largest contributor to the supervisory ceiling, is a workshop paper, is not peer-reviewed, and contains no measured values. The conclusion is narrower than a dismissal and sharper than a caveat. The genuinely empirical, peer-reviewed, quantitative core of the supervisory-control literature is considerably smaller than its citation counts suggest, and a paper presenting the ironies of automation as an empirical finding would be overstating its evidence base. The material carrying the most weight here is drawn from elsewhere: the visual-search psychophysics of Wolfe and colleagues, the robotics span measurements of Crandall and colleagues, the clinical override datasets, the two meta-analyses of human–AI combination, and the two conference results on self-verification. A reader auditing this paper should audit those first. 11.4 What could not be verified, and what remains to be done The following items were used, qualified or declined on the basis of incomplete verification, and each is recorded so that a reader can locate the paper's soft joints without re-deriving them. The texts of Articles 14 and 26 of the EU AI Act were read from a mirror rather than the Official Journal PDF, and should be checked against it before submission. The postponement dates attributed to the Digital Omnibus on AI rest on a single reading of the amending text; the instrument's existence, title, adoption and entry-into-force dates were confirmed directly, but the dates on which the high-risk obligations bind were not independently corroborated. No slip opinion was obtained for any of the four Delaware decisions, so nothing is quoted from them anywhere in this paper and each is cited for the proposition alone, in this paper's own words; the received formulations reached the paper through a partial Justia retrieval and a law-firm case summary. The consequence for the argument Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 85 is confined and is stated where it arises: the further step taken in Section 10.5, that claiming oversight one does not perform is itself part of what the violation consists in, rests on the misrepresentation element of the Boeing record as reported at that secondary level, and is advanced with that dependence attached rather than as an independently verified holding. No Delaware decision applying Caremark to AI oversight was verified to exist as at August 2026, and the paper says so rather than implying an extension that has not occurred. The Japanese Companies Act article numbering and the ใ€Œ้€šๅธธๆƒณๅฎšใ•ใ‚Œใ‚‹ใ€ formulation could not be retrieved from primary sources, and both require confirmation; the comparative argument of Section 10.6 depends on neither, since it requires only that the Japanese standard key on foreseeability, which is a property of the structure of that standard rather than of its article numbering or of any particular rendering of the received formulation. The ๅคงๅ’Œ้Š€่กŒ case was not verified and is not relied upon. The normative text of ISO/ IEC 42001:2023 is paywalled and unread, so no control language is quoted and no claim is made about its Statement-of-Applicability structure; the standard is accordingly excluded from the structural finding of Section 10.4 rather than counted there as an instrument found silent on quantity, that finding being stated over the instruments whose text this paper's research pass actually inspected. Several empirical items were declined for the same reason. Cummings and Mitchell (2008) is frequently cited for an operator-utilisation ceiling; no quantitative result could be obtained from it, so the record is cited and no figure appears. The long-tail take-over range attributed to Eriksson and Stanton (2017) could not be confirmed and is not used. Elish's (2019) case studies were confirmed from the journal record rather than the article body, so the paper refers to them in general terms and quotes nothing from her treatment of them. The reasoning of the tribunal in Moffatt v. Air Canada was reached through secondary analyses because the primary database blocks automated retrieval, and is described rather than quoted. The characterisation of Robodebt frequently quoted in press coverage was not verified against the report and does not appear here. Two further items are misattributions the paper deliberately declines to use, and the declining matters more than it might appear. The first is the resumption figure of "23 minutes 15 seconds", almost universally attributed in management writing to Mark, Gonzalez and Harris (2005); it does not appear in that paper, whose figure is 25 minutes 26 seconds, itself conditional on same-day resumption and therefore biased downward. The second is the reading of Mark, Gudith and Klocke (2008) as showing that interruptions slow people down. It showed the opposite on time: interrupted work was completed faster, and the cost appeared as stress, workload, frustration and compressed output. That correction is not pedantry, because the misreading and the finding point in opposite managerial directions. If interruption showed up as delay, an insolvent oversight channel would announce itself in the throughput metrics a management system already watches. Because it shows up as strain and compressed quality, it does not. 11.5 Competing explanations the paper has not excluded Four alternative accounts of the same evidence remain open, none disposed of here. The first is selection. Firms that adopt on-the-loop configurations may differ systematically from firms that do not — in the risk profile of their decision classes, in their regulatory exposure, in the maturity of their control functions. If so, the documented failures in Section 9 may not identify oversight design as the cause of anything; they may identify the kind of organisation that arrives at nominal oversight, with the arrangement and the failure both consequences of a prior disposition. Nothing here excludes that possibility, because the cases are selected on the outcome: they are in the record because they failed badly enough to produce a parliamentary inquiry, a royal commission or Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 86 a regulatory order. The comparison class of on-the-loop configurations that did not fail is not observed. The second is immaturity rather than structure. Every capability figure in Section 1.4 is a moving target, and one preprint estimate places the doubling of agent time horizons on the order of three months. The constraints this paper describes may be the constraints of a transitional period — exception rates high because the systems are immature, action counts high because the scaffolding is crude, and a sufficiently capable agent generating so few exceptions that the solvency constraint never binds. The paper's answer is narrower than it once was, and the amended Proposition 11 is the reason. The proposition is not constructed to survive arbitrary capability improvement unconditionally. The marginal value of a verifier is the product of its decorrelation from the generator and its detection probability on the error class, and capability improvement moves those two terms in opposite directions. Decorrelation is a joint property of the generator–verifier pair rather than a capability of either, and no improvement on either side supplies or removes it; but the falling exception prevalence that improvement produces lowers a human verifier's detection probability through Proposition 7. The human's oversight value is therefore neither eliminated by capability growth nor preserved by it. The proposition survives capability improvement only where the firm operates an instrument that holds detection probability up as prevalence falls — the argument is made in full at Section 7.5, and the instrument is the fourth design proposal of Section 10.7, the injection of synthetic exceptions at elevated prevalence with immediate ground-truth feedback. Against the immaturity objection this is a conditional answer rather than a decisive one, and the condition is one almost no deployment presently satisfies: a firm whose exception rate is falling because its models are improving, operating no prevalence-restoring instrument, is exactly the case the objection describes, and what the paper offers against it is a design obligation rather than a refutation. The answer is honest but not conclusive, and a reader should weigh whether it is sufficient. It defends the independence claim conditionally; it does not defend the capacity claims of Sections 4 and 5, which would weaken if exception prevalence fell far enough — although Section 5 gives the reason to expect that falling prevalence degrades rather than rescues detection. The third is the reverse of the paper's own claim. Dessein (2002) implies that veto-based oversight is usually value-destroying against a better-informed agent unless the incentive conflict is extreme. Read through that result, the low measured effectiveness of human oversight documented in Section 6 is not evidence that oversight should be repaired but evidence that firms should delegate more rather than supervise better. The paper's constructive proposal is compatible with that reading: moving decisions into the constitutive class and out of the control class is delegation plus objective alignment, which is what Dessein recommends. That compatibility is a point in the proposal's favour and simultaneously a limitation on the evidence, because the evidence assembled here does not discriminate between the paper's account and Dessein's. Both predict that control oversight underperforms; they disagree about why, and no measurement offered here separates them. The fourth is that verification independence may be manufacturable by technical means. Definition 5 is a statistical condition, not a statement about species. An adversarially trained verifier, a formal method with a soundness guarantee, or an ensemble of genuinely heterogeneous models — different architectures, corpora, vendors and input channels — might satisfy it. The paper's own evidence supports the possibility: Stechly and colleagues (2025) show that substituting a sound external verifier for self-critique moves Game of 24 from 3 to 38 per cent, Graph Colouring from 2 to 37 per cent and Blocksworld from 55 to 87 per cent, and that verifier is not human. If such a verifier can be shown decorrelated with its generator on the error classes that matter, Proposition 10 continues to hold and the inference that the exogenous verifier must be human does not follow. This Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 87 is the point on which the paper is most likely to be misread as a defence of human beings. The claim is about independence, not about species. A demonstrably decorrelated automated verifier would satisfy Definition 5, and its construction is a research programme rather than an objection — one whose progress would narrow, though not close, the residual this paper calls non-delegable. 11.6 What the theory explains, and the awkward fact about it The theory's central claim is that firms cannot state the parameters of their own oversight. It therefore explains the world largely through an absence. No firm known to this paper publishes an oversight solvency statement, no instrument surveyed in Section 10 requires one, and the observation that the quantity is missing is the observation the theory rests on. That is an uncomfortable evidentiary position, and it should be named rather than finessed. Because no such statement exists, the theory has not been tested against a case where one does. It has not been shown that a firm which measures λ, IT, c, A and N and publishes them thereby behaves differently, discovers insolvency it did not suspect, or improves any outcome. The theory predicts all three; it has demonstrated none. This mirrors an acknowledged position in the series' capital-allocation paper, where a construct's principal evidence was likewise the non-existence of the measurement it proposed, and the same discipline applies here. A theory whose principal evidence is the absence of a measurement is weaker than one whose evidence is a measurement, and it is weaker in a way that argument cannot repair. It is repaired by someone taking the measurement. The paper should be read accordingly, and Appendix A exists to make taking it cheap enough that it might be. 11.7 Conflict of interest and self-citation disclosure The five-role formulation examined by this paper was advanced by the author's own firm, VURA Capital Innovation Holdings, Inc., in a press release dated 29 June 2026. The firm offers advisory services in the design of the arrangements the paper analyses. The interest is disclosed in Section 1.1 and is restated here in the form required of an academic disclosure. The paper's conclusions do not vindicate the formulation. Two of the five roles are shown to carry attention costs the model does not price: governance and exception intervention are control functions in the sense of Definition 6, and their cost rises with the throughput the deployment exists to increase. The closed-loop claim is not supported by the peer-reviewed literature relied on in Section 7, in which self-correction without external feedback left accuracy flat or lower in every reported cell, substituting self-critique for a sound external verifier degraded performance in two of three planning domains, and a supervisory agent built on the same base model as the agent it supervised approved at approximately eight times the rate at which it denied. The series' self-citation discipline is applied explicitly. Each citation of the author's other working papers is classified as an axiomatic foundation, as observational evidence, or as a design proposal. Future Value Theory (Kadowaki, 2026a) and Enterprise Redefinition (Kadowaki, 2026b) are axiomatic foundations, adopted as premises rather than as findings established. Enterprise Redefinition Observed (Kadowaki, 2026c) is observational evidence, and is the only self-citation carrying evidentiary weight. From Job Description to Purpose Description (Kadowaki, 2026d) and Brain Capital Management (Kadowaki, 2026e), supplying respectively the role-design apparatus used in Section 8 and the capacity construct that Section 4 formalises together with the disclosure architecture that Section 10 and Appendix B adapt, are design proposals. Redefinition of Capitalism (Kadowaki, 2026g) is cited in Section 1.2 as methodological precedent only. Design proposals are not counted as support for any claim anywhere in this paper. The propositions that inherit a designproposal premise are named: Propositions 3, 4 and 5 rest on a capacity construct first proposed in Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 88 Kadowaki (2026e), and Proposition 12's reconstruction of the five roles uses the role-design apparatus of Kadowaki (2026d). Each is warranted independently by the evidence cited in the section that states it, and none rests on the self-citation alone. A working paper is unrefereed at the time of circulation, and this one is published by an organisation with a commercial interest in its subject matter; the evidence-provenance asymmetry set out in Section 1.2 applies to it on the same terms as to any other unrefereed source. Readers should weigh the disclosed interest when assessing the paper's framing, its selection of evidence and its choice of emphasis. 11.8 The proposal against its own critique Section 4.6 used Bevan and Hood (2006) and Strathern (1997) to close off the intensification route, on the ground that a measurement system distorts what it measures and itself goes unaudited. The paper then proposed, in Section 10.7 and Appendix B, a published metrics regime — λ, IT, c, A, N, utilisation, override rate, override trend, path dependence, injection rate, catch rate and the routing-recall estimate R — and never turned the same instrument on it. Bevan and Hood's three effects apply to it directly. Threshold effects appear the moment a solvency figure is published: a firm reports ρ just below unity and has no incentive to go further, so a diagnostic becomes a target that performance bunches beneath. Ratchet effects appear as soon as periods are compared: a unit disclosing attention headroom invites a larger exception allocation next quarter, so the rational response is to consume it first. Output distortion is the most serious, since every quantity here is a proxy for a quality that is not itself measured, and that quality is what the oversight exists to produce. The specific gaming moves should be named, since an unnamed vector is one a firm can adopt without noticing. The arrival rate λ is lowered by re-routing exceptions away from the named overseer, or more cheaply by redefining what counts as an exception, the routing threshold being the firm's own to set. Interaction time IT is lowered by measuring to first touch rather than to closure, a substitution that can halve the reported figure without touching the work. Uncommitted attention A is inflated by reporting nominal availability rather than the subtractive residual Appendix A specifies, the overstatement Proposition 4 identifies. The solvency statement as a whole is satisfied by splitting a decision class into several, since every term is defined per class, and a class divided is a load divided on paper while the overseer's day is unchanged. A published override rate becomes a target whose cheapest satisfaction is not better adjudication but routing fewer and easier exceptions, which improves the number and degrades the control. And the catch rate on injected synthetic exceptions — the item on which Section 7.5's argument depends — is inflated simply by making the injections detectable, the least effortful route to a good figure and one that destroys the entire evidentiary value of the exercise. The Bevan and Hood analysis applies with particular force to the last of the quantities added to the regime, the routing-recall estimate R, because R is the one figure a firm can improve without altering anything its router does. Recall is a ratio whose denominator is the consequential errors committed in the class, so narrowing what the firm counts as routed — or, equivalently, narrowing the class of errors against which recall is scored — raises the reported figure while the rule selects exactly what it selected before. That is the same classboundary move as the subdivision of a decision class, which is precisely the move the audit trail below is designed to catch, and it is a reminder that the trail's stable published identifier does more work than any other requirement in the protocol. Two of these are auditable from outside the firm on the published figures alone, and the rest are not. The injection catch rate is auditable if, and only if, the injections are drawn by a third party from a distribution the firm does not control, which converts a self-report into an outside Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 89 measurement. The decision-class boundary is auditable if the classification is published and held stable, since a class that subdivides between periods is visible in the taxonomy. The remainder — what counts as an exception, whether IT runs to first touch or to closure, what is subtracted to produce A — are internal accounting choices no external reader can verify from the numbers themselves, and the only control available at that level is that the definitions be published with the numbers and not moved. Part of that residue is now answered rather than merely conceded, and the answer should be credited before what remains of it is stated. Appendix B requires, as a condition of any Tier 1 declaration, a minimum audit trail: a stable published identifier for each decision class, with any change between periods disclosed as a restatement setting the prior boundaries beside the revised ones; and item-level timestamps for the routing event, for first touch and for closure, retained rather than aggregated. Those two requirements convert two of the vectors named above. Classsplitting ceases to be an invisible move, because a load divided on paper must be declared as a restatement to be divided at all. And the endpoint of the interaction-time measurement ceases to be an internal convention, because the substitution that halves the reported figure by stopping the clock at first touch is contradicted by a record the firm is required to keep. Both are converted by requiring records rather than by requiring candour, which is what makes them enforceable against a firm disinclined to either. The residue that survives is the part that matters most, and no extension of the trail reaches it. A log establishes that the reported numbers describe what the firm did; it does not establish that the decision classes were drawn in good faith in the first place. A firm may keep an impeccable record of a taxonomy designed from the outset to distribute an unmanageable load across classes that each look solvent, and the trail will show every timestamp in order. What counts as an exception remains the firm's to define, and the definition is prior to every quantity the protocol collects. That is a limit on the instrument rather than a defect of the trail, and it is the reason the paper does not claim that a completed declaration establishes an adequate control environment, but only that an incomplete or unstable one is visible as such. The evidence most damaging to the paper's own instrument has so far been omitted from it entirely. Becker, Rush, Barnes and Rein (2025) ran a randomised controlled trial of sixteen experienced opensource developers across 246 tasks on mature repositories on which they averaged five years of work, randomising each task to AI-allowed or AI-disallowed. The study is a preprint, not peerreviewed. Its measured effect was a 19 per cent increase in completion time when the tools were allowed, against the same developers' post-hoc estimate that the tools had made them 20 per cent faster. METR published a partial walk-back on 24 February 2026, stating that the original estimate no longer represents current conditions and reporting new preliminary figures whose confidence intervals cross zero and which METR itself calls only very weak evidence. The sign of the effect is therefore not established; what survives is the belief gap, practitioners being wrong about the direction of an effect on their own work, by some forty percentage points, while living through it. A paper whose entire prescription is voluntary self-measurement and self-publication must say that self-report in this domain has been measured and found invalid, and must draw the consequence rather than note the irony. The consequence is a rule of acceptance for Appendix B: every quantity must be derived from a record system maintained for another purpose — routing logs, workflow timestamps, assignment records, the injection harness — never from a practitioner's estimate or a manager's judgement. The METR result does not merely warn against self-report at the margin; it establishes that estimate and measurement can differ in sign. This is the same acceptance rule the series' fifth working paper applied to a different disclosure problem, and that parallel is recorded as Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 90 a design proposal of the author's rather than as evidence, in keeping with Section 11.7. It supports no proposition; it records only that the rule was not invented for this occasion. 11.9 The disclosure paradox and the cost of the protocol The paper's central recommendation is one a firm's counsel would forbid, and that should be stated rather than left for the reader to discover. Section 10.5 concludes from the Boeing findings, in express dependence on the misrepresentation element of that record as reported at secondary level, that claiming oversight one does not perform is itself part of what the violation consists in. Appendix A's worked illustration, on the author's own illustrative parameters, returns a configuration insolvent by a factor of roughly 2.7 at median interaction time and by a factor of eight at the ninetieth percentile. A firm that publishes an honest solvency statement showing insolvency, and continues to operate the decision class unchanged, has not merely disclosed a weakness; it has manufactured, dated and signed the documentary record for a claim against itself, in a form establishing that the board knew. The exposure should be described as narrowly as Section 10.5 describes it, since the narrower version is the one that survives: what the disclosure bears on is the accuracy of what the firm has represented about its own control environment, and not the insolvency itself, which is a condition under Definition 4 and not a breach of any duty. Under Section 10.5's own logic, publishing honestly and not halting is worse than publishing nothing, and counsel advising otherwise would be advising badly. This is a strong and probably decisive reason why the disclosure will not be adopted voluntarily, and the paper's proposal that firms adopt it voluntarily is weaker for it. What follows is an argument for making the disclosure a regulatory instrument rather than a voluntary one. Where a quantity is mandated of every firm in a class, publishing it is compliance rather than confession, and the asymmetry penalising the honest discloser disappears. Section 10.4 establishes that the slot is empty: across the instruments whose text this paper's research pass inspected, exactly one states a measurable human-oversight quantity, and that one is a headcount. Every other specifies process and documentation, so a mandated oversight quantity would displace nothing already in force. The regulatory slot is both empty and available, and this paper's argument is better read as identifying what should occupy it than as advice to a firm acting alone. The honest counter-consideration is the one Section 11.8 has just made: a mandated disclosure of an inequality a firm cannot satisfy would in practice be satisfied by redefining its terms — the exception threshold moved, the class subdivided, the interaction time measured to first touch — so mandation converts a problem of incentive into one of definitional discipline rather than dissolving it. A regime mandating the quantities without freezing their definitions would produce compliance and no information. The second objection the paper has not answered is the practitioner's: that the protocol is unimplementable at reasonable cost. Not one of Appendix B's eleven items is priced anywhere in this paper, and that omission should be conceded before it is addressed. An order-of-magnitude statement is nonetheless possible, and it does not flatter the proposal. It is the reason the protocol is now presented in two tiers. The eight items of Tier 1 are declarative: they require a firm to state a classification, name a holder of authority, record a register entry, describe a reporting route, disclose a reliance, or read quantities off record systems it already maintains. Their cost is a policy exercise conducted once and revisited on a governance cycle, and it is small. The three items of Tier 2 are continuous and expensive in the one resource the deployment exists to release. Syntheticexception injection with ground-truth feedback consumes the overseer's adjudication time on items of no business value, and grader time to establish the ground truth on each; scheduled returns to unassisted operation consume production throughput directly, executing a decision class without Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 91 the assistance installed to accelerate it; and the residual uniform sample required of a firm operating a triage layer consumes review capacity on the stream the routing rule declined to route, which is the very work the triage was installed to remove. The third is the newest and is the one whose cost is most exactly the benefit claimed: a firm reporting a workload reduction from its routing must give part of that reduction back in order to establish what the routing costs it in errors never presented. All three fall on the operation the automation exists to speed up, and all three recur every period. The predictable adoption pattern follows, and the paper concedes it rather than arguing against it. A firm minimising cost adopts Tier 1 and declines Tier 2, and its published protocol will be substantially complete and evidentially empty — because the three items declined are the three carrying the actual evidentiary warrant, the injection regime resting on the only intervention shown to move the low-prevalence miss rate, the unassisted schedule on the only one shown to restore failure detection, and the residual sample being the only route by which the recall that bounds a triage layer's efficacy can be estimated at all. Tier 1 describes a configuration; Tier 2 tests whether it works. Appendix B nonetheless advises adopting Tier 1 first, on the ground that a declared loop position and a published independence register are what make an absent Tier 2 visible from outside the firm. The protocol has never been implemented in full by any firm known to this paper, and is therefore untested as an instrument rather than merely unvalidated in its effects. It also yields something more useful than a caveat. The adoption pattern is itself a testable prediction of the paper's own theory: if firms adopt the protocol, the theory predicts that Tier 1 will be adopted at a materially higher rate than Tier 2, and that within Tier 2 the residual sample will be declined most often by the firms claiming the largest reduction in review workload. Both predictions are refutable by counting. Table 9 closes the section. It is the map of verification for the paper as a whole: it records, for each of the thirteen propositions, the unit at which it is stated, how it was derived, and in one clause the condition that would refute it. Where a proposition inherits a premise from a self-cited design proposal, the derivation cell says so, and Section 11.7 gives the independent warrant relied on instead. Proposition Unit of analysis Derivation Refutation condition (abbreviated) P1 Position determination Decision class Definitional; conjectural on scaling Approvals remain determinative where consequence latency has fallen below measured human response latency. P2 Rubberstamping Decision class Imported from formal theory Override rate and override accuracy are independent of the human's information disadvantage. P3 Oversight solvency Oversight configuration Imported from measured evidence (inherits designproposal premise) A persistently insolvent exception channel shows no decline in perexception handling quality. P4 Substitution, not addition Overseer Imported from measured evidence on attention (inherits design-proposal premise) A duty added to an unchanged workload matches a dedicated overseer's detection performance. Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 92 Proposition Unit of analysis Derivation Refutation condition (abbreviated) P5 Span Overseer Imported from measured evidence; the exogenous-arrival bound derived from Definition 4 (inherits design-proposal premise) Where the exception stream is endogenous to the supervised processes, measured knowledgework span materially exceeds the robotics range without greater neglect time. P6 Reliability– supervisability tension Oversight configuration Imported from measured evidence Failure-detection performance is invariant to reliability constancy under multiple-task load. P7 Rarity Overseer Imported from measured evidence An intervention that does not alter observed prevalence eliminates the low-prevalence deficit. P8 Error decorrelation Oversight configuration Conjectural extension with a measured component Overseers demonstrably correlated with the system on the relevant error class nonetheless exceed the accuracy of a system whose baseline advantage over them exceeds the reported sampling error. P9 Verifiability precedence Decision class Conjectural extension Decisions with counterfactual ground truth admit contemporaneous verification that improves accuracy. 

P10 No dependent verifier Control loop Imported from measured evidence Iterated verification by a verifier demonstrably correlated with the generator on the relevant error class improves accuracy on a task class with a non-trivial error rate. P11 Independence is necessary and not sufficient Overseer Conjectural extension with a measured component Verifier value increases with agreement holding accuracy constant; or detection probability on rare error classes is invariant to prevalence under a regime carrying no prevalence-restoring instrument. P12 Scaling dichotomy Firm Asymptotics derived from Definitions 3, 4 and 6, and the bound on non-uniform routing from the same definitions with Definition 5; conjectural only in the premise imported from Proposition 7, that detection probability increases in prevalence in knowledge work (inherits design-proposal premise) A control mechanism has per-period attention cost bounded in T and the share of consequential errors it removes bounded below — whether through coverage, through routing recall, or through any combination of the two — without transferring control to a verifier correlated with the generator. P13 Tracing without tracking Firm Conjectural; historical Formal accountability on an overseer lacking channel, latency or authority is followed by measured error-rate improvement. Table 9. The thirteen propositions: analytical unit, derivation, and refutation. Human-on-the-Loop · WP No. 8 11. Limitations, Falsification, Competing Explanations, and Interests 93 12. Conclusion This paper began with a proposition offered as a management model: artificial intelligence executes the work, and the human defines purpose, values and capital allocation, exercises governance, and intervenes at exceptions. The proposition was taken seriously enough to be tested rather than endorsed or dismissed, and what survives is narrower, more demanding and more useful than what was proposed. The argument runs as a chain, and the chain is worth restating in its final form, because each link constrains the next. The first link is that loop position is measured, not chosen. A configuration is on the loop only where the human's end-to-end response latency is smaller than the interval between the system's commitment to an action and the realisation of that action's consequence. Both quantities are properties of a deployed system; neither is a property of a policy document. A firm therefore does not decide to be on the loop, and the decision it believes it is taking when it adopts an oversight posture has already been taken for it by the throughput of the system it deployed. Raising throughput converts an in-the-loop arrangement into an on-the-loop one, and then into an out-ofthe-loop one, without any policy changing and without the conversion appearing anywhere in the reporting, because approval rates do not fall as this happens. They hold, or rise, and are read as evidence that the system performs well. The second link is that oversight is a consumable resource rather than a role. It has a capacity — uncommitted attention, interaction time, context re-acquisition cost, neglect time, queueing delay — and a capacity can be exhausted. An exception channel whose arrival-rate-weighted cost exceeds the attention available to it is insolvent, and the critical property of insolvency is the form its failure takes. It does not present as refusal: nothing is rejected, no queue builds, no deadline is missed. It presents as acceptance — override, dismissal and non-response rise until the constraint is satisfied by driving interaction time toward zero. This is why the arrangement is stable, why it is invisible from a dashboard, and why the observable symptom of failure — a high and rising override rate — is routinely reported as a sign that the system works. The third link is the one that makes the problem dynamic rather than static. The two properties that make an automated decision class worth deploying, its reliability and its throughput, are the two properties that most reliably destroy the supervisor's ability to detect its failures. Constancy of reliability lowers detection of the failures that do occur. Falling exception prevalence raises the miss rate through a shift in the observer's decision criterion rather than any loss of perceptual sensitivity, a shift that resists both payment and instruction. Rising throughput lowers the attention available per exception. Three mechanisms, from three unrelated literatures, pushing the same way as the system improves. The blades close by default, though not by necessity: event rate, feedback provision, scheduled unassisted operation and workload allocation are organisational variables, and the design proposals exist because they are. The fourth link is that oversight nonetheless works, under conditions that are known, narrow and rarely present where it is mandated: where the overseer's errors are decorrelated from the system's on the relevant error class, and where verifying an output costs less than producing it. An independent information channel — evidence the system does not receive — is the strongest and most auditable source of that decorrelation rather than the condition itself, and it is what the paper recommends because decorrelation is otherwise unobservable to a firm lacking labelled ground truth. Where neither holds, the human modulates the level of reliance without improving its discrimination — which is a change in the confidence attached to a decision and not a change in its quality. The uncomfortable observation is not that oversight never helps, but that the conditions under which it helps and the conditions under which it is required are close to disjoint, since Human-on-the-Loop · WP No. 8 12. Conclusion 94 oversight is imposed precisely where a model has been deployed because it outperforms the incumbent process, and on decisions whose ground truth is a counterfactual future. The fifth link carries the paper's deepest claim, and it is deliberately not the claim usually made on the human's behalf. A control loop cannot be closed by a verifier statistically dependent on its generator. Where generator and verifier share a base model, a training corpus, an input channel, a vendor or an incentive, the loop's error rate is bounded below by their shared error rate, and iteration can move the output away from correctness rather than toward it. The human's contribution is therefore independence, not judgement. That matters because it is the only justification for human oversight in this paper that survives improvements in the supervised system — but it survives them only where an instrument maintains the overseer's detection probability as prevalence falls. The overseer's value is the product of two terms. Independence is a joint property of a pair, and no amount of capability on either side supplies or removes it; detection probability enjoys no such protection, and a better model is a lower exception prevalence, which lowers detection. The value is therefore neither eliminated nor preserved by capability. It has to be maintained. The corollary is unwelcome. A firm that succeeds in getting its overseers to agree with the model more often has degraded its control environment, and its agreement statistics will report that degradation as an improvement. The chain terminates in a dichotomy. Because the attention cost of fixing a system's objective function is independent of how often the system acts, while the attention cost of monitoring executions and intercepting exceptions at a fixed standard of detection rises with it, the five roles do not stand or fall together. Purpose, values and capital allocation are constitutive: exercised ex ante, revised on a governance cycle, and unchanged in cost whether the policy they fix is applied ten thousand times or ten million. That part of the proposition survives intact, for a reason the proposition itself does not state. Governance and exception intervention are control: exercised contemporaneously, rising in cost with the variable the deployment exists to increase if their standard of detection is to be held, and therefore reaching a throughput beyond which they are insolvent. Holding their cost flat by reviewing a constant sample does not escape this. It converts an unaffordable cost into an undeclared efficacy: what falls in place of the budget is the share of the system's consequential errors that the oversight actually removes, and a firm improving its system faster than it enlarges its review is losing that share faster still. They are not thereby deleted from the model. They are shown to be unpriced, and what is asked of a firm claiming them is a solvency statement it does not currently produce. The reconstruction is a design instruction, not a prohibition. As much of the two control roles as design permits should migrate into the constitutive class: a rule fixed in advance rather than a case reviewed as it arises; a constraint compiled into the system's action space rather than a violation caught afterwards; an allocation of which cases receive human attention rather than a review of every case. The first two convert a cost that scales into one that does not; the third changes which part of the caseload the surviving attention is spent on rather than exempting that attention from scaling, and is the only one of the three that raises how often error is present in what the human is shown, and with it the chance of the error being caught, as Section 8.3.1 sets out. What that third migration buys is nonetheless capped, and the cap is the share of the firm's errors that the rule actually places in front of a human: an error the rule never presents has not been prevented by it, and no diligence on the part of the reviewer recovers what was never shown. The migration moves the exposure rather than eliminating it, and moves it out of a quantity the firm measures, the time its overseers have, and into one it usually does not, the reach of its own routing. A firm claiming the benefit therefore owes that second figure alongside the saving it reports, and owes it from the stream the rule declined to route, since the cases the rule selected can say nothing about the cases it Human-on-the-Loop · WP No. 8 12. Conclusion 95 passed over. Each is a redesign of routing rather than a demand for a better model or a more diligent reviewer. The question this paper puts to a board follows directly, and it is not the question boards are currently asked. It is not whether the firm is on the loop. The firm is on the loop, or past it, whether or not it decided to be, the decision having been taken by its throughput rather than its governance committee. The question is which decisions the firm has succeeded in moving into the class whose oversight cost does not grow — which constraints are compiled rather than checked, which routing rules fixed rather than applied case by case, which limits set rather than monitored. And for everything that remains in the control class, which will always be something, the question is whether the firm can state the numbers that would make its claimed oversight solvent: the exception arrival rate, the interaction time, the context re-acquisition cost, the uncommitted attention actually available after other assigned work, and the override rate that reveals whether the claim is being met. A firm that can answer the first question has done the design work. A firm that can answer the second has priced what it is claiming. A firm that can answer neither is not exercising oversight; it is describing itself as exercising oversight, and the difference is the whole subject of this paper. It is worth naming at the end what the title calls the non-delegable residual, because the argument has arrived at it from three directions and the three grounds of non-delegability are not the same. The first is structural: verification independence. A loop closed by a verifier statistically dependent on its generator does not converge on correctness, and no improvement in the generator supplies the missing decorrelation, because independence is a relation between two error distributions and not a property of either. What cannot be delegated here is not the judgement but the standing outside. The second is logical: constitutive oversight. Purpose, admissible means, and the allocation of capital and attention across purposes fix the objective function by which everything else is evaluated, and there is no further objective by reference to which the choice of objective could itself be delegated; the attempt regresses. The third is doctrinal: fiduciary accountability. The duty to establish and use a system of oversight falls on the board and is not discharged by designating someone to hold it, which is why Section 9 finds that naming an unequipped overseer allocates liability without transferring control, and why the international guidance recorded in Section 2 states flatly that accountability cannot be transferred to machines. What is not in the residual should be said with equal clarity, because this argument is easily read as a stronger claim than it is. Contemporaneous control oversight — watching executions and intercepting exceptions — is not non-delegable in any of those three senses. It is not that a machine may not do it. It is that at throughput a human cannot, and that the firms which say they do have not priced what they are claiming. The residual is smaller than the five roles suppose, and it is the part that survives every improvement the technology is capable of. One thing would advance this field further than any additional argument, including this one: a single deployed enterprise agent system whose oversight parameters are measured and published. Human-on-the-Loop · WP No. 8 12. Conclusion 96 Appendix A. The Oversight Solvency Calculation A.1 The unit of analysis The calculation set out here is a specification a firm could execute within a reporting quarter from data its systems already emit. It is a proposal of the author, not a finding, and the parameters it requires have never been estimated for a deployed agent system, as Section 11.1 records. The unit of analysis is the decision class, not the system, the model, the team or the department. A decision class is a set of decisions sharing an objective, a consequence structure and a routing rule: refund authorisations above a stated value, contract clause amendments, credit adjustments. Every term in the constraint varies by decision class and by no other cut, so an aggregate computed over a system operating five decision classes with five consequence latencies and five overseers corresponds to no configuration that exists. A firm that cannot enumerate its decision classes cannot perform this calculation. A.2 Measuring each term One convention governs everything below and should be fixed before any term is measured, because the calculation is uninterpretable without it. Every quantity is stated per period — a working day is the natural period for most decision classes and is used throughout the illustration — and the two sides of the solvency constraint are both durations per period. The arrival rate λ is a count per period; the interaction time IT and the re-acquisition cost c are durations per exception, so their sum multiplied by λ is a duration per period; and the uncommitted attention A is likewise a duration, the time available for oversight per period, not a headcount, a fraction or a percentage. Stated that way the constraint of Definition 4 divides through into a single dimensionless number, and Section 4 states it in that form. A second convention follows from Section 4.7 and governs every term operationalised below: none of these quantities is a constant of the overseer, since attention falls and effective handling and re-acquisition times rise as the period wears on and as interruptions accumulate, so each measurement is a value at a point within the period rather than a property of it. Every figure produced by this protocol should therefore record when in the period it was taken, and should be taken at least twice, early and late, the difference being the fatigue term Section 4.7 identifies. The exception arrival rate λ is taken from the system's own logs, not from the design document. Two counts must be kept separate: exceptions routed to a human, which consume attention, and exceptions the system resolved without routing. Their ratio is the parameter a firm will change to become solvent, and its trend indicates whether the routing threshold has been quietly loosened. λ should be expressed per named overseer, since a rate aggregated across a pool says nothing about whether any individual in it is solvent. The interaction time IT is the elapsed time from routing to closure — not to acknowledgement, not to first touch; the distinction drawn in Section 3.2 between motor handover and restored supervisory competence applies directly. It should be reported at the median and the ninetieth percentile rather than as a mean, on the justification of the mean–standard-deviation relationship of Section 3.2, where the correlation between the mean and the standard deviation of take-over times was r = 0.82 (Spearman's rank correlation 0.73, n = 397): variance scales with the central tendency, so a configuration designed against a mean is systematically under-designed. The context re-acquisition cost c is the additional time required to restore the working state of the displaced task. Firms will not initially have it, so a stated default is needed: the controlled resumption-lag measurements of Section 3.3 place it between roughly 0.95 s and 1.9 s per switch, Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 97 asymptoting between 13 and 23 s, and a firm with no local measurement should adopt the upper end of that range and say so. The warning matters more than the default. Laboratory values are lower bounds, because laboratory tasks preserve the environmental cues from which a suspended goal is retrieved. Field observation of knowledge work gives same-day resumption at 25 min 26 s, itself biased downward. The laboratory figure reported without that caveat is a floor presented as an estimate. The available attention A is stated as time available for oversight per period, and is measured as the residual of the overseer's assigned workload rather than as their nominal availability. This is the term firms most systematically overstate, and Proposition 4 turns on it: because attention is fixed in the short run, an oversight duty assigned to a role from which no work has been removed reallocates capacity rather than adding it. The test is subtractive — contracted time per period less every other duty carrying a delivery expectation — and the corresponding disclosure is the list of duties removed when the oversight duty was added, which for most deployments will be empty. One consequence of the subtractive test must be stated in advance so that it is not read as a failure of the instrument, and Section 4 now states it: where the subtraction returns a value at or near zero, that is the result of the test and not a defect in it. An overseer retaining a full operational role has, on the plain reading of Definition 3, close to no uncommitted attention, and the correct response is not to find a more generous measure but to recognise that A must be provisioned, by removing work, rather than measured. A firm that cannot name the duties it removed has not measured a small A; it has measured a zero. The neglect time NT is the interval over which the supervised process's expected output quality remains above the firm's stated threshold without attention, and is the only term that is a property of the supervised system rather than of the human. Where it was measured directly it rose from approximately 30 s with two supervised processes to over 120 s with eight. In an agentic deployment its analogue is the number of hard-to-reverse actions taken before an error becomes costly to undo, so a firm that has architected reversibility into its action space has bought neglect time directly. A caution belongs with this term and nowhere else in the protocol, because it determines whether the term can be measured at all: NT and the quality threshold it is defined against both require knowledge of the quality of the unattended trajectory, and therefore a verification signal available at or near the time of action. Where Proposition 9's condition holds — where ground truth is a counterfactual future that may never come to exist — NT has no operational definition, and a firm should record that fact rather than produce a figure. In such decision classes the span bound of Section A.3 cannot be evaluated and should not be reported, and the solvency test reduces to the arrival-rate condition alone, which remains well defined because it is stated over elapsed time rather than over outcome quality. Section 11.1 sets out why this is a definitional limit rather than a data gap. The wait time WT is the queueing delay incurred when several supervised processes demand attention at once. It is zero in the single-process case and grows with the number of processes, which makes the span bound tighten faster than the ratio of neglect time to interaction time suggests. A.3 The two inequalities ρ = λ·(IT + c) / A ≤ 1 N ≤ NT/(IT + WT) + 1 Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 98 The first is Definition 4's solvency condition divided through by A, and is the form Section 4 uses. Both sides of the underlying inequality are durations per period, so their ratio ρ is dimensionless: it is the utilisation of the overseer's oversight capacity, the share of the available oversight time that the arrival-rate-weighted cost of exceptions consumes. A value of ρ = 1 is not a comfortable operating point but the boundary, since it leaves no capacity for variance, and the ninetiethpercentile interaction time exists in this specification precisely because a configuration designed against a median is designed against the wrong statistic. Values above unity are the multiple by which demand exceeds capacity, and are reported as such. The second inequality bounds concurrency and is evaluable only where NT is defined, as Section A.2 records. A configuration failing either is insolvent under Definition 4, and fails by acceptance rather than refusal — which is why the calculation must be published alongside the symptoms below. A.4 The three observable symptoms that must be published alongside A firm could satisfy both inequalities on paper and still operate an insolvent channel, since the terms most easily overstated are the ones favourable to it. Three observable statistics should be published alongside them. The first is the override or dismissal rate: the proportion of routed exceptions closed without the recommended action being altered. Its anchor is the clinical record of Section 4.4, where override rates ran from 49 to 96 per cent across seventeen studies in 2006 and 46.2 to 96.2 per cent across twenty-three articles in 2020. A high rate is not by itself proof of failure — in a review of 600 overrides, 53 per cent were judged appropriate — but a rate that is not published is a solvency claim that cannot be tested. The second is the trend in that rate. A level speaks to the calibration of the channel; a rise says the constraint is binding and is being satisfied by driving interaction time toward zero. The dose– response relation is measurable: acceptance fell by 30 per cent for each additional reminder per encounter (IRR = 0.70, p < .001), while general workload showed no significant association with acceptance. That null is the operationally important half: it is alert volume per encounter, not busyness, that drives desensitisation, so the remedy is a lower arrival rate rather than a lighter workload — λ, not A. The third is the path dependence statistic: the probability that an alert type is overridden given that its first instance was overridden. In the same cohort this was 87.9 per cent for best-practice advisories and 99.9 per cent for drug alerts, against 51.9 and 58.4 per cent where the first instance had been accepted. It is the sharpest single measure of whether a channel is being used or merely processed. A value approaching unity means the disposition of an exception class was fixed by its first instance and every later instance is ratified rather than judged — a state that reports as high throughput and low latency in every other metric collected. A.5 A worked illustration The following numbers are illustrative and not measured. No parameter below should be treated as an empirical estimate for any decision class in any firm; as Section 11.1 states, none has been estimated for a deployed agent system. 

Term Illustrative value Source of the value in practice λ, exceptions routed to the named overseer 40 per working day Routing logs Exceptions resolved without routing 360 per working day Same logs Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 99 Term Illustrative value Source of the value in practice IT, routing to closure, median 6 minutes Workflow timestamps, closure not acknowledgement IT, routing to closure, 90th percentile 22 minutes As above c, context re-acquisition cost 2 minutes per switch Local measurement; laboratory default is a floor A, oversight time available per day 120 minutes of a 480- minute day Contracted time less other assigned duties NT, neglect time 40 minutes Interval to threshold quality loss; defined only where the class is verifiable WT, queueing delay 4 minutes Timestamps under simultaneous demand Constraint 1: ρ = λ·(IT + c) / A ≤ 1 At median IT ρ = 40 × (6 + 2) / 120 = 320/120 ≈ 2.7 Insolvent; demand ≈2.7× capacity At 90th-percentile IT ρ = 40 × (22 + 2) / 120 = 960/120 = 8.0 Insolvent; demand 8× capacity λ consistent with ρ ≤ 1 at median IT λ ≤ 15 per day The routing threshold the attention implies Constraint 2: N ≤ NT/(IT + WT) + 1 Implied span 40/(6 + 4) + 1 = 5 concurrent processes A bound on concurrency, not a target Table A1. Worked illustration of the solvency calculation for a single decision class. All parameter values are illustrative and not measured. Two features are worth drawing out, neither dependent on the particular numbers. The gap between the median and the ninetieth percentile changes ρ by a factor of three, which is why a mean is the least informative statistic available. And the binding remedy is a routing decision rather than a supervisory one: bringing ρ to unity at the stated attention requires the arrival rate to fall from forty to fifteen, a decision taken ex ante and therefore constitutive in the sense of Definition 6. Nothing about the overseer changes. That the implied span of five falls inside the range of four to six measured in Section 4.3 is an artefact of the illustrative inputs, not corroboration. A.6 Sensitivity of the constraint to the transferred parameters Section 11.1 concedes that every constant used above was measured somewhere other than knowledge work. The concession is worth more if it is made structural, and the form of the constraint permits that. Inspection of ρ = λ(IT + c)/A shows that the transferred constants enter at exactly one place and enter it in a particular way: IT and c are additive with one another and their sum multiplies λ. Error in either is neither damped nor traded against the other, and the two cannot be estimated jointly by fitting the ratio. What follows is that the quantity governing the sensitivity of the entire calculation is not IT or c taken separately but the ratio c/IT, which decides whether context re-acquisition is a rounding error appended to the handling of an exception or the dominant cost of handling one at all. The same ratio governs the efficacy result of Section 8.3.1, where the number of items a fixed attention budget permits to be examined per period is n = A/(IT + c): a firm wrong about c/IT is wrong about how much oversight its attention buys as well as about whether it can afford the oversight it has assigned. The laboratory anchor must then be stated at its true strength rather than at its convenient one. Monk, Trafton and Boehm-Davis (2008) measured resumption lags of roughly one to two seconds, Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 100 rising logarithmically with interruption length and asymptoting between thirteen and twenty-three seconds of interruption, on a computer-based hierarchical task in which the environmental cues supporting goal retrieval were preserved throughout. Mark, Gonzalez and Harris (2005) observed twenty-four knowledge workers in the field and recorded same-day resumption of an interrupted working sphere averaging twenty-five minutes and twenty-six seconds, a figure itself biased downward by conditioning on resumption occurring at all. The two differ by three orders of magnitude, and the constraint set out in this appendix is evaluated with a parameter whose plausible range in knowledge work spans them. That fact, and not any particular value drawn from within the range, is the sensitivity result. A calculation in which one of two additive terms is uncertain by a factor of a thousand does not have an uncertain answer; it has an answer whose sign, in the sense of solvent or insolvent, is determined by an assumption the firm has not made explicit. Archetype Character of the supervised work IT (elicited range) c (elicited range) c/IT (assumed) NT (elicited range) Dominant term Exception queue in customer service, dedicated overseer High volume, low value per item, dispositions largely reversible; the displaced task is the next item in the same queue Assumed 2– 10 min. No anchor in knowledge work; the measured 15–17 s of the robotics studies is two orders below and is not used as an estimate Assumed 0.5–2 min. Placed near the laboratory end (1–2 s lag, 13–23 s asymptote) because retrieval cues are preserved and the displaced context is the queue itself 0.05–0.5 Assumed 30– 240 min, inferred from reversibility of the disposition rather than from any measured interval λ and IT Credit or underwriting review inside a loan officer's caseload Overseer retains a full operational role; the displaced task is a different customer file carrying its own working state Assumed 10–40 min. No anchor Assumed 5– 25 min, upper end anchored to the field workingsphere resumption figure of 25 min 26 s 0.3–2.5 No anchor and no definition. Proposition 9's counterfactual condition holds for default outcomes, so NT is undefined rather than unmeasured c, the switch Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 101 Archetype Character of the supervised work IT (elicited range) c (elicited range) c/IT (assumed) NT (elicited range) Dominant term Software changereview queue Overseer is normally also an author; the displaced task is their own work in progress, with the deepest working state of the five Assumed 5– 30 min. No anchor Assumed 5– 25 min, same field anchor at the upper end 0.3–3 Assumed minutes to hours where automated tests and cheap reversion exist; no anchor c, the switch Incidentresponse escalation Overseer pre-empted from all other work for the duration; a substantial share of arrivals is exogenous to the supervised processes Assumed 15–90 min to closure. No anchor Assumed 5– 25 min, incurred at entry and again at exit; same field anchor 0.1–1 Assumed seconds to minutes; the only archetype plausibly of the order of the measured robotics range of 30–120 s, and the resemblance is assumed rather than shown NT and the exogenous arrival rate Strategic or capitalallocation review Exercised on a governance cycle; constitutive in the sense of Definition 6 rather than control Assumed hours to days. No anchor of any kind at this scale Assumed hours to days. No anchor of any kind at this scale Order 1, indeterminate Undefined. Proposition 9's condition holds outright Neither; the class sits on the constitutive side, where attention cost is invariant in throughput Table A2. Scenario sensitivity of the solvency constraint across five knowledge-work archetypes. Every range in this table is a start-of-period value, and every range is elicited or assumed for the purpose of sensitivity analysis. No cell is a measurement, and no cell should be cited as an estimate for any decision class in any firm. Where a range is anchored, the anchor is named and is itself a measurement taken in another domain; where no anchor exists, the cell says so. One qualification governs the whole table and is stated before the results are drawn from it. Every range above is a value for an overseer at the start of a period, and Section 4.7 sets out why that matters: none of these parameters is constant within a period, and all of them drift in the direction that tightens the constraint. The governing ratio c/IT is itself time-varying for the same reason, and it rises as the period wears on, because the term in the numerator is the one cumulative interruption degrades faster. Interaction time lengthens with fatigue, but re-acquisition cost is the quantity the evidence shows to be a function of the interruption history directly — resumption lag rose from 1,322 ms after a no-task interruption to 1,789 ms after a high-demand one, with roughly three times the error rate, so an overseer entering a switch already displaced from a demanding exception pays more for it than the same overseer paid at the start of the day. The consequence for the table is that Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 102 the archetypes with the highest interruption frequency are the ones whose tabulated ratios understate the constraint most: the customer-service exception queue, whose tabulated c/IT of 0.05– 0.5 is the lowest in the table but which absorbs the largest number of switches per period, and incident-response escalation, where the exogenous share of arrivals is substantial and each entry and exit carries its own re-acquisition cost. In both rows the start-of-period ratio is the most favourable value the row will take, and a firm reading either as a period-wide characterisation has read the table as more optimistic than it is. No dynamic form is proposed here, for the reasons Section 4.7 gives; the point is only that the direction of the error is known and is the same in every row. Four analytical results follow from the table, and they, rather than any cell in it, are what the subsection is for. The first is that where the overseer is dedicated, c is small relative to IT, the ratio c/ IT is well below unity, and the constraint is governed by the arrival rate and the interaction time. A firm in that position can size its configuration from telemetry it already holds, because the two terms that dominate ρ are both recorded in the routing system, and the term it has not measured contributes little enough that a generous assumption about it does not change the verdict. This is the comfortable case, and it is also the rarer one. The second is that where the overseer retains an operational role — the second, third and fourth rows of the table — c is of the same order as IT or larger, and the constraint is governed by the switch rather than by the work. The managerial consequence is sharper than the observation and runs against the ordinary instinct. If c is of the order of IT or above, halving the interaction time of an exception barely moves ρ, because the term that was halved was never the larger one; a firm that invests in tooling, templates and better exception summaries to cut IT is optimising the smaller addend. Batching, by contrast, divides the number of switches rather than the cost of each, so routing exceptions into scheduled blocks moves ρ a great deal, and moves it by a factor that grows as c/IT grows. The two interventions look comparable on a slide and are not comparable in the constraint. The countervailing consideration is stated in Section 8.3: batching converts an interception into a post-hoc review, and is therefore available only where the consequence latency of the decision class permits it, which is Definition 1 deciding what the solvency arithmetic is allowed to recommend. The third is that NT is the parameter with no knowledge-work anchor at all. Every other term in the constraint has at least a measured analogue in some domain; NT has one only in the robotics studies, and Section 11.1 now establishes something stronger than the absence of a measurement. Where Proposition 9's counterfactual-ground-truth condition holds, NT is not merely unmeasured but undefined, because both the interval and the quality threshold it is stated against presuppose knowledge of the unattended trajectory. Two of the five rows above are in that position and a third is partly so. The span bound is therefore evaluable only in verifiable decision classes, and in the others the solvency test reduces to the arrival-rate condition alone, as Section A.2 records. The fourth follows from the other three. For most firms the practical output of this model is not a number but an ordering: which of the three regimes distinguished in Section 4.2 the configuration occupies — the endogenous regime, in which every exception originates in a supervised process, so that the span bound and the solvency condition are one constraint written twice and the remedy is to supervise fewer processes or to make each safer to leave alone; the exogenous regime, in which part of the stream arrives independently of those processes, so that solvency binds first, each unit of exogenous arrival rate consumes NT processes of span directly, and reducing concurrency will not relieve it; or the mixed and non-stationary regime, in which the exogenous term is bursty and correlated with the failures the supervision exists to catch, so that a configuration solvent in the mean can be insolvent throughout the interval in which the exceptions actually arrive — and, Human-on-the-Loop · WP No. 8 Appendix A. The Oversight Solvency Calculation 103 within that regime, which of IT and c dominates. Both questions are answerable without a point estimate of either parameter, and the answers determine which remedy is worth attempting. A firm that knows only its ordering knows more than a firm holding a precise ρ computed from imported constants, because the ordering is robust to the factor of a thousand and the ρ is not. One measurement would reduce the indeterminacy more than any other, and it is the cheapest of those available. It is a within-firm estimate of c for an overseer who retains an operational role. It requires no new instrumentation and no research design: ordinary work-tracking telemetry already records when a task was opened, when it was closed, and what was worked on in between, so the estimate is obtained by comparing time-to-completion of the displaced task on occasions when an exception intervened with time-to-completion of comparable tasks on occasions when none did, the difference being c plus the interaction time already counted in IT. It satisfies the acceptance rule of Section 11.8, being drawn from a record system maintained for another purpose rather than from anyone's estimate of their own resumption cost. Until some firm reports it, the position of c within the three orders of magnitude between the laboratory and the field is the largest single unknown in this appendix, and every worked figure in Section A.5 should be read as conditional on an assumption about it that no one has yet tested. Appendix B. A Declaration Protocol for a Firm Claiming Human-on-theLoop The following is a protocol rather than an argument: a numbered set of statements a firm should be able to make and evidence, one set per decision class. It is a design proposal of the author. Its purpose is to make an oversight claim falsifiable from outside the firm, which Section 9 identifies as the property that restores the marginal return on repairing oversight. A firm unable to complete an item should publish the gap, since an acknowledged gap is a governance statement and a silent one is not. One rule of acceptance governs every item and is argued for in Section 11.8: each quantity must be derived from a record system maintained for another purpose — routing logs, workflow timestamps, assignment records, the injection harness — and none may be satisfied by a practitioner's estimate or a manager's judgement, self-report in this domain having been measured and found capable of differing from the measurement in sign. The eleven items are grouped into two tiers, because they do not cost the same thing and a firm entitled to know the price before it commits should be told which items it is committing to. Tier 1 comprises the eight items a firm can complete from a policy exercise and records it already keeps, and then republish on a governance cycle: a classification stated, an authority named, a register entry recorded, a reporting route described, a reliance disclosed, and quantities read off routing systems and workflow timestamps that exist for other reasons. Tier 2 comprises the three items that require ongoing measurement or the consumption of production capacity, and whose cost therefore recurs in every period and falls on the operation the automation was installed to accelerate. The evidentiary warrant sits in Tier 2. The Tier 1 items describe a configuration; the Tier 2 items test whether it works — two of them resting on the only intervention shown to move the low-prevalence miss rate and the only one shown to restore automation-failure detection, and the third being the only means by which a firm operating a triage layer can establish the recall that bounds what that layer removes. The prediction of Section 11.9 follows directly: a firm minimising cost will adopt Tier 1 and decline Tier 2, and a protocol so completed is substantially complete and evidentially empty — a declaration of position rather than a control. Tier 1 is nonetheless worth adopting first, and the sequencing advice should be given honestly rather than withheld to force the harder items: declaring the loop position and the independence register is what makes the absence of Tier 2 Human-on-the-Loop · WP No. 8 Appendix B. A Declaration Protocol for a Firm Claiming Human-on-the-Loop 104 visible to a reader outside the firm, and an absence that can be seen is a governance fact, whereas an unstated one is not. The minimum audit trail: what makes a Tier 1 declaration checkable Tier 1 is the cheap tier and is for that reason the gameable one. Section 11.8 names the vectors against the paper's own instrument, and each of them operates on a Tier 1 quantity: the arrival rate is lowered by re-routing exceptions away from the named overseer or by redefining what counts as an exception; interaction time is halved by measuring to first touch rather than to closure; uncommitted attention is inflated by reporting nominal availability in place of the subtractive residual; and the solvency statement as a whole is satisfied by splitting a decision class into several, every term being defined per class. None of these moves is detectable in the published figures themselves, because each produces a well-formed number. What distinguishes a declaration that can be checked from one that can only be read is therefore not the quantities but the records standing behind them, and a firm making a Tier 1 declaration should be able to place four kinds of record before a third-party auditor or a regulator. The first concerns the unit of account. Each declared decision class must carry a stable published identifier, that identifier must not change within a reporting period, and any change between periods must be disclosed as a restatement stating both the prior and the revised boundaries of the class. This is what defeats class-splitting: a class that subdivides is visible in the taxonomy rather than in the arithmetic, and a firm that must publish the old and new boundaries side by side cannot divide a load on paper without saying that it has. The second concerns the timing of every routed item. Each item routed to a human must carry timestamps for the routing event, for first touch and for closure, retained at the item level rather than as a period aggregate, so that interaction time cannot be reported from first touch alone and the choice of endpoint is visible to a reader rather than being an internal accounting convention. The same records are what make Section 4.7's earlyand-late reporting possible, since they carry the time within the period at which each item arrived. The third concerns the attention figure. A must be derived by subtraction from a roster, assignment or workload system maintained for another purpose — contracted time in that system less every other duty carrying a delivery expectation — in accordance with the acceptance rule Section 11.8 takes from the series' fifth working paper, and never from a practitioner's estimate or a manager's judgement of how much of the period is available for oversight. The randomised evidence recorded in Section 11.8 is the warrant: self-report in this domain has been measured against the record and found capable of differing from it in sign, which disqualifies it as a source for a quantity on which a solvency claim turns. The fourth concerns duration. The records supporting all three must be retained for at least two reporting periods, so that a trend and not merely a level can be audited; the override rate and the utilisation ratio are quantities whose movement carries the diagnostic content, and a single period of retained data permits no movement to be seen. The limit of this should be stated in the same breath. An audit trail of this kind establishes that the reported numbers describe what the firm actually did; it does not establish that the decision-class boundaries were drawn in good faith in the first place, which remains a judgement no log can settle. Tier 1 (eight items) — declarative: obtainable from a policy exercise and existing records The declared loop position and the measured latencies supporting it. Publish τsys, τH , the method by which each was measured, the resulting classification under Definition 1, and the remeasurement interval. Supporting evidence: loop position is measured rather than declared, and τH is not constant — in the only deployed dataset available, reaction time increased with 1. Human-on-the-Loop · WP No. 8 Appendix B. A Declaration Protocol for a Firm Claiming Human-on-the-Loop 105 cumulative autonomous operation even among paid professional safety drivers whose sole task was to watch, so a firm measuring τH once at deployment has measured the most favourable moment in the system's life. The solvency statement. Publish λ, IT at the median and ninetieth percentile, c, A stated as oversight time available per period, NT, WT, the implied span N, the realised utilisation ρ = λ(IT + c)/A at both the median and the ninetieth percentile, and the two inequalities of Appendix A evaluated against them, together with the override rate, its trend, and the path-dependence statistic. Every quantity must carry the point in the period at which it was measured, and the utilisation ratio must be reported at two measurement points, early and late in the period, the difference between the two being the fatigue term: none of A, IT, c and the overseer's detection probability is a constant within a period, and all four drift in the direction that tightens the constraint, so a single mid-period figure records as solvent a channel that may stand at ρ = 0.8 early and ρ = 1.4 late. This is the requirement Section 4.7 states and Appendix A.2 imposes on every term of the calculation, and the item-level timestamps required by the minimum audit trail above are what make the partition obtainable without additional instrumentation. Supporting evidence: insolvency presents as acceptance rather than refusal, so no metric a firm currently collects would reveal it, and fourteen years of interface improvement in clinical alerting moved override rates essentially not at all. The independence register entry. For every control relied upon, state whether the verifier shares a base model, a training corpus, an input channel, a vendor or an incentive with the generator it verifies, and the error classes for which independence is therefore not claimed. Wherever a triage or routing layer selects which items reach a human at all, the same five relations must be recorded for the router as for the verifier, because Definition 5 bites one layer earlier: a routing rule built on the same base model as the generator, trained on the same corpus or reading the same input channel fails to route precisely the error classes the generator is most prone to commit, and fails to route them non-randomly. 

The router's dependence is the more damaging of the two, for the reason the second consequence of Section 8.3.1 gives: an error a correlated verifier examines and misses is at least in the review record, whereas an error a correlated router never selects is outside the examined stream altogether, recoverable by no diligence and no additional attention, and leaves no trace of a review that was never attempted. Supporting evidence: Definition 5, and the field observation that a supervisory agent built on the same base model as the agent it supervised authorised requests about eight times as often as it denied them — an outcome its publisher attributed to the two sharing blind spots because they were the same model. The decision-order design. State whether the overseer records an independent judgement before the system's output is displayed, and publish the rate of divergence between that prior judgement and the final disposition. Supporting evidence: the predict-then-see ordering reduced disparity in algorithmic influence by 81.5 per cent and in deviation by 73.9 per cent relative to showing the score first, the ordering rather than the content carrying the effect. Its second merit bears on Proposition 8: the recorded unassisted judgement is itself an independent channel, which the amended proposition identifies not as the condition under which oversight improves accuracy but as the strongest and most auditable source of the error decorrelation that is that condition. Decorrelation is what the proposition requires; an independent channel is the way a firm can most readily produce it and most readily evidence that it has. The placement of any automated monitor within the three-lines architecture, and the identity of the third-line function. State where each automated monitor sits, and name the third-line function, which cannot be an automated component of the operation it evaluates: 2. 3. 4. 5. Human-on-the-Loop · WP No. 8 Appendix B. A Declaration Protocol for a Firm Claiming Human-on-the-Loop 106 under the Three Lines Model an AI monitor is a first- or second-line control, third-line independence being a principle rather than a configuration option. Automated monitoring may carry the ongoing evaluation, but the separate evaluation, and the judgement of what constitutes a deficiency, must be held by a named human function. Supporting evidence: Ghosh, Shetty, Bansal and Nath (2022), a peer-reviewed empirical study of incidents in a large-scale cloud service, report that in that heavily instrumented production environment roughly 55 per cent of high-severity incidents were first detected by automated monitors and 29 per cent first reported by external customers — some 45 per cent were not first caught by monitoring. The escalation path and the named holder of the authority to halt. Name the individual holding authority to disregard, reverse or halt the decision class, and evidence that the authority is unencumbered: what approval must be obtained before exercising it, what the exercise costs the holder in throughput terms, and how many times it was exercised in the reporting period. Supporting evidence: the third failure condition of Definition 2, and the deployer duty to assign oversight to natural persons possessing the necessary competence, training, authority and support — two of whose four elements are organisational rather than cognitive, and are the two a firm can withhold without saying so. The board-level reporting route, its frequency and its content. State which board or committee receives the foregoing items, at what frequency and in what form, and specifically whether the solvency statement, the override rate and its trend, and the independence register are among them, or whether what reaches the board is an assurance that oversight is in place. Supporting evidence: the binding duty today is corporate law rather than AI law, and for a mission-critical risk the monitoring system must be board-level, dedicated and actually used; a board that publicly represented that it monitored a safety risk in ways it did not assembled the fact pattern that settled for $237.5 million. Claiming oversight one does not perform is itself part of what the violation consists in — a step Section 10.5 advances in express dependence on that record as reported at secondary level, and only within the boundary Section 10.5.1 states — which makes what is reported upward — quantities, or assurances — the item worth disclosing. That boundary should be read with this item rather than assumed away: the threshold is bad faith and not negligence, so neither an operational failure of the supervised system nor insolvency under Definition 4 is by itself a breach of any duty, and what this item makes checkable is the accuracy of what the firm has represented about its own control environment. The anti-pattern declaration. State whether the firm relies for oversight on model explanations, confidence displays alone, a supervisor agent on the same base model, instructions to be more careful, or response-time service levels as a proxy for oversight quality, and if so on what basis, given that each has been directly tested against the failure mode it is offered to address and has failed. This item exists because the other ten items can be satisfied in form by a firm that has installed exactly these measures, whose documentation would then be indistinguishable from that of a firm with a functioning control. Tier 2 (three items) — continuous: requiring ongoing measurement or the consumption of production capacity The synthetic-exception injection rate and catch rate. State the rate at which known synthetic exceptions are injected into the overseer's stream, the ground-truth feedback given on each judgement, and the resulting catch rate, reported separately from that on genuine exceptions. Supporting evidence: high-prevalence bursts with full feedback are the only intervention in Wolfe and colleagues (2007) that moved the low-prevalence miss rate at all, where payoff manipulation and forced slowing had no effect — but the warrant must be stated at its true 6. 7. 8. 9. Human-on-the-Loop · WP No. 8 Appendix B. A Declaration Protocol for a Firm Claiming Human-on-the-Loop 107 strength, as Section 5.3 sets out. In Experiment 7 (N = 14) the low-prevalence miss rate was 21 per cent against 18 per cent in the matched high-prevalence condition, a difference not significant at p = .312; the 46 per cent customarily set against that 21 per cent is the lowprevalence baseline of Experiment 1a, a different experiment with different observers, so the headline contrast is across experiments rather than within a design. A failure to reject the null at N = 14 is weak evidence of elimination. The within-design replication that would establish it — high- and low-prevalence conditions with and without feedback, at adequate power — has not been run, and this declaration item is therefore in part a request that firms operating an injection regime generate it, since a firm running injections at scale is generating exactly that design as a by-product and need only report it. The schedule of returns to unassisted operation. State the frequency and duration of scheduled periods in which the decision class is executed without automated assistance, and the detection performance measured before and after. Supporting evidence: a ten-minute scheduled return to manual control significantly improved automation-failure detection, and the benefit persisted after automation resumed. The related proposal in the series' Brain Capital Management paper is a design proposal, not support. The residual uniform sample of the unrouted stream, and the routing recall it estimates. Where any rule selects which items reach a human, publish the rate at which items the rule declined to route are drawn uniformly at random and reviewed anyway, the basis on which that sample was sized — which is the precision required of an estimate of the rule's recall, and not the number of errors the sample is expected to catch — and the resulting estimate of recall, reported alongside any reduction in review workload claimed for the triage layer, so that the two figures are read together rather than one without the other. Supporting evidence: the sixth step of Section 8.3.1, in which the share of consequential errors a triage design removes is bounded above by the recall of its router, so that recall is the parameter on which the design's efficacy turns; and the reason the sample cannot be dispensed with, which is that recall is not observable from the routed stream. Everything a firm sees of its routing rule — its precision, the outcomes of the reviews it commissions, the catch rate achieved within it — is conditioned on having been routed, whereas recall is a statement about what was not, so the only route to it is to look at the stream the rule rejected. The sample is a permanent cost of the design rather than a transitional one, because recall is a property of a rule operating on a distribution that drifts, and a rule validated once is not thereby validated later. Where the router is an instance of the supervised system, is trained on the same data or reads the same inputs, Definition 5 applies to it and the item should record that dependence in the same terms as the independence register. The boundary condition must be declared rather than left implicit: where the decision class is one in which ground truth does not arrive after the fact, sampling the unrouted stream yields items whose correct disposition is itself unknown, the estimate cannot be produced by this or any other means, and the firm should state plainly that the central parameter of its triage layer is unknown and that the workload reduction it reports is therefore unpriced in errors. 10. 11. Human-on-the-Loop · WP No. 8 Appendix B. A Declaration Protocol for a Firm Claiming Human-on-the-Loop 108 References Agarwal, N., Moehring, A., Rajpurkar, P., & Salz, T. (2023). Combining human expertise with artificial intelligence: Experimental evidence from radiology (NBER Working Paper No. 31422). National Bureau of Economic Research. https://doi.org/10.3386/w31422 [Working paper; not peer-reviewed.] Aghion, P., & Tirole, J. (1997). Formal and real authority in organizations. Journal of Political Economy, 105(1), 1–29. https://doi.org/10.1086/262063 AI Incident Database. (2025). Incident 1152: Replit AI coding agent deletes production database during code freeze. Responsible AI Collaborative. https://incidentdatabase.ai/cite/1152/ [Grey literature, corroborated by trade reporting; no regulatory or judicial record was located.] Altmann, E. M., & Trafton, J. G. (2002). Memory for goals: An activation-based model. Cognitive Science, 26(1), 39–83. https://doi.org/10.1207/s15516709cog2601_2 Amnesty International. (2021, October). Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal (Index EUR 35/4686/2021). [NGO grey literature; not peer-reviewed.] Ancker, J. S., Edwards, A., Nosal, S., Hauser, D., Mauer, E., & Kaushal, R. (2017). Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Medical Informatics and Decision Making, 17, Article 36. https://doi.org/10.1186/s12911-017-0430-8 [A published Correction exists (PMID 31739801); figures should be checked against the corrected version.] Anthropic. (2025a). Project Vend: Can Claude run a small shop? Anthropic and Andon Labs. https:// www.anthropic.com/research/project-vend-1 [Company technical report; not peer-reviewed.] Anthropic. (2025b, December 18). Project Vend: Phase two. Anthropic Frontier Red Team. https:// www.anthropic.com/research/project-vend-2 [Company technical report; not peer-reviewed, and no net profit figure is disclosed.] Autoriteit Persoonsgegevens. (2021, December). Decision imposing an administrative fine of €2.75 million on the Belastingdienst (Dutch Tax and Customs Administration) for discriminatory and unlawful data processing. [Primary regulatory record.] Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779. https://doi.org/ 10.1016/0005-1098(83)90046-8 [A five-page “Brief Paper”; an analytical essay with no sample, design, data or statistics.] Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). ACM. https://doi.org/ 10.1145/3411764.3445717 Barres, V., Dong, H., Ray, S., Si, X., & Narasimhan, K. (2025). τ²-Bench: Evaluating conversational agents in a dual-control environment (arXiv:2506.07982). Sierra Research. https://arxiv.org/abs/2506.07982 [Preprint; not peer-reviewed.] Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced opensource developer productivity (arXiv:2507.09089). METR. https://arxiv.org/abs/2507.09089 [Preprint; not peer-reviewed, and superseded in part by a METR revision of 24 February 2026 whose preliminary estimates have confidence intervals crossing zero.] Ben-Michael, E., Greiner, D. J., Huang, M., Imai, K., Jiang, Z., & Shin, S. (2025). Does AI help humans make better decisions? A statistical evaluation framework for experimental and observational studies. Proceedings of the National Academy of Sciences, 122(38), e2505106122. https://doi.org/10.1073/pnas. 2505106122 Bevan, G., & Hood, C. (2006). What’s measured is what matters: Targets and gaming in the English public health care system. Public Administration, 84(3), 517–538. https://doi.org/10.1111/j.1467-9299.2006.00600.x [Documentary case analysis of a natural policy experiment; no sample and no counterfactual.] Bloom, N., Garicano, L., Sadun, R., & Van Reenen, J. (2014). The distinct effects of information technology and communication technology on firm organization. Management Science, 60(12), 2859–2885. https://doi.org/ 10.1287/mnsc.2014.2013 Human-on-the-Loop · WP No. 8 References 109 Bo, J. Y., Wan, S., & Anderson, A. (2025). To rely or not to rely? Evaluating interventions for appropriate reliance on large language models. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25) (pp. 1–23). ACM. https://doi.org/10.1145/3706598.3714097 Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188, 1–21. https://doi.org/10.1145/3449287 Cameron, L. D., & Rahman, H. A. (2022). Expanding the locus of resistance: Understanding the co-constitution of control and resistance in the gig economy. Organization Science, 33(1), 38–58. https://doi.org/10.1287/ orsc.2021.1557 Casner, S. M., Geven, R. W., Recker, M. P., & Schooler, J. W. (2014). The retention of manual flying skills in the automated cockpit. Human Factors, 56(8), 1506–1516. https://doi.org/10.1177/0018720814535628 Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., Chopra, B., Tiwari, R., Keutzer, K., Parameswaran, A., Klein, D., Ramchandran, K., Zaharia, M., Gonzalez, J. E., & Stoica, I. (2025). Why do multi-agent LLM systems fail? In Advances in Neural Information Processing Systems: Datasets and Benchmarks Track (NeurIPS 2025). Preprint arXiv:2503.13657. Center for AI Safety. (2026, July 1). A significant increase in digital labor automation [Blog post]. https://safe.ai/ blog/significant-increase-in-digital-labor-automation [Organisation blog post; not peer-reviewed.] Childcare Allowance Parliamentary Inquiry Committee [Parlementaire ondervragingscommissie Kinderopvangtoeslag]. (2020, December 17). Ongekend onrecht [Unprecedented injustice]. House of Representatives of the States General. English text: Venice Commission CDL-REF(2021)073. [Official parliamentary inquiry report; not peer-reviewed.] Committee of Sponsoring Organizations of the Treadway Commission. (2013). Internal control—Integrated framework. COSO. [Professional framework; not peer-reviewed and paywalled; the component and principle names were verified through secondary reproduction of the official list.] Commodity Futures Trading Commission & Securities and Exchange Commission. (2010, September 30). Findings regarding the market events of May 6, 2010: Report of the staffs of the CFTC and SEC to the Joint Advisory Committee on Emerging Regulatory Issues. [Primary regulatory record.] Crandall, J. W., & Cummings, M. L. (2007). Identifying predictive metrics for supervisory control of multiple robots. IEEE Transactions on Robotics, 23(5). [Figures read from the authors’ final preprint; the published page range was not independently verified.] Crandall, J. W., Cummings, M. L., Della Penna, M., & de Jong, P. M. A. (2011). Computing the effects of operator attention allocation in human control of multiple robots. IEEE Transactions on Systems, Man, and Cybernetics — Part A: Systems and Humans, 41(3), 385–397. https://doi.org/10.1109/TSMCA.2010.2084082 Cummings, M. L., & Mitchell, P. J. (2008). Predicting controller capacity in supervisory control of multiple UAVs. IEEE Transactions on Systems, Man, and Cybernetics — Part A: Systems and Humans, 38(2), 451–460. https://doi.org/10.1109/TSMCA.2007.914757 [Bibliographic record only; no quantitative result from this paper was verified and none is cited.] De-Arteaga, M., Fogliato, R., & Chouldechova, A. (2020). A case for humans-in-the-loop: Decisions in the presence of erroneous algorithmic scores. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20). ACM. https://doi.org/10.1145/3313831.3376638 Delaware General Corporation Law, §§ 102(b)(7), 144 and 220. Delaware Senate Bill 21 (2025), signed 25 March 2025, amending Delaware General Corporation Law §§ 144 and 220. [Verified at secondary level only; it did not amend the oversight duty.] Dessein, W. (2002). Authority and communication in organizations. The Review of Economic Studies, 69(4), 811–838. https://doi.org/10.1111/1467-937X.00227 Dixit, V. V., Chand, S., & Nair, D. J. (2016). Autonomous vehicles: Disengagements, accidents and reaction times. PLOS ONE, 11(12), e0168054. https://doi.org/10.1371/journal.pone.0168054 Docherty, B. (2012, November 19). Losing humanity: The case against killer robots. Human Rights Watch & Harvard Law School International Human Rights Clinic. https://www.hrw.org/report/2012/11/19/losinghumanity/case-against-killer-robots [NGO report; not peer-reviewed, and it cites no prior source for the taxonomy.] Drew, B. J., Harris, P., Zègre-Hemsey, J. K., Mammone, T., Schindler, D., Salas-Boni, R., Bai, Y., Tinoco, A., Ding, Q., & Hu, X. (2014). Insights into the problem of alarm fatigue with physiologic monitor devices: A Human-on-the-Loop · WP No. 8 References 110 comprehensive observational study of consecutive intensive care unit patients. PLOS ONE, 9(10), e110274. https://doi.org/10.1371/journal.pone.0110274 Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human-robot interaction. Engaging Science, Technology, and Society, 5, 40–60. https://doi.org/10.17351/ests2019.260 [The case pairing was verified from the journal record and secondary summaries only; the article body could not be retrieved.] Endsley, M. R. (2017). From here to autonomy: Lessons learned from human–automation research. Human Factors, 59(1), 5–27. https://doi.org/10.1177/0018720816681350 Eriksson, A., & Stanton, N. A. (2017). Takeover time in highly automated vehicles: Noncritical transitions to and from manual control. Human Factors, 59(4), 689–705. https://doi.org/10.1177/0018720816685832 [Record verified; the widely circulated long-tail range could not be confirmed and is not used.] Falk, A., & Kosfeld, M. (2006). The hidden costs of control. American Economic Review, 96(5), 1611–1630. https://doi.org/10.1257/aer.96.5.1611 Flight deck automation erodes fine-motor flying skills among airline pilots. (2016). Human Factors. https:// doi.org/10.1177/0018720816640394 [Authorship, volume and pagination were not independently verified; cited only as an unresolved contrary finding, and no figure is quoted from it.] Fok, R., & Weld, D. S. (2024). In search of verifiability: Explanations rarely enable complementary performance in AI-advised decision making. AI Magazine, 45(3), 317–332. https://doi.org/10.1002/aaai. 12182 [Peer-reviewed review and position article, not primary research; volume and issue should be confirmed against the publisher.] Gershon, P., Mehler, B., & Reimer, B. (2023). Driver response and recovery following automation initiated disengagement in real-world hands-free driving. Traffic Injury Prevention, 24(4), 356–361. https://doi.org/ 10.1080/15389588.2023.2189990 Ghosh, S., Shetty, M., Bansal, C., & Nath, S. (2022). How to fight production incidents? An empirical study on a large-scale cloud service. In Proceedings of the 13th Symposium on Cloud Computing (SoCC ’22). ACM. https://doi.org/10.1145/3542929.3563482 Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45, 105681. https://doi.org/10.1016/j.clsr.2022.105681 Green, B., & Chen, Y. (2019a). Disparate interactions: An algorithm-in-the-loop analysis of fairness in risk assessments. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19) (pp. 90–99). ACM. https://doi.org/10.1145/3287560.3287563 Green, B., & Chen, Y. (2019b). The principles and limits of algorithm-in-the-loop decision making. Proceedings of the ACM on Human-Computer Interaction, 3(CSCW), Article 50, 1–24. https://doi.org/10.1145/3359152 [Participant counts were taken from the author-hosted version; verify them against the ACM version.] Group of Governmental Experts on Emerging Technologies in the Area of Lethal Autonomous Weapons Systems. (2019). Guiding principles affirmed by the Group of Governmental Experts (Report of the 2019 session, CCW/GGE.1/2019/3, Annex IV). United Nations Office at Geneva. [Principle text verified from ICRC commentary hosted by UNODA rather than from the UN document itself.] Guadalupe, M., & Wulf, J. (2010). The flattening firm and product market competition: The effect of trade liberalization on corporate hierarchies. American Economic Journal: Applied Economics, 2(4), 105–127. https://doi.org/10.1257/app.2.4.105 Holmström, B. (1979). Moral hazard and observability. The Bell Journal of Economics, 10(1), 74–91. https:// doi.org/10.2307/3003320 Holmstrom, B., & Milgrom, P. (1991). Multitask principal–agent analyses: Incentive contracts, asset ownership, and job design. The Journal of Law, Economics, and Organization, 7(special issue), 24–52. https://doi.org/ 10.1093/jleo/7.special_issue.24 Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large language models cannot self-correct reasoning yet. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024). Preprint arXiv:2310.01798. In re Boeing Co. Derivative Litigation, 2021 WL 4059934 (Del. Ch. 7 September 2021), C.A. No. 2019-0907-MTZ; settled for $237.5 million. [Slip opinion not obtained; the record is described from a law-firm case summary and nothing is quoted.] In re Caremark International Inc. Derivative Litigation, 698 A.2d 959 (Del. Ch. 1996). Human-on-the-Loop · WP No. 8 References 111 Institute of Internal Auditors. (2020, July). The IIA’s three lines model: An update of the three lines of defense. IIA. [Professional framework; not peer-reviewed, and the currently hosted copy may carry a later revision date.] ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. International Organization for Standardization, December 2023. [Paywalled; the normative text was not read, and no claim is made about what the standard requires.] Jamieson, G. A., & Skraaning, G. (2018). Levels of automation in human factors models for automation design: Why we might consider throwing the baby out with the bathwater. Journal of Cognitive Engineering and Decision Making, 12(1), 42–49. https://doi.org/10.1177/1555343417732856 Kadowaki, N. (2026a). Future value theory (VURA Working Paper Series No. 1). VURA Capital Innovation Holdings, Inc. ISSN 2761-011X. https://doi.org/10.5281/zenodo.21255663 Kadowaki, N. (2026b). Enterprise redefinition (VURA Working Paper Series No. 2). VURA Capital Innovation Holdings, Inc. ISSN 2761-011X. https://doi.org/10.5281/zenodo.21718637 Kadowaki, N. (2026c). Enterprise redefinition observed (VURA Working Paper Series No. 3). VURA Capital Innovation Holdings, Inc. ISSN 2761-011X. https://doi.org/10.5281/zenodo.21850022 Kadowaki, N. (2026d). From job description to purpose description (VURA Working Paper Series No. 4). VURA Capital Innovation Holdings, Inc. ISSN 2761-011X. https://doi.org/10.5281/zenodo.21872816 Kadowaki, N. (2026e). Brain capital management (VURA Working Paper Series No. 5). VURA Capital Innovation Holdings, Inc. ISSN 2761-011X. https://doi.org/10.5281/zenodo.21942227 Kadowaki, N. (2026g). Redefinition capitalism (VURA Working Paper Series No. 7). VURA Capital Innovation Holdings, Inc. ISSN 2761-011X. https://doi.org/10.5281/zenodo.21945763 Kellogg, K. C., Valentine, M. A., & Christin, A. (2020). Algorithms at work: The new contested terrain of control. Academy of Management Annals, 14(1), 366–410. https://doi.org/10.5465/annals.2018.0174 [Integrative review and theory-building essay; not primary empirical research, and it reports no effect sizes.] Krakowski, I., Kim, J., Cai, Z. R., Daneshjou, R., Lapins, J., Eriksson, H., Lykou, A., & Linos, E. (2024). Human-AI interaction in skin cancer diagnosis: A systematic review and meta-analysis. npj Digital Medicine, 7, 78. https://doi.org/10.1038/s41746-024-01031-w Kwa, T., West, B., Becker, J., Deng, A., et al. (2025). Measuring AI ability to complete long tasks (arXiv: 2503.14499). METR. https://arxiv.org/abs/2503.14499 [Preprint; not peer-reviewed.] Lång, K., Josefsson, V., Larsson, A.-M., Larsson, S., Högberg, C., Sartor, H., Hofvind, S., Andersson, I., & Rosso, A. (2023). Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): A clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study. The Lancet Oncology, 24(8), 936–944. https://doi.org/10.1016/S1470-2045(23)00298-X Louw, T., Markkula, G., Boer, E., Madigan, R., Carsten, O., & Merat, N. (2017). Coming back into the loop: Drivers’ perceptual-motor performance in critical events after automated driving. Accident Analysis & Prevention, 108, 9–18. https://doi.org/10.1016/j.aap.2017.08.011 Marchand v. Barnhill, 212 A.3d 805 (Del. 2019). [Slip opinion not obtained; cited for the proposition only, and nothing is quoted.] Mark, G., Gonzalez, V. M., & Harris, J. (2005). No task left behind? Examining the nature of fragmented work. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’05) (pp. 321–330). ACM. https://doi.org/10.1145/1054972.1055017 Mark, G., Gudith, D., & Klocke, U. (2008). The cost of interrupted work: More speed and stress. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’08) (pp. 107–110). ACM. https:// doi.org/10.1145/1357054.1357072 Matthias, A. (2004). The responsibility gap: Ascribing responsibility for the actions of learning automata. Ethics and Information Technology, 6(3), 175–183. https://doi.org/10.1007/s10676-004-3422-1 Mazeika, M., Gatti, A., Hendrycks, D., et al. (2025). Remote Labor Index: Measuring AI automation of remote work (arXiv:2510.26787). Center for AI Safety & Scale AI. https://arxiv.org/abs/2510.26787 [Preprint; not peer-reviewed, and produced jointly by organisations with an interest in the result.] METR. (2026, January 29). Time horizon 1.1 [Blog post]. https://metr.org/blog/2026-1-29-time-horizon-1-1/ [Organisation blog post; not peer-reviewed.] Human-on-the-Loop · WP No. 8 References 112 Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, 14 February 2024). [Primary tribunal record, reached through secondary legal analyses because CanLII blocks automated retrieval; the tribunal’s reasoning is not quoted verbatim.] Molloy, R., & Parasuraman, R. (1996). Monitoring an automated system for a single failure: Vigilance and task complexity effects. Human Factors, 38(2), 311–322. https://doi.org/10.1177/001872089606380211 [Detection percentages were not obtainable and none is quoted.] Monk, C. A., Trafton, J. G., & Boehm-Davis, D. A. (2008). The effect of interruption duration and demand on resuming suspended goals. Journal of Experimental Psychology: Applied, 14(4), 299–313. https://doi.org/ 10.1037/a0014402 Monperrus, M. (2026). The end of code review: Coding agents supersede human inspection (arXiv: 2606.13175v1). KTH Royal Institute of Technology. [Preprint, and explicitly a position paper presenting no new empirical study.] Nanji, K. C., Slight, S. P., Seger, D. L., Cho, I., Fiskio, J. M., Redden, L. M., Volk, L. A., & Bates, D. W. (2014). Overrides of medication-related clinical decision support alerts in outpatients. Journal of the American Medical Informatics Association, 21(3), 487–491. https://doi.org/10.1136/amiajnl-2013-001813 National Institute of Standards and Technology. (2023, January). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). [Voluntary framework.] National Institute of Standards and Technology. (2024, July). Artificial intelligence risk management framework: Generative artificial intelligence profile (NIST AI 600-1). [Voluntary framework.] Nissenbaum, H. (1996). Accountability in a computerized society. Science and Engineering Ethics, 2(1), 25–42. https://doi.org/10.1007/BF02639315 Ocasio, W. (1997). Towards an attention-based view of the firm. Strategic Management Journal, 18(S1), 187– 206. https://doi.org/10.1002/(SICI)1097-0266(199707)18:1+<187::AID-SMJ936>3.0.CO;2-K Olsen, D. R., & Goodrich, M. A. (2003). Metrics for evaluating human-robot interactions. In Proceedings of PERMIS 2003 (NIST Performance Metrics for Intelligent Systems Workshop). [Workshop paper; not peerreviewed, and it reports no measured values.] Onnasch, L., Wickens, C. D., Li, H., & Manzey, D. (2014). Human performance consequences of stages and levels of automation: An integrated meta-analysis. Human Factors, 56(3), 476–488. https://doi.org/ 10.1177/0018720813501549 Parasuraman, R., Bahri, T., Molloy, R., & Singh, I. (1992). Adaptive automation and human performance: II. Effects of shifts in the level of automation on operator performance (Report NAWCADWAR-92036-60). Naval Air Warfare Center Aircraft Division, Warminster, PA. Defense Technical Information Center accession ADA254127. [Government technical report; not peer-reviewed.] Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055 [Integrative review; it reports no sample of its own.] Parasuraman, R., Molloy, R., & Singh, I. L. (1993). Performance consequences of automation-induced “complacency”. International Journal of Aviation Psychology, 3(1), 1–23. https://doi.org/10.1207/ s15327108ijap0301_1 [Bibliographic record verified; the article was not independently obtained and no figure is quoted from it.] Parasuraman, R., Mouloua, M., & Molloy, R. (1996). Effects of adaptive task allocation on monitoring of automated systems. Human Factors, 38(4), 665–679. https://doi.org/10.1518/001872096778827279 Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics — Part A: Systems and Humans, 30(3), 286–297. https://doi.org/10.1109/3468.844354 Patwardhan, T., Dias, R., Proehl, E., Kim, G., Wang, M., Watkins, O., Posada Fishman, S., Aljubeh, M., Thacker, P., Fauconnet, L., Kim, N. S., Chao, P., Miserendino, S., Chabot, G., Li, D., Sharman, M., Barr, A., Glaese, A., & Tworek, J. (2025). GDPval: Evaluating AI model performance on real-world economically valuable tasks (arXiv:2510.04374). OpenAI. https://arxiv.org/abs/2510.04374 [Preprint; not peer-reviewed, and produced by a party with a commercial interest in the result.] Police and Criminal Evidence Act 1984 (United Kingdom), section 69 (repealed in 1999 by the Youth Justice and Criminal Evidence Act 1999). Human-on-the-Loop · WP No. 8 References 113 Poly, T. N., Islam, M. M., Yang, H.-C., & Li, Y.-C. (2020). Appropriateness of overridden alerts in computerized physician order entry: Systematic review. JMIR Medical Informatics, 8(7), e15653. https://doi.org/ 10.2196/15653 Post Office Horizon IT Inquiry. (2025, July 8). Final report, volume 1 (Chair: Sir Wyn Williams; HC 1119). [Primary inquiry record.] Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J., & Wallach, H. (2021). Manipulating and measuring model interpretability. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). ACM. https://doi.org/10.1145/3411764.3445315 Qin, X., Zhou, X., Chen, C., Wu, D., Zhou, H., Dong, X., Cao, L., & Lu, J. G. (2025). AI aversion or appreciation? A capability–personalization framework and a meta-analytic review. Psychological Bulletin, 151(5), 580–599. https://doi.org/10.1037/bul0000477 Rajan, R. G., & Wulf, J. (2006). The flattening firm: Evidence from panel data on the changing nature of corporate hierarchies. The Review of Economics and Statistics, 88(4), 759–773. https://doi.org/10.1162/rest. 88.4.759 Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), OJ L, 2024/1689, 12.7.2024; in force 1 August 2024. [The text of Articles 14 and 26 was read from a faithful mirror of the Official Journal rather than from the Journal itself.] Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 as regards the simplification of the implementation of harmonised rules on artificial intelligence (Digital Omnibus on AI), adopted 8 July 2026, in force 27 July 2026. [Existence, title, adoption and entry-into-force dates verified at EUR-Lex; the postponement dates rest on a single reading of the amending text.] Royal Commission into the Robodebt Scheme. (2023, July 7). Report (Commissioner Catherine Holmes AC SC). [Official royal commission report; not peer-reviewed, and its details were corroborated through secondary sources because the Commission’s own site blocked automated retrieval.] Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful human control over autonomous systems: A philosophical account. Frontiers in Robotics and AI, 5, Article 15. https://doi.org/10.3389/frobt.2018.00015 Securities and Exchange Commission. (2013, October 16). In the matter of Knight Capital Americas LLC (Administrative Proceeding File No. 3-15570, Release No. 34-70694). [Primary regulatory record.] See, J. E., Howe, S. R., Warm, J. S., & Dember, W. N. (1995). Meta-analysis of the sensitivity decrement in vigilance. Psychological Bulletin, 117(2), 230–249. https://doi.org/10.1037/0033-2909.117.2.230 Sheridan, T. B., & Verplank, W. L. (1978). Human and computer control of undersea teleoperators (Technical report). Man–Machine Systems Laboratory, Department of Mechanical Engineering, Massachusetts Institute of Technology. OCLC 8544670. [Grey literature; not peer-reviewed, and cited through the route stated in Section 2.1.] Simon, H. A. (1971). Designing organizations for an information-rich world. In M. Greenberger (Ed.), Computers, communications, and the public interest (pp. 37–72). The Johns Hopkins Press. [Book chapter; not peer-reviewed. The quotation is at p. 40.] Stechly, K., Valmeekam, K., & Kambhampati, S. (2025). On the self-verification limitations of large language models on reasoning and planning tasks. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025). Preprint arXiv:2402.08115. Stone v. Ritter, 911 A.2d 362 (Del. 2006). [Slip opinion not obtained; cited for the proposition only, and nothing is quoted.] Strathern, M. (1997). “Improving ratings”: Audit in the British University system. European Review, 5(3), 305– 321. https://doi.org/10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4 [Anthropological essay; no data, and Strathern restates rather than originates the formulation.] Szalma, J. L., Hancock, P. A., Warm, J. S., Dember, W. N., & Parsons, K. S. (2006). Training for vigilance: Using predictive power to evaluate feedback effectiveness. Human Factors. https://doi.org/ 10.1518/001872006779166343 [Bibliographic record located only; content, volume, issue and pages were not verified, and no figure is quoted from it.] Tschandl, P., Rinner, C., Apalla, Z., et al. (2020). Human–computer collaboration for skin cancer recognition. Nature Medicine, 26(8), 1229–1234. https://doi.org/10.1038/s41591-020-0942-0 Human-on-the-Loop · WP No. 8 References 114 United States Air Force. (2009, May 18). Unmanned aircraft systems flight plan 2009–2047. Headquarters, United States Air Force. Defense Technical Information Center accession ADA505168. United States Department of Defense. (2023, January 25). Autonomy in weapon systems (DoD Directive 3000.09, reissuance cancelling the 2012 directive). Office of the Under Secretary of Defense for Policy. Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. https://doi.org/10.1038/ s41562-024-02024-1 van der Sijs, H., Aarts, J., Vulto, A., & Berg, M. (2006). Overriding of drug safety alerts in computerized physician order entry. Journal of the American Medical Informatics Association, 13(2), 138–147. https:// doi.org/10.1197/jamia.M1809 VURA Capital Innovation Holdings, Inc. (2026, June 29). AIใ‚จใƒผใ‚ธใ‚งใƒณใƒˆๆ™‚ไปฃใ€็ตŒๅ–ถ่€…ใฏใ€ŒHuman-on-the-Loop็ตŒ ๅ–ถใ€ใธ [In the age of AI agents, executives move to “Human-on-the-Loop management”] [Press release]. PR TIMES. https://prtimes.jp/main/html/rd/p/000000009.000183394.html Warm, J. S., Parasuraman, R., & Matthews, G. (2008). Vigilance requires hard mental work and is stressful. Human Factors, 50(3), 433–441. https://doi.org/10.1518/001872008X312152 [Narrative review; it reports no sample of its own.] Whiteford, P. (2021). Debt by design: The anatomy of a social policy fiasco – Or was it something worse? Australian Journal of Public Administration, 80(2), 340–360. https://doi.org/10.1111/1467-8500.12479 [Frequently miscited to the Australian Journal of Social Issues.] Wickens, C. D. (2018). Automation stages & levels, 20 years after. Journal of Cognitive Engineering and Decision Making, 12(1), 35–41. https://doi.org/10.1177/1555343417727438 Wickens, C. D., Clegg, B. A., Vieane, A. Z., & Sebok, A. L. (2015). Complacency and automation bias in the use of imperfect automation. Human Factors, 57(5), 728–739. https://doi.org/10.1177/0018720815581940 [Volume, issue and pages were not independently confirmed.] Willison, S. (2026, February 19). SWE-bench Verified: The February 2026 standardised re-run [Blog post]. https:// simonwillison.net/2026/Feb/19/swe-bench/ [Blog post reporting on a continuously updated leaderboard artifact; not peer-reviewed.] Wolfe, J. M., Horowitz, T. S., & Kenner, N. M. (2005). Rare items often missed in visual searches. Nature, 435(7041), 439–440. https://doi.org/10.1038/435439a Wolfe, J. M., Horowitz, T. S., Van Wert, M. J., Kenner, N. M., Place, S. S., & Kibbi, N. (2007). Low target prevalence is a stubborn source of errors in visual search tasks. Journal of Experimental Psychology: General, 136(4), 623–638. https://doi.org/10.1037/0096-3445.136.4.623 Xu, F. F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z. Z., Zhou, X., Guo, Z., Cao, M., Yang, M., Lu, H. Y., Martin, A., Su, Z., Maben, L., Mehta, R., Chi, W., Jang, L., Xie, Y., Zhou, S., & Neubig, G. (2025). TheAgentCompany: Benchmarking LLM agents on consequential real world tasks. In Advances in Neural Information Processing Systems: Datasets and Benchmarks Track (NeurIPS 2025). Preprint arXiv: 2412.14161. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). τ-bench: A benchmark for tool-agent-user interaction in real-world domains. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025). Preprint arXiv:2406.12045. Yuan, M., Zhou, Z., Xiong, X., et al. (2026). OSWorld 2.0: Benchmarking computer use agents on long-horizon real-world tasks (arXiv:2606.29537). https://arxiv.org/abs/2606.29537 [Preprint; not peer-reviewed.] Zhang, B., de Winter, J., Varotto, S., Happee, R., & Martens, M. (2019). Determinants of take-over time from automated driving: A meta-analysis of 129 studies. Transportation Research Part F: Traffic Psychology and Behaviour, 64, 285–307. https://doi.org/10.1016/j.trf.2019.04.020 Ziegelmeyer, A., Schmelz, K., & Ploner, M. (2012). Hidden costs of control: Four repetitions and an extension. Experimental Economics, 15(2), 323–340. https://doi.org/10.1007/s10683-011-9302-8 Japanese-language sources follow, ordered by the Japanese title under which each is cited in the text. AIไบ‹ๆฅญ่€…ใ‚ฌใ‚คใƒ‰ใƒฉใ‚คใƒณ๏ผˆ็ฌฌ1.2็‰ˆ๏ผ‰[AI Guidelines for Business, version 1.2]. (2026, March 31). ็ทๅ‹™็œใƒป็ตŒๆธˆ็”ฃๆฅญ ็œ [Ministry of Internal Affairs and Communications & Ministry of Economy, Trade and Industry]. [Nonbinding soft law.] Human-on-the-Loop · WP No. 8 References 115 ไผš็คพๆณ• [Companies Act], Act No. 86 of 2005 (Japan), Article 362(4)(vi) and Article 362(5). [Numbering taken from the secondary literature and not confirmed against e-Gov, which was not retrievable in the research pass.] ไผš็คพๆณ•ๆ–ฝ่กŒ่ฆๅ‰‡ [Ordinance for Enforcement of the Companies Act], Article 100 (Japan). [Numbering taken from the secondary literature and not confirmed against e-Gov, this provision having been renumbered over time.] ้‡‘่žๅบ [Financial Services Agency]. (2026, March). AIใƒ‡ใ‚ฃใ‚นใ‚ซใƒƒใ‚ทใƒงใƒณใƒšใƒผใƒ‘ใƒผ๏ผˆ็ฌฌ1.1็‰ˆ๏ผ‰— ้‡‘่žๅˆ†้‡ŽใซใŠใ‘ใ‚‹ AIใฎๅฅๅ…จใชๅˆฉๆดป็”จใฎไฟƒ้€ฒใซๅ‘ใ‘ใŸๅˆๆœŸ็š„ใช่ซ–็‚นๆ•ด็† [AI discussion paper, version 1.1: An initial mapping of the issues in promoting the sound use of AI in the financial sector]. [Explicitly non-binding; it sets no supervisory expectations.] ไบบๅทฅ็Ÿฅ่ƒฝ้–ข้€ฃๆŠ€่ก“ใฎ็ ”็ฉถ้–‹็™บๅŠใณๆดป็”จใฎๆŽจ้€ฒใซ้–ขใ™ใ‚‹ๆณ•ๅพ‹๏ผˆAIๆŽจ้€ฒๆณ•๏ผ‰[Act on the Promotion of Research, Development and Utilisation of AI-Related Technologies], ไปคๅ’Œ7ๅนดๆณ•ๅพ‹็ฌฌ53ๅท [Act No. 53 of 2025] (Japan), promulgated 4 June 2025 and fully in force 1 September 2025. [Twenty-eight articles and no penal provisions; the article numbering comes from practitioner analyses because e-Gov and the ่ก†่ญฐ้™ข text were not retrievable in the research pass.] ไบบๅทฅ็Ÿฅ่ƒฝๅŸบๆœฌ่จˆ็”ป๏ผˆ็ฌฌโ…กๆœŸ๏ผ‰[Second-Term AI Basic Plan], Cabinet decision of 14 July 2026 (Japan). [The substantive text on human supervision was not retrievable in the research pass, and nothing is attributed to it.] ๆ—ฅๆœฌใ‚ทใ‚นใƒ†ใƒ ๆŠ€่ก“ไบ‹ไปถ [Nihon System Gijutsu case], ๆœ€้ซ˜่ฃๅˆคๆ‰€็ฌฌไธ€ๅฐๆณ•ๅปท [Supreme Court of Japan, First Petty Bench], judgment of 9 July 2009, ๆฐ‘้›†63ๅทป6ๅท1442้ . [The formulation ใ€Œ้€šๅธธๆƒณๅฎšใ•ใ‚Œใ‚‹ใ€ is reported from the secondary literature and is not quoted from the ๆฐ‘้›† report.] Human-on-the-Loop · WP No. 8 References 116

© 2026 by Vura Capital Innovation Holdings Inc.
โ€‹ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ

bottom of page