top of page

Human on the Loop (HOTL): Advanced Human-AI Collaboration Processes

What is Human on the Loop (HOTL)? A New Model for Human-AI Collaboration

industry transformation on human on the loop

The evolution from generative AI to autonomous AI agents is bringing corporate management to a major turning point. Until now, many companies have operated under a structure in which humans execute work while systems and AI support them—a Human-in-the-Loop model, where humans remain continuously involved in business processes to make decisions and provide approvals.

However, in an era where AI agents can autonomously execute and coordinate activities across sales, marketing, analytics, customer service, software development, operations, and other functions, a management model in which humans individually approve and execute every decision will make it increasingly difficult to achieve both speed and scale.

VURA Capital Innovation defines this fundamental shift as the transition from Human-in-the-Loop to Human-on-the-Loop (HOTL) Management.

Human-on-the-Loop Management is a next-generation management model for the AI era in which AI autonomously executes tasks and operations, while humans define purpose, values, and capital allocation; oversee the system as a whole; and intervene only when necessary—controlling the system from outside the operational loop.

VURA Capital Innovation believes that this transition will become a core element of Enterprise Redefinition in the AI era—the redesign of business models and organizations in response to technological evolution.

Why Human-on-the-Loop Management, and Why Now?

AI agents are now evolving toward Closed-loop AI—systems that autonomously observe the outcomes of their own actions, identify issues, improve their approach, and execute again.

Traditional AI adoption has largely followed a Human-in-the-Loop model, with humans serving as the central hub for input, evaluation, and execution. Today, however, autonomous loops in which multiple AI agents collaborate to complete entire workflows are becoming practical realities.

Examples include AI coding agents autonomously identifying and fixing bugs, AI sales agents continuously optimizing proposals based on outcomes, and AI systems monitoring infrastructure and automatically recovering from failures.

When autonomous cycles operate at this extraordinary speed, a management model that requires humans to remain involved in every decision loop risks turning humans themselves into the organization's greatest bottleneck.

In the AI era, the role of management is no longer primarily to make individual operational decisions. It is to define purpose, values, and capital allocation, and to design and govern the overall system composed of both humans and AI.

Executives must therefore evolve from being individual decision-makers within the operational loop to becoming architects and overseers of the overall system in which AI agents autonomously create value.

Executives Must Move from Inside the Loop to Above the Loop

The Role of Management Is Shifting from Execution to Governance

VURA Capital Innovation views the evolution of industry and management models as follows.

Industrial Age | Human-in-the-Process

Humans as “Executors”

Since the Industrial Revolution, humans have been integral to the process itself: humans execute the work, and humans manage the work.

Digital Age | Human-in-the-Loop

Humans as the “Core of the Workflow”

With digitalization and the early adoption of AI, systems generate recommendations while humans remain responsible for approval and decision-making.

AI Agent Age | Human-on-the-Loop

Humans as “Governors”

AI agents autonomously execute and coordinate work, while humans define purpose, values, and capital allocation, overseeing the system from outside the operational loop and intervening when necessary.

The Evolution of Industry and Management Models

In Human-on-the-Loop Management, the essential roles of humans can be summarized into five areas:

1. Purpose — Defining Purpose
Determining the direction the company should pursue and the future it intends to create.

2. Values — Defining Values
Establishing the ethical principles and corporate values that guide AI behavior—creating the guardrails that prevent undesirable or uncontrolled AI actions.

3. Capital Allocation
Allocating financial, human, technological, and other resources to the areas where they can generate the greatest value and return.

4. Governance — Oversight and Governance
Establishing mechanisms to ensure that both AI and humans operate effectively, responsibly, and safely.

5. Exception Intervention
Making the final human decision when unexpected situations, exceptional circumstances, or significant risks arise.

A Theory of Oversight Capacity, Its Failure Modes, and the Non-Delegable Residual in the Age of AI Agents

​

You did not decide to be on the loop. Speed put you there.

When AI agents start executing decisions, management says: "A human is watching in the end, so it is fine." This paper asks when that sentence is true. Its finding: in many organizations, it no longer is.

VURA Working Paper No.8. This page explains the core of the theory. The full paper (English original, 116 pages; Japanese edition, 111 pages) is linked at the end. A Japanese book edition, Naze, AI ējento jidai, hito no gabanansu ga hatan suru no ka? ("Why Human Governance Breaks Down in the Age of AI Agents"), was written from it.

This paper in three minutes

  • Oversight is a resource, not a role. A day's attention is finite and is spent as it is used. Whether there is enough can be calculated.

  • When it runs out, people do not say "I cannot review this." They approve. A high approval rate is not evidence of safety but a symptom of insolvency.

  • Humans fail to monitor for three reasons. Reaction lag (latency); one person can watch only 4 to 6 things at once (capacity); and the better the AI gets, the rarer its failures, and the rarer a thing is, the more often it is missed (rarity).

  • "Let AI watch the AI" does not close the loop. What is needed is not more capability but errors that are independent of the generator's. Independence cannot be bought with intelligence.

  • The management question is not "are we on the loop?" It is "which decisions have we moved to the side where cost stays flat as volume grows?"

  • This paper falsifies its own company's proposal. Of the "five roles that remain human" VURA published, two are not part of the non-delegable residual.

The strongest claim: oversight is not a role but a resource you use up

"Human-on-the-Loop" is usually described as a posture. Stand outside the loop, intervene when necessary, and you are safe.

This paper rewrites it as an authority structure and a finite resource, not a posture.

An analogy: a bank account. Oversight is not a title but a balance. Every case reviewed is a debit. When the balance is gone, the title does not let you pay.

The solvency condition ρ ≤ 1 — one division tells you whether you have enough

Oversight takes time. Each exception needs dialogue time IT (time spent engaging with the case) plus an associated cost c (switching, documentation). What the supervisor has left after all other duties is unbound attention A. If exceptions arrive at rate λ (lambda), the condition for oversight to hold fits in one inequality.

Solvency condition ρ = λ(IT + c) / A ≤ 1

What is divided by what. The numerator, λ(IT + c), is the time needed to review every exception of the day properly (cases × time per case). The denominator, A, is the time actually available. So ρ (rho) = time needed ÷ time available. At 1 or below, you keep up. Above 1, you do not. The larger the number, the deeper the breakdown: ρ = 3 means you need three times the time you have.

The span constraint — how many agents one person can watch

The number of agents N one person can oversee also has a ceiling. Neglect time NT (how long an agent can safely be left alone) divided by the sum of dialogue time IT and resumption delay WT (the lag before a returning human is actually working the case) bounds the effective supervisory span — the number of agents one person can genuinely look after.

Span constraint N ≤ NT / (IT + WT) + 1

Read it as the time you can leave things alone, divided by the effort one case costs. Agents that need less handling raise N; long IT and WT lower it.

The important point is that A only ever appears by subtraction. Start from 480 minutes. Subtract the supervisor's own duties, then meetings, then travel and interruptions. What remains is A. The book edition draws this subtraction and says: "A result where A is close to zero is not a calculation error. It is the answer."

Insolvency shows up as approval, not refusal

This is the most counter-intuitive part. An analogy: a desk piled with documents awaiting sign-off. When the pile gets tall, people do not say "I can no longer read these." They stamp faster. From the outside, that decline is indistinguishable from careful reading of every page. Throughput holds; only the substance drains away.

Proposition 5 | Insolvency manifests first as quality loss, not as throughput loss

In plain terms — When oversight breaks, the case count does not fall. What falls is the substance of each case.

The traces are everywhere in clinical alert systems.

  • Of 157,483 alerts, 52.6% were overridden (53% of those overrides were appropriate; the paper presents the evidence in both directions)

  • Acceptance of drug alerts is under 1%. Each additional alert cuts acceptance by 30% (IRR 0.70), and someone who has overridden an alert once overrides it again with 87.9% / 99.9% probability

  • Intensive care: 187 alarms per bed per day; 88.8% of arrhythmia alarms are false positives

  • Override rates were 49–96% in a 2006 survey and still 46.2–96.2% in 2020. Fourteen years of improvement efforts barely moved the range

A high approval rate is not evidence of safety. It is a symptom of insolvency.

The starting point: falsifying our own proposal

This paper has an origin ordinary management papers lack. On 29 June 2026 VURA itself issued a press release, "In the age of AI agents, executives move to Human-on-the-Loop management," proposing five roles that remain human: purpose, values, capital allocation, governance, and exception intervention.

Written up as "the good five," they failed against all four evidence bases: human factors engineering, human–computer interaction, agent evaluation, and administrative records. So the theory was rebuilt from the counter-evidence.

The paper concludes that two of the five roles the company proposed — governance and exception intervention — are not part of the non-delegable residual. That is not self-criticism. It is method.

The conflict of interest is disclosed as a structure in Section 11.7: the party proposing how oversight should be measured stands to profit from that measurement.

The term comes from defense doctrine, not academia

A second overlooked point is the term's origin. Its technical content corresponds to Sheridan & Verplank's (1978) level 6 of automation: "the human is given a limited time to veto before automatic execution." But the phrase "on the loop" appears in neither management science nor human factors. It comes from the US Air Force UAS Flight Plan (2009), which glossed "man on the loop" as supervisory control, and Human Rights Watch's Losing Humanity (2012), which fixed the three-way classification.

And US Department of Defense Directive 3000.09 (2023 revision) does not adopt this vocabulary — it writes "operator-supervised autonomous weapon system."

A management paper must not claim doctrinal backing for the term. This paper uses it as the name of an authority structure, not as an authority.

Three reasons humans cannot monitor

Latency — response-time SLAs measure the wrong quantity

A meta-analysis of 129 studies and 520 observations puts mean time from takeover request to action at 2.72 seconds (range 0.69–19.79). The interesting finding: the larger the time budget, the slower the reaction — by 1.35 seconds. On real roads, manual control returns in 2.4 seconds, but gaze behavior does not begin to recover until 6.1 seconds. The hands are back; the looking is not.

The decisive report: takeover time and the quality of the avoidance maneuver were nearly independent. Fast responders did not cope better. Putting "seconds to respond" on a monitoring dashboard is measuring a quantity that should not be measured.

Capacity — the effective span is 4 to 6

In measured supervisory spans, dialogue time per case is roughly 15–17 seconds and the effective span is 4 to 6. At eight agents performance collapses and robot losses rise twelvefold.

Worse, automating the allocation of attention did not improve performance. Having a machine tell the operator which agent to look at next did not work. Compliance exceeded 95% and subjective workload fell, yet no effect on performance was detected (F(2,33)=0.50, p=0.609), and 67% of participants disliked the aid.

Rarity — the smarter the AI, the less humans notice its failures

In the classic visual-search studies, as target prevalence falls from 50% to 10% to 1%, the miss rate rises from 7% to 16% to 30%, reaching 52% when targets are extremely rare. What matters is why. Sensitivity d′ (the ability to discriminate) did not get worse (1.97 → 2.55). What moved was the decision criterion c, from 0.13 to 1.14 — the height of the bar for saying "it is there."

People did not stop seeing. They became readier to say "nothing there." Manipulating rewards did not help; telling people to slow down did not help. The only intervention that worked combined high-prevalence bursts (practice blocks with targets deliberately made frequent) with full feedback. The paper treats this as weak evidence — the experiment had N=14 and was not significant (p=.312), and the 46% → 21% improvement is a between-experiment comparison.

Proposition 7 | As failures become rarer, the probability of detection falls

In plain terms — The AI's improvement itself breaks oversight. The paper calls this the oversight scissors. The upper blade is rising AI performance; the lower blade is falling human detection.

On average, oversight does not work. But the conditions under which it works are known

A 2024 meta-analysis reports a synergy effect of g = −0.23 for human–AI combinations: the combination loses to whichever of human-alone or AI-alone is better. Augmentation — simply adding AI to a human — gives g = +0.64. Decision tasks: −0.27. Generation tasks: +0.19. And where the AI is stronger than the human: −0.54.

Neither explanations nor confidence displays are significant moderators. The hope that "if the AI gives its reasons, the human can judge" is not supported. Publication bias is detected on the optimistic side only (Egger β=1.96, p=0.002). A separate meta-analysis of 163 studies and N=82,078 shows how context-dependent the picture is: algorithm aversion d = −0.50, algorithm appreciation d = +0.27.

So what was different in the cases that worked? None of them was an override design. A human marking the AI's answer right or wrong is the form that does not work.

  • In the child-protection hotline case, the human held an independent information channel the model did not have

  • In a randomized trial of breast-cancer screening (80,033 women), detection stayed at 6.0 versus 5.0 per 1,000 while reading workload fell 44%. That is a triage design — the AI sorts, and humans concentrate on what needs seeing — not a human stamping "approved" on outputs

  • Another study shows that none of human, AI, or human-plus-AI dominates

Here the paper weakens its own proposition. It originally set "an independent information channel" as the condition. That was falsified by a report in which crowd workers reading the same inputs outperformed the AI alone as a team (0.89 versus 0.84). The condition was generalized to error decorrelation — humans and AI not being wrong in the same places in the same way. An independent channel is the most auditable source of decorrelation, not the condition itself.

"Let AI watch the AI" does not close the loop

Then should oversight be handed to AI? This is the paper's deepest layer (Propositions 10 and 11).

  • In systematic tests of self-correction, every condition was flat or worse (GPT-4 on GSM8K: 95.5 → 89.0)

  • Under self-critique, accuracy fell from 5% to 3% and from 16% to 2%. With a sound external verifier, it rose to 38% and 37%. With an LLM as verifier, the false-negative rate was 95.8% — most errors slipped through

  • In a taxonomy of multi-agent failures, failure rates ran 41–86.7%, and verification failures accounted for 23.5% (incorrect verification 9.1%). The fixes that worked were organizational design, not better models (assigning final decision authority: +9.4%; adding a goal-verification stage: +15.6%)

  • In the second phase of an autonomous-operation experiment, a "boss" agent on the same base model approved versus rejected at roughly 8:1. The developer's own explanation: they "share the same blind spots"

An analogy: proofreading your own manuscript. Typos survive because the assumptions you had when writing are the ones you have when reading.

Here, though, the paper breaks its own argument and rebuilds it. The first draft defined a "closed loop" as self-observation without external signal. Observations from the environment are external signals, so that definition was self-contradictory. Real agents read test suites and database states. So the question moves from "is there feedback?" to "is the feedback path independent of the generator's error distribution?" Tests an agent wrote for itself are not independent (correlated blind spots). A compiler is.

Central thesis Humans are needed not because their judgment is better. A closed loop requires a verifier that is statistically independent of the generator, and independence is not supplied by improvements in capability. But — oversight value = independence × detection probability. Capability gains preserve the first term and erode the second (through rarity). Independence is necessary, but not sufficient.

In plain terms — The human's contribution is not intelligence but being wrong differently from the AI. That alone is not enough: actually noticing is required at the same time. It is a product, so if either factor is zero, the whole is zero.

The scaling dichotomy — the management question changes

From here the paper splits the five roles by how their cost grows. One axis: what happens to cost when throughput T (cases handled per day) rises tenfold, a hundredfold.

  • Constitutive oversight — setting the frame in advance (purpose, values, capital allocation). Not incurred per case, so cost does not change with volume. Written O(1) (flat)

  • Controlling oversight — checking after the fact (approval, sampling inspection, exception handling). Incurred per case, so it grows with volume

RoleTypeCost against throughput TNon-delegable?

PurposeConstitutive oversightO(1) (flat)Non-delegable (logical grounds)

ValuesConstitutive oversightO(1) (flat)Non-delegable (logical grounds)

Capital allocationConstitutive oversightO(1) (flat)Non-delegable (logical grounds)

GovernanceControlling oversightGrowsNot non-delegable

Exception interventionControlling oversightGrowsNot non-delegable

Proposition 12 | Controlling oversight cannot be bounded under a fixed detection guarantee A fixed-size sampling inspection looks like O(1) cost, but coverage n/T thins out as O(1/T), and undetected exposure p(T−n) grows as Ω(T). Further, falling prevalence within the sample degrades the detection probability itself.

In plain terms — The promise "we sample 100 cases a day" is emptied the moment volume rises. What you can keep shrinks; what you cannot keep swells.

How to read the symbols.

  • Coverage n/T — the share of all T cases inspected. Hold n at 100 while T goes from 1,000 to 100,000 and 10% becomes 0.1%

  • O(1/T) — "shrinks in inverse proportion to T." As T grows, coverage approaches zero

  • Undetected exposure p(T−n) — uninspected cases multiplied by the error rate p: the volume of errors that get out unseen

  • Ω(T) — "grows at least in proportion to T." It does not level off

Supervisory effectiveness η (eta) = R · s · d(π). R is the routing rate (the share of cases the AI sends to a human), s the sampling rate (the share of routed cases actually inspected), d(π) the detection probability (the chance of noticing an error at prevalence π). All three are at most 1, so the product is always smaller than where you started. As throughput grows, all three fall together; if any one approaches zero, η approaches zero.

The management question is not "are we on the loop?" but "what have we moved into the O(1) class?"

The non-delegable residual rests on three different grounds

The non-delegable residual is what stays with humans no matter how much is delegated to AI.

GroundKindRelation to the five roles

Independence of verificationStructuralOutside the five roles. The property that makes a supervisor effective

Constitutive settings (purpose, values, capital allocation)LogicalThree of the five roles

Fiduciary dutyLegalAbove the five roles. Does not disappear whoever is appointed

Fiduciary duty is the duty owed by someone entrusted with another party's assets or interests.

"Not non-delegable" does not mean "safe to hand to the machine." It means that under throughput, humans cannot do it. That is exactly why those two rows — and only those two — require measurement.

What binds today is not AI regulation but company law

Waiting for AI regulation is not an option. The EU AI Act, Article 14(4)(b), names automation bias (the tendency to accept machine output uncritically), and the n ≥ 2 requirement in Article 14(5) (verification by two people) is the only quantitative requirement in any normative document the paper surveyed. However, Regulation (EU) 2026/1744 (the Digital Omnibus, adopted 8 July 2026, in force 27 July 2026) postponed the Annex III high-risk obligations to 2 December 2027. As of August 2026, Article 14 is not yet applicable.

In Japan, the AI Promotion Act (Act No. 53 of 2025) carries no penalties. The AI Business Operator Guidelines v1.2 (31 March 2026) name automation bias, but the verb is "consider," not "implement." The Financial Services Agency's AI discussion paper v1.1 observes practice this way:

"Because a ringi-sho (internal approval document) can materially affect customers, for example in lending decisions, the great majority of operations require human involvement (Human in the loop)."

Standards such as ISO/IEC 42001 and the NIST AI RMF are entirely procedural, with no quantitative requirements at all.

Even so, something binds today: the directors' duty of oversight. In Delaware, Marchand v. Barnhill (2019) and In re Boeing (2021, a US$237.5 million settlement) established the logic that claiming to conduct oversight you are not conducting is itself a breach. In Japan, the counterpart is Article 362(4)(vi) of the Companies Act and the "ordinarily foreseeable" standard of the Japan System Technology case (Minshū vol. 63, no. 6, p. 1442).

And this is the paper's sharpest legal contribution. By naming automation bias, the published guidelines themselves construct the "foreseeability" on which liability turns.

No penalty does not mean no liability.

So what do you measure, and what do you disclose?

The paper proposes an eleven-item disclosure protocol: Tier 1 (declarative), eight items; Tier 2 (continuous), three. Tier 2:

  1. Synthetic exception injection rate and capture rate — are supervisors catching the exceptions deliberately planted among the cases?

  2. Return-to-unassisted-operation schedule — when, and how much, do you train by running without AI assistance?

  3. A residual uniform sample of the non-routed stream, and the routing recall it estimates — randomly sample the cases the AI did not send to a human, and measure what should have been routed but was not

The book edition breaks this into three questions a board can use tomorrow.

  1. How many minutes of unbound attention A per day does the person responsible for oversight have left after other duties?

  2. As the sampling denominator has grown, how has supervisory effectiveness moved?

  3. Of the decisions delegated to AI, how many zero rules have been built into the action space in advance, making the action impossible? (A zero rule is a constraint that removes an option from the start rather than stopping it afterward.)

The third question matters most. The first two diagnose the present; only the third measures progress toward constitutive oversight.

What the paper admits it cannot yet show

This working paper has not been reviewed by an external academic journal. It went through three rounds of internal review and revision, to version 1.3. Because it has not passed independent third-party review, the basis for trusting it lies not in its provenance but in the fact that all 13 propositions carry a falsification condition.

The weaknesses it discloses:

  • Some numbers were deliberately not used. The widely circulated "23 minutes 15 seconds to resume after an interruption" is not in the original source (the correct figure, under a different condition, is 25 minutes 26 seconds), and a frequently cited utilization ceiling could not be verified, so no number from it is cited. "1.9–25.7 seconds" and "a 93% error rate" could not be traced to primary sources and are not used. The full text of ISO/IEC 42001 could not be read, so nothing from it is placed in quotation marks

  • NT in the span constraint is in principle unmeasurable under Proposition 9, and the paper says so. An unmeasurable quantity remains in the formula

  • Two of the nominal-oversight cases are "removal of oversight" and are explicitly n = 1 cases

  • The paper criticizes its own proposal (Section 11.8). The classic critique that a metric corrupts once it becomes a target applies to its own eleven items

  • The disclosure paradox (Section 11.9) — honestly disclosing solvency creates a record of one's own breach of the duty of oversight. The more honest, the more exposed. The cost of implementing the protocol is also not accounted for

  • An unresolved reviewer comment on the internal consistency of Appendix Table A2 remains (a fix would require inventing numbers, so it is held)

  • Verification against primary Japanese legal sources (e-Gov, Minshū), the EU Official Journal, and Delaware case records is in progress

Implications for executives and directors

For executives. Check whether you report a high approval rate as evidence of safety. It may be a symptom of insolvency. Measure instead unbound attention A (the time actually available per day) and supervisory effectiveness η (routing rate × sampling rate × detection probability).

For boards. The sentence "a human gives final confirmation" is, as of 2026, already a legal risk factor. If claiming oversight you do not conduct constitutes a breach, what needs checking is not the org chart but how many seconds per case you can actually spend.

For those responsible for AI adoption. The goal is not more monitoring but moving controlling oversight into constitutive oversight. Instead of adding approvals after the fact, build zero rules into the action space in advance. And choose triage designs, not override designs. Those two things.

Human on the Loop Working Paper

As enterprises deploy autonomous AI agents, Human-on-the-Loop (HOTL) is touted as the ideal governance model. This paper proves that real-time human monitoring naturally breaks down as AI scales, establishing 4 core insights:

  • Physical Feasibility, Not Policy: Loop position is set by execution speed versus human response time; high throughput physically forces oversight out-of-the-loop.

  • Exhaustible Resource: Human attention is finite. Overloaded capacity leads to "insolvent" oversight, failing through silent rubber-stamping.

  • Low-Prevalence Decay & Blind Spots: Accurate AI makes human error detection harder, while supervisor AIs share identical blind spots. The human's true value is statistical verification independence.

  • The Scaling Dichotomy: Real-time monitoring costs in attention, whereas Constitutive Oversight (ex-ante rule-setting) costs —making it the only sustainable model at scale.

Human on the Loop Working Paper

No. 8

bottom of page