top of page

Ageless Management

Ageless Management Cover

This paper develops Ageless Management, a theory of how organizations can redesign the value of human experience in the age of generative AI and longer working lives. It argues that as AI reduces the marginal cost of tasks associated with fluid-intelligence-type processing, the relative value of accumulated domain knowledge, contextual judgment, metacognition, and experience-based pattern recognition may increase. The paper proposes that the value of experience may shift from a production premium toward an audit premium: experienced individuals may create value by detecting errors, contextual inconsistencies, and hidden risks in AI-generated outputs. It integrates three mechanisms—cognitive complementarity between generations, the audit value of experience, and the lifelong extension of Brain Capital—and introduces the concept of generational decorrelation.

Rather than claiming that age itself creates superior judgment, the paper formulates these mechanisms as falsifiable hypotheses and proposes an empirical research agenda. Ageless Management ultimately reframes demographic aging as a question of how organizations can combine AI and diverse human cognitive capabilities to improve judgment, oversight, and long-term value creation.

Ageless Management

No. 9

V U R A W O R K I N G P A P E R S E R I E S — N o . 9

Ageless Management

A Theory of Cognitive Complementarity, the Audit Value of Experience,

and the Lifelong Extension of Brain Capital in the Age of AI

ใ‚จใƒผใ‚ธใƒฌใ‚น็ตŒๅ–ถ โ€• AIๆ™‚ไปฃใซใŠใ‘ใ‚‹่ช็Ÿฅ็š„็›ธ่ฃœๆ€งใ€็ตŒ้จ“ใฎ็›ฃๆŸปไพกๅ€คใ€ใใ—ใฆ่„ณ่ณ‡ๆœฌใฎ็”Ÿๆถฏๆ‹กๅผตใฎ็†่ซ–

Naoki Kadowaki

VURA Capital Innovation Holdings, Inc.

August 2026

Version 1.3 (August 2026). Revised in response to three rounds of internal review. This working paper is an

unrefereed draft; comments are welcome. Contact: VURA Capital Innovation Holdings, Inc. The views

expressed are the author's own and do not represent the official position of the affiliated organization.

Abstract

This paper presents a theory of Ageless Management: a management regime that removes the

variable of chronological age from decisions about role allocation, evaluation, participation,

and exit, and instead allocates roles dynamically on the basis of measured cognitive

characteristics, accumulated domain experience, health status, and the individual's own

volition. Whereas conventional senior-employment and diversity initiatives have rested on a

paradigm of accommodation and compensation, this paper envisions a management model

that, through bidirectional complementarity between heterogeneous cognitive abilities

mediated by AI, converts multigenerational and diverse talent into co-creating agents of a

managerial resource: brain capital. The paper's principal axis is the axis of age and experience

linking super-seniors (aged 60 to their 90s) and youth; the extension to socially marginalized

groups is positioned as a secondary context of generalization, applicable only insofar as the

logic of the cognitive bottleneck and AI complementation applies. The theory's point of

departure is the following asymmetry. Generative AI lowers the marginal cost of fluid

intelligence (Gf)-type tasks and raises the relative marginal value of crystallized intelligence

(Gc)-type tasks (contextual interpretation and interpersonal judgment, and above all

evaluative and audit-type tasks)—although this rise is conditional on verification actually

detecting errors. Under this shift, the "cognitive bottleneck"—whereby difficulty in performing

the Gf-type components of a job bundle has forced exit from the job as a whole—can be

removed. At the same time, the supervision of AI output retains a residual that cannot be

delegated to AI (Human on the Loop: HOTL), and the supply side of supervision as a scarce

resource comes into question. This paper conceptualizes this scarce input operationally as

"experiential audit capacity"—the capacity to detect contextual errors, practical risks, and

ethical risks in AI output—measured by detection performance on verification tasks. The

claim that this capacity is formed not by chronological age but by the interaction of domainspecific

knowledge and operational schemata (the products of long-term domain experience)

with generalized crystallized intelligence and metacognition is placed not in the definition but

as a refutable proposition, and constitutes the core of the theory. This formation has boundary

conditions: when AI output conforms to the auditor's own received views, the synergy of

confirmation bias and automation bias means experience can instead impede detection, and

in domains where technological change is rapid and the half-life of domain knowledge is

short, the audit effectiveness of experience can decline. Further, the paper formalizes

"generational decorrelation"—the tendency of supervisor pools that have lived through

different era environments and technological generations to be less likely to share blind spots

in judgment—as an organizational supply source of oversight value. Its realization is explicitly

conditioned on three requirements: a channel condition, that auditors have unmediated

access to raw output; an organizational condition, that power gradients are flattened so that

objections reach decision-making; and a model condition, that the AI output under audit is not

monopolized by a single foundation model and cross-verification by models from different

developers is used in parallel. The claims are consistently conditional. The three social

problems of population aging, constrained youth participation, and the exclusion of socially

marginalized groups convert into untapped supply sources of brain capital (stock K ×

utilization rate u) only when the conditions of AI-mediated complementarity and Brain Safety

Working Paper | Ageless Management in the AI Era 2

are met; the conversion is not automatic. Moreover, the outcomes of multigenerational staffing

are evaluated not solely by the quality of ideas but as net benefit—including the avoidance of

excessive risk and the reduction of rework, and net of the cost of decision time required for

verification. The paper presents eight definitions; twelve propositions, each with a refutation

condition—including the superiority of dynamic role allocation over fixed, chronological-agebased

role allocation (Proposition 12)—and three testable hypotheses (brain health,

multigenerational team outcomes, and the non-compressibility of experience). Most of the

propositions are untested: this paper is not a report of established facts but a proposal of

theory equipped with a map of verification.

Keywords: Ageless Management, crystallized intelligence, fluid intelligence, cognitive complementarity,

experiential audit capacity, generational decorrelation, brain capital, multigenerational ecosystem,

Human on the Loop, Brain Safety, healthy life expectancy, dynamic role allocation

JEL classification: J14, J24, J26, M12, M14, M54, O33, I31

1. Introduction

1.1 Demographic Structure and "the Domain of Longevity"

Contemporary corporate management stands in the midst of a demographic transformation

without precedent in human history. Rising life expectancy and declining fertility are trends

common to the advanced economies, and employment systems premised on the linear threestage

sequence of education, work, and retirement carry over a design conceived for an era

when life expectancy was substantially shorter than it is today. International policy

frameworks have made this transition an explicit theme: in December 2020 the United

Nations General Assembly declared the start of the Decade of Healthy Ageing (2021–2030),

designating as action areas the development of environments that support older people's

functional ability and the combating of ageism (WHO 2020; international organization

document). An extension of life expectancy does not in itself mean an extension of healthy

life expectancy, but the continuing expansion of "the domain of longevity"—the segment of

life once designated as post-retirement—cannot be ignored as a change in the preconditions

of management.

The three-stage model is embedded in firms not as mere custom but as a bundle of

institutions. Simultaneous mass hiring of new graduates designs the transition from

education to work as a one-time gateway; seniority-based treatment uses years of tenure as a

proxy variable for ability; and mandatory retirement age (teinen) systems enforce the

transition from work to retirement uniformly by chronological age. These institutions all

share the implicit assumption that age is a sufficiently good predictor of ability and role.

Public discussion of multi-track, multi-stage lives has spread, but the work of reformulating

that demand as a theory of corporate role allocation—at the level of who is allocated what,

Working Paper | Ageless Management in the AI Era 3

under which contractual form, and with what protections—has not yet been adequately

carried out. This paper takes that gap as its subject.

Moreover, this implicit assumption is itself inconsistent with the findings of cognitive

science. Cognitive aging is not a uniform decline but is asynchronous across abilities.

According to cross-sectional data from large web samples, abilities such as processing speed

peak in early adulthood, whereas the peak of accumulation-based abilities such as

vocabulary is observed in the late 60s to early 70s, and there is substantial heterogeneity in

the peak ages of cognitive abilities (Hartshorne & Germine 2015; the cross-sectional design

and the inclusion of cohort effects are discussed in detail in Section 2). If, at a given age,

abilities that decline coexist with abilities that are preserved or still growing, then an

institution that allocates roles by the single variable of chronological age is discarding too

much information. And it is this paper's contention that the discarded information contains

precisely those cognitive assets that remain unrecovered as managerial resources.

This change should not be discussed solely as a question of the quantity of labor supply.

This paper's concern is a question of quality: the cognitive assets accumulated in the

expanded "domain of longevity"—long-term domain experience, crystallized intelligence (Gc),

and metacognition—are not reaching work under the current mechanisms of role allocation.

Mandatory retirement systems and age-based treatment allocate roles by the proxy variable

of chronological age rather than by measured individual cognitive characteristics. As

discussed below, the distributions of cognitive characteristics across age groups overlap

substantially, and individual differences can exceed age differences. Uniform allocation by

chronological age can be rational only when the cost of measuring individual characteristics

is high and no alternative allocation criterion exists. The diffusion of generative AI is

undermining this very premise—that is this paper's point of departure.

One thing, however, should be made clear at the outset. This paper does not claim that

older people do not decline. The average age-related decline in fluid intelligence (Gf) is a

robust finding of cognitive science, and this paper accepts it as a premise (Section 2). Nor does

the paper assert that working makes people healthy. The relationship between work and

brain health faces difficulties of causal identification, and this paper treats it as a hypothesis

to be tested (H1) (Sections 3 and 9). What this paper interrogates is not the fact of aging itself

but a design problem of the allocation mechanism: what aspect of aging forces exit from

roles, and whether that forcing factor can be removed by AI.

The reason for setting the level of analysis at corporate management (micro) should also be

stated. Responses to population aging have so far been discussed mainly at the level of

institutional policy on pensions and mandatory retirement (macro) and at the level of

individual health and reskilling (individual). But the place where cognitive assets actually

meet tasks is the organization. Who is allocated which role, to which tasks AI is assigned, how

supervision is designed, and what contractual forms and protections participation is given—

each of these is a managerial decision, and institutional policy merely supplies the constraints

Working Paper | Ageless Management in the AI Era 4

outside them. Even if macro institutions change, unless the firm's theory of role allocation

changes, the cognitive assets of "the domain of longevity" will not reach work. Conversely,

once firm-level design is established, it can supply concrete requirement specifications to

institutional reform debates (Section 7).

1.2 The Limits of Conventional Senior-Employment and D&I Models

The senior re-employment and diversity and inclusion (D&I) initiatives undertaken at many

firms have, despite the legitimacy of their ideals, carried a structure that connects poorly to

value creation. Conventional senior employment has typically been framed in the context of

legal compliance and corporate social responsibility, and has tended to remain within a

compensatory, protective mindset—"offset age-related decline and assign routine work"—a

mindset of restoring a minus to zero. In this configuration, the individuals concerned feel

they are not expected to contribute, and the organization in turn perceives rigidifying labor

costs. The paradigm of accommodation and compensation rests on an accounting that books

its targets as costs, and to that extent it is not sustainable.

This impasse appears at three levels. At the level of motivation, the shrinking of roles saps

the individual's cognitive engagement and reproduces a self-perception of being written off.

At the financial level, continued employment unconnected to value creation is treated as a

cost center and becomes a target of cuts with every business cycle. At the strategic level, both

senior employment and D&I are marginalized as initiatives outside the core business and

never become design variables of the business model. It has also been pointed out that D&I

initiatives have fallen into an isomorphic impasse. So long as the targets are framed as beings

to be protected, the success of the initiative depends on the continuation of goodwill and

budget, not on the success of the business. What is needed is not an additional layer of moral

persuasion but the identification of pathways by which diverse cognitive assets connect to

outcomes, and the design of the conditions under which those pathways function.

This paper locates two structural causes at the root of the impasse. The first cause is that job

design has failed to resolve the "cognitive bottleneck." A job is not a demand for a single

ability but a bundle of Gf-type components (processing speed, working memory, learning of

novel procedures) and Gc-type components (accumulated knowledge, contextual

interpretation, interpersonal judgment). When performance of the Gf-type components

becomes difficult, exit from the job as a whole is forced even if high levels of Gc-type ability

are retained. Abilities that are possessed do not reach work—this is what this paper calls the

cognitive bottleneck (formally, Definition 3 in Section 5), and conventional routine-work reemployment

has not removed this bottleneck but merely circumvented the problem by

moving people to low-load jobs that do not touch it. When generative AI sharply lowers the

marginal cost of Gf-type components, this bottleneck itself becomes removable—this is one of

the central theoretical claims of this paper (previewing Propositions 1–2).

Working Paper | Ageless Management in the AI Era 5

The second cause is that the unconditional claim that diversity generates value in itself has

not been empirically supported. Facing this point squarely, without papering over it, is a

condition of the integrity of this paper's argument. Meta-analytic results on the relationship

between age diversity and team outcomes are consistent in hovering near a zero average

effect. Joshi & Roh (2009), in a meta-analysis of 39 studies and 8,757 teams, estimated the

direct effect of age diversity on performance at r = −.06, and Schneid, Isidor, Steinmetz &

Kabst (2016), in a quantitative review of 74 studies, concluded that the overall relationship

between age diversity and team outcomes is nonsignificant (the sole exception being

turnover, r = .11, whose direction is if anything unfavorable). In the most recent and largest

registered-report meta-analysis, Wallrich et al. (2024) (615 reports, 2,638 effect sizes), the

effect of demographic diversity as a whole was r = .014, effectively zero.

Yet the same literature also identifies the conditions under which effects appear. Backes-

Gellner & Veen (2013), using large-scale linked German data, reported that age diversity has a

positive effect on firm productivity when, and only when, the firm is engaged in creative

rather than routine tasks. Wegge et al. (2012), from a research program covering three

industries and more than 745 teams, identified high task complexity, a climate low in age

discrimination, and positive appraisal of diversity as success conditions, and Wallrich et al.

(2024) likewise confirmed that the diversity–performance relationship becomes more positive

when task complexity and dependence on creative divergence are high. In other words, what

the evidence has rejected is the unconditional celebration of diversity, not conditional

complementarity. This paper theorizes the content of those conditions as the state in which AI

substitutes for and complements Gf-type components so that heterogeneous cognitive assets

become connectable across individuals—AI-mediated complementarity—and formalizes it as

a moderator of diversity's performance effect (previewing Proposition 6). The empirical

headwind against diversity is, for this paper, not a refutation but a demand to identify the

missing mediating mechanism.

On the other hand, evidence exists suggesting that homogeneity has its costs as well. In

mock-jury experiments, diverse groups have been reported to exchange a wider range of

information than homogeneous groups, with majority-group members themselves making

fewer factual errors (Sommers 2006). This, however, is laboratory research manipulating

racial diversity, and replication with age diversity has not been confirmed—this paper treats

it strictly as indirect evidence. The reason this point gains weight in the age of AI is that the

humans supervising AI output tend to be composed of homogeneous groups sharing the same

training and information environment as the AI's own era. Supervisors who learned from the

same teaching materials and grew up in the same technological environment are also likely

to share the errors they overlook. Given WP8's formalization that the value of a supervision

channel depends on the decorrelation of errors (Kadowaki 2026h), people who have lived

through different era environments, technological generations, and failure cases can carry

managerial significance not merely as objects of fairness considerations but as supply sources

Working Paper | Ageless Management in the AI Era 6

of supervisory independence. This is the preview of what this paper calls generational

decorrelation, and whether it holds is a matter for testing (Proposition 5, Hypothesis H3).

Further, data from the early diffusion phase of generative AI show a distribution that reads

as a headwind for older age groups. According to Bick, Blandin & Deming (2024), based on a

representative U.S. survey, workplace generative AI usage rates as of late 2024 were 34.5%

among those aged 18–29 and 34.6% among those aged 30–39, but 16.7% among those aged 50–

64—those in their 50s and above at roughly half the level of younger workers. In a fivecountry

survey by the employment-support NGO Generation (2024; grey literature), 90% of

U.S. hiring managers said they would consider candidates under 35 for AI-related roles, while

only 32% would consider those over 60; and in an AARP survey (2026; grey literature), only

12% of U.S. workers aged 50 and over had received AI training while 49% said they wanted to

learn—a 37-point gap between motivation and opportunity. These disparities are not

evidence of ability differences. Differences in usage rates include differences in opportunity,

training, and job composition, and the skew in hiring intentions is not even a measured

finding of discrimination from an audit study. Left unaddressed, however, AI can act to widen

the divide along age lines. The same technology can become, depending on design, either a

device for removing the cognitive bottleneck or an amplifier of new exclusion—identifying

the managerial design variables that determine which way it tips is the practical motivation

of this paper.

There is one more headwind: compression pressure on the value of experience itself.

Analyzing the introduction of a generative AI assistant to 5,172 customer-support agents,

Brynjolfsson, Li & Raymond (2025) reported that resolutions per hour rose 15% on average,

but the gains were concentrated among less experienced, less skilled workers (about +30%),

while the most experienced group saw only a small speed improvement and a slight decline

in quality. In a field experiment with 758 BCG consultants (Dell'Acqua et al. 2023; figures from

the working paper version), improvement among lower performers likewise exceeded that

among higher performers on tasks within AI's capability frontier. If AI compresses surfacelevel

differences in skill and experience, the reading that the veteran's premium will

disappear looks natural. But the same experiment also showed that on tasks outside AI's

capability frontier, the probability of reaching a correct answer fell by 19 percentage points in

the AI-using group. What is being compressed is the difference in the production of

deliverables; there is no evidence that the difference in the ability to verify the validity of AI

output against context has been compressed. On the contrary, the more unverified AI output

flows into organizations, the higher the relative value of verification capacity rises. Reading

the headwind data carefully leads not to a negation of this paper's claim but directly to the

question of what is compressed and what is not (Section 4, previewing Proposition 3).

Nor does the vulnerability of the AI transition lie only at the upper end of the age

distribution. An asymmetry has been noted whereby youth, who lose the very entry point for

accumulating experience as entry-level jobs are substituted by AI, are especially vulnerable

(Ayalon 2026). At the upper end, accumulated experience fails to reach work; at the lower

Working Paper | Ageless Management in the AI Era 7

end, the opportunity to accumulate experience itself withers—the ability to address both ends

of the problem simultaneously is the reason this paper adopts the framework of a

multigenerational ecosystem theory rather than a senior-employment theory.

To summarize. The limits of the conventional models derive not from a shortage of

goodwill but from a shortage of theory. Employment extension that leaves the cognitive

bottleneck in place does not let ability reach work. Unconditional celebration of diversity

carries no persuasive force against the near-zero average effects of the meta-analyses. And

the disparity data of the generative AI transition are an ambivalent signal, foretelling

widening exclusion if neglected and connected complementarity if designed for. This paper

rereads each of these headwinds as a demand to specify the conditions, and fixes those

conditions in the form of definitions, propositions, and refutation conditions. This is the

paper's undertaking: to carry out the shift from the paradigm of accommodation and

compensation to a paradigm of co-creation as theory, not slogan.

Evidence-grade note: among the figures in this section on usage rates, hiring intentions, and training

opportunity, Bick, Blandin & Deming (2024) is an NBER working paper (unrefereed; representative

survey), while Generation (2024) and AARP (2026) are surveys by an NGO and a membership

organization (grey literature). All are descriptive statistics based on self-report, not measurements of

ability or productivity.

1.3 Research Questions

From the above problem awareness, this paper poses the following central question. Under

what conditions does the complementation of cognitive abilities by generative AI render role

allocation based on chronological age unnecessary, and enable the cognitive assets of

multigenerational and diverse talent to be converted into the managerial resource of brain

capital? This question decomposes into three subquestions.

RQ1 (removal of the bottleneck): Does AI complementation of Gf-type components

remove the cognitive bottleneck and let retained Gc-type abilities reach work? And when

it does, how does the relative marginal value of Gc-type tasks change (Propositions 1–2)?

RQ2 (the supply side of supervision): For the non-delegable residual of supervising AI

output, does experiential audit capacity—the capacity to detect contextual errors,

practical risks, and ethical risks in AI output—constitute a scarce supply source? Is that

capacity formed by the interaction of long-term domain experience with Gc-type ability

and metacognition? And can generationally heterogeneous supervisor pools

organizationally supply the decorrelation of errors (Propositions 3–5, Hypothesis H3)?

RQ3 (conditional resource conversion): Under what conditions—AI-mediated

complementarity, Brain Safety, and institutional design not confined to the employment

boundary—does the multigenerational ecosystem increase brain capital (stock K ×

utilization rate u) bidirectionally? And what happens when those conditions are absent

(Propositions 6–11, Hypotheses H1–H2)?

Working Paper | Ageless Management in the AI Era 8

The three subquestions stand in a cumulative relationship. If RQ1 does not hold, the job exit

of older workers is the consequence of general ability decline, and this paper's undertaking

loses its foundation. If RQ1 holds but RQ2 does not, supervision in the age of AI becomes a

generic skill unrelated to the extent of experience, and no value specific to multigenerational

staffing arises. If RQ1 and RQ2 hold but the conditional design of RQ3 is absent, expanded

participation becomes indistinguishable from an expansion of unprotected labor supply. That

is, this paper's theory is a three-story structure—a theory of ability (RQ1), a theory of

supervision (RQ2), and a theory of institutions (RQ3)—and the propositions are arranged so

that each story can be refuted independently.

A limitation on the scope of the questions should also be attached. What this paper

addresses is the domain of knowledge work in which the evaluation, supervision, and

contextual judgment of AI output determine outcomes. No generalization is claimed to jobs

whose core is physical skill, or to domains where AI adoption itself does not advance.

Moreover, the term "super-seniors (aged 60 to their 90s)" is used throughout this paper as a

description of an age interval and does not imply that the people within that interval are

homogeneous—on the contrary, the very magnitude of individual differences is the ground

for rejecting allocation by chronological age (previewing Proposition 12).

This paper answers these questions in the mode of a conceptual paper with measurement

proposals. That is, it fixes the theory in testable form through verbatim definitions (8) and

propositions with refutation conditions (12), and draws a map of verification through

hypotheses with explicit identification strategies (3). Most of the paper's propositions are

untested, and the paper does not itself deliver empirical resolution. The value of a theory, on

this view, lies not in declaring it correct but in specifying in advance what evidence would

lead to its rejection.

Terminological discipline is also made explicit in the introduction. The structure in which

humans supervise the operation of AI is written throughout this paper as Human on the Loop

(HOTL), and the expression Human-in-the-Loop is used verbatim only when quoting primary

sources. In this paper the term HOTL is a description of function and is not intended to confer

any doctrinal authority. "Super-seniors (aged 60 to their 90s)" is used as a description of an

age interval, and "socially marginalized groups" as a general term for people who have

experienced exclusion from institutions and markets; neither carries a connotation of

protected status. The concepts requiring definition—Ageless Management, Gf-type/Gc-type

tasks, the cognitive bottleneck, experiential audit capacity, generational decorrelation, AImediated

complementarity, the multigenerational ecosystem, Brain Safety—are fixed as

verbatim definitions in Section 5; the introduction confines itself to previewing their main

points.

Working Paper | Ageless Management in the AI Era 9

1.4 Positioning of This Paper and Adjacent Research

This paper is the ninth in the VURA Working Paper Series and sits in the third layer of the

series' three-layer architecture—Layer 1, era structure (macro): Redefinition Capitalism

(Kadowaki 2026g); Layer 2, social structure (meso): the Self-Defined Society (Kadowaki 2026f);

Layer 3, corporate management (micro): the body of management theories (Figure 1). The

paper inherits its framework from three papers. First, Future Value Theory, which inverts the

origin of value from past cash flows to future value (Kadowaki 2026a). Second, Brain Capital

Management, which conceives brain capital as "stock (K) × utilization rate (u)" and places its

protection and disclosure at the foundation of management (Kadowaki 2026e). Third, the

Human on the Loop (HOTL) theory, which formalizes the value of AI supervision as

independence × detection probability and generalizes the value of a supervision channel to

the "decorrelation of errors" (Kadowaki 2026h). In addition, the framework of enterprise

redefinition and its observation (Kadowaki 2026b, 2026c) and the role-design theory shifting

from the distribution of work to the distribution of purposes and questions (Kadowaki 2026d)

are used for positional reference only.

The content of the inheritance should be specified. From BCM (Kadowaki 2026e), this paper

inherits the measurement framework that conceives brain capital as the product of stock (K)

and utilization rate (u), and the discipline of distinguishing the assessment of states from the

execution of measures. This paper extends that framework from the employees of a single

firm to the participants of a multigenerational ecosystem not confined to the employment

boundary—curbing the depreciation of K among super-seniors, forming K early among youth

and socially marginalized groups, and raising u across the organization (previewing

Proposition 8). From FVT (Kadowaki 2026a), the paper inherits the evaluative inversion that

places the origin of value in future value creation rather than past track records; this is the

premise for bringing onto the evaluation agenda the cognitive assets of youth with few years

of track record and of participants with gaps in their résumés. However, the pathway by

which diversity of thought propagates through discontinuous innovation to ESG evaluations

and the cost of capital is treated in this paper not as a verified fact but as a pathway

hypothesis (Section 6).

Consistency with WP8 (Kadowaki 2026h) is, above all, the lifeline of this paper. WP8 showed

that human supervision of AI often fails—automation bias, alarm fatigue, constraints on

supervisory capacity. Accordingly, this paper does not claim that humans are good at

supervision, nor does it treat the unverified statement that older people are good at AI

supervision as an established fact. The paper's claim takes the following form. Since a

residual of supervision that cannot be delegated to AI exists, the supply side of supervision as

a scarce resource must be interrogated, and experiential audit capacity is a candidate scarce

input. And generational decorrelation is a supply-side response to the decorrelation condition

that WP8 set out. Even if independence is secured, however, detection probability is not

thereby guaranteed. That is precisely why this paper submits this claim itself as a hypothesis

for testing (H3). Note that self-citations within the series are handled in three categories

Working Paper | Ageless Management in the AI Era 10

—"inherited," "referenced," and "tested in this paper"—and claims lacking independent

supporting evidence outside the series are not made to appear established through chains of

self-citation (Section 10). In particular, the extension of HOTL's decorrelation condition to the

generational axis (Proposition 5) and the relationship between BCM's discussion of learning

protection and the non-compressibility of experience (H3) are newly claimed by this paper

and are untested.

On top of this inheritance, the temporal robustness of the theory is declared in advance.

What this paper's propositions rest on is not the transient performance gap of the current

model generation—"today's LLMs are poor at contextual verification." They rest on the

structure of verification independence formalized by WP8 (Kadowaki 2026h, Propositions 10

and 11): whatever the intelligent system, a verifier whose errors are correlated with its own

training distribution cannot self-supply statistical independence in detecting its own

systematic blind spots. The independence of verification is not supplied by improvements in

capability. Because this structure does not depend on the performance profile of any

particular model generation, advances in AI self-correction and self-verification do not, in

themselves, render this paper's propositions obsolete. What changes is the relative

composition of the supply sources of decorrelation—humans with heterogeneous experience,

and models of different lineages—and this paper's framework persists as that allocation

problem. This declaration does not, however, mean that the human contribution within that

allocation is invariant. The possibility that the contribution shrinks, and the means of

detecting it, are addressed head-on in Section 10.5.

Figure 1 The series' three-layer architecture. This paper is No. 9 in Layer 3—corporate management

(micro). Note: the placement is architectural and is not a claim of logical dependence.

The single-point novelty of this paper lies in the following. Whereas conventional senioremployment

and D&I theory has been a paradigm of accommodation and compensation (cost/

Layer 1 Epochal structure (macro) Redefinition Capitalism (RCap)

โ‘ฆ Redefinition Capitalism (2026g)

Layer 2 Social structure (meso) Self-Defined Society (SDS)

โ‘ฅ Self-Defined Society (2026f)

Layer 3 Enterprise management (micrEon)terprise redefinition and the theory system

โ‘  Future Value Theory (2026a) — value theory

โ‘ก Enterprise Redefinition (2026b) — transformation framework

โ‘ข Enterprise Redefinition Observed (2026c) — evidence

โ‘ฃ From JD to Purpose Description (2026d) — role design

โ‘ค Brain Capital Management (2026e) — human foundation

โ‘ง Human on the Loop (2026h) — oversight structure

โ‘จ Ageless Management — this paper: extension along the age axis (extends โ‘ค; bridges โ‘  and โ‘ง)

Working Paper | Ageless Management in the AI Era 11

CSR), this paper presents a management model that converts multigenerational and diverse

talent into co-creating agents of the managerial resource of brain capital, through

bidirectional complementarity between heterogeneous cognitive abilities mediated by AI. A

terminological note is in order. When this paper says bidirectional complementarity, it refers

to the bidirectionality specified in Definition 6 (Section 5)—that the benefits of

complementation (substitution of Gf-type components) and the supply of audit (provision of

Gc-type verification) flow mutually among participants. The "bidirectional capital formation"

of Proposition 8 (that brain capital increases in all participating strata) is a distinct claim, and

the two are used separately. The theoretical contribution condenses into three operations.

First, the operation of re-attributing the supervisory value of older workers from "age" to "the

interaction of long-term domain experience × Gc × metacognition (= experiential audit

capacity)." Age is merely a function of the time that makes accumulation possible, and that is

exactly why this paper is "ageless (independent of age)" rather than a celebration of seniority.

Second, the operation of formalizing generational heterogeneity as an organizational supply

source of the "decorrelation of errors." Third, the operation of converting the three social

problems of population aging, constrained youth participation, and the exclusion of

marginalized groups into untapped supply sources of brain capital, conditional on AImediated

complementarity and Brain Safety. Conversely, the paper also makes explicit the

points on which it claims no novelty. The idea of conceiving AI as a prosthesis for Gf-type

abilities belongs to the lineage of assistive-technology and prosthetics research (Section 2),

and the idea of taking marginalized groups as objects of inclusive design has precedents in

inclusive design. This paper's contribution lies not in the invention of the individual insights

but in the composition that integrates them into a refutable management theory.

Here, the paper explicitly declares the primary–secondary relationship of its scope. The

principal axis of this paper is the axis of age and experience linking super-seniors (aged 60 to

their 90s) and youth. The core of the definitions and propositions—the formation of

experiential audit capacity (Proposition 4), generational decorrelation (Proposition 5)—is

formulated on this axis, and the verification designs (H1–H3) also target this axis. By contrast,

the extension to socially marginalized groups is a secondary context of generalization,

applicable only insofar as the logic of the cognitive bottleneck (Definition 3)—the structure

whereby difficulty in performing some components of a job bundle prevents retained

abilities from reaching work—and AI complementation applies. What must be stressed is that

this generalization does not claim identity of mechanism. Age-related change in Gf-type

components, the constraints on experience accumulation among youth still in development,

and sensory, motor, and cognitive accessibility constraints are heterogeneous as mechanisms,

and transplanting the argument of the aging axis unmodified onto the developmental axis or

the accessibility axis would be an overgeneralization that erases the mechanisms specific to

each axis and the contexts of the people concerned. What this paper claims is limited to the

commonality of the allocation structure—that difficulty in performing some components of a

bundle forces exclusion from the whole bundle—and the applicability of the logic that AI

Working Paper | Ageless Management in the AI Era 12

complementation can act on that structure. The particulars of marginalized-group

participation grounded in the mechanisms, institutions, and participatory research specific to

each axis remain a task beyond this paper.

In addition to the primary–secondary relationship of scope, the meta-structure of the

theory's lineage is also declared. This paper traverses several bodies of literature—cognitive

science (Section 2), health epidemiology (Section 3), and institutional theory (Section 7)—but

these are not parallel pillars. The principal axis of this paper is the theory of firm competitive

advantage. The paper connects to the lineage of the resource-based view, which holds that a

managerial resource becomes a source of sustained competitive advantage when it is

valuable, rare, inimitable, and non-substitutable (Barney 1991), and of dynamic capabilities

theory, which explains competitive advantage by the capacity to integrate and reconfigure

such resources and competences under changing environments (Teece, Pisano & Shuen 1997),

in the following way. Experiential audit capacity (Definition 4) is a path-dependent product

that can be accumulated only over time through long-term domain experience (Propositions

3 and 4); it cannot be procured instantly on the market, and, if Proposition 3 is correct, cannot

be replicated by AI—in that sense it is a candidate for a rare and inimitable supervisory

resource. And the composition that organizes this resource—the supervisory portfolio that

designs the decorrelation of errors (Section 6), dynamic role allocation (Proposition 12)—sits

at the level of dynamic capabilities that reconfigure resources under change. FVT's (Kadowaki

2026a) evaluative inversion toward future value also connects with this paper on this lineage

of competitive advantage. By contrast, institutional theory (Section 7) and the health

discussion (Section 3) are positioned as the institutional context and the complementary

assets that make the attainment of competitive advantage possible. The protective institutions

for non-employment participation are the institutional context that enables connection to the

resource of experiential audit capacity, and brain health and Brain Safety (Section 8) are the

complementary assets that prevent the resource's depreciation and sustain its utilization. The

descriptions in Sections 3 and 7 should therefore be read not as free-standing health theory

or policy theory but as the specification of the conditions for competitive advantage. This

structure is reconfirmed in the conclusion (Section 11).

The paper's normative anchor is also made explicit. When this paper says "brain capital," it

is not a resource concept that decomposes human beings into parts of brain function and

selects and reuses only the parts with market value. Such a reading—the commodification of

human beings, or a eugenic reading that sorts people by cognitive characteristics—this paper

explicitly rejects. The paper's concept of capital connects to Sen's (1999) capability approach,

which conceives development as the expansion of the substantive freedoms to live the lives

people have reason to value, and the protection and extension of brain capital are positioned

as the foundation of cognitive dignity and human flourishing. That is, measurement and

allocation are not devices for ranking people but means of returning to individuals the

freedom of participation that the crude proxy variable of chronological age has taken from

them. This normative premise functions not as decoration but as a design constraint. That

Working Paper | Ageless Management in the AI Era 13

Brain Safety (Section 8) includes protection from exploitation among its requirements, and

that Section 10 discusses the misuse of measurement and its diversion into selection as risks

of this paper itself, are consequences of this anchor. Even so, the danger that measurement

converts into a new selection device does not disappear, and this paper takes on that danger

as self-criticism in Section 10.

The differences from adjacent research are as follows. Dawson et al. (2022) advocated

investment in "Late-Life Brain Capital" and called for a value shift integrating the knowledge

and experience of older people into the economy and society. This is the direct predecessor

framework of this paper, but it remains an invited-article advocacy piece and lacks

definitions, propositions, refutation conditions, and verification designs at the level of a

management regime. This paper operationalizes that advocacy into a theory that firms can

implement and test. Ayalon (2026) discusses intergenerational relations in the workplace in

the age of AI and points out the asymmetry that youth may be especially vulnerable owing to

the substitution of entry-level jobs. Where that article discusses, from the standpoint of

ageism and social policy, how AI generates intergenerational tension, this paper formalizes,

from the standpoint of management and organization, the design conditions under which

multigenerational staffing generates value—the two stand in a relationship of

complementarity, not opposition. Nielsen (2023) discussed a division of labor in which older

knowledge workers use experience to winnow the options AI generates ("wise winnowing"),

but this is an essay without empirical data, and this paper translates that intuition into

testable propositions (Propositions 3–5) and an experimental design (H3). A publication with

the similar title "Ageless Collaboration" also exists, but this paper does not rely on it. Finally,

the research gap is made explicit. Whether expertise improves AI supervision is itself

contested—there are reports that experts, too, overlook AI's errors—and, a fortiori, no study

operationalizing age or tenure to measure AI-supervision performance could be found within

the scope of this paper's search (as of August 21, 2026). This paper's propositions give this gap

a testable form.

1.5 Structure of This Paper

The structure of this paper is as follows. Section 2 organizes the cognitive-science

foundations: the Gf/Gc distinction, the heterogeneity of ability-specific peak ages, the lineage

of selective optimization with compensation (SOC) and the extended mind, and the

controversy over cognitive reserve—presented as controversy. Section 3 examines the

evidence on work and brain health with evidence grades, identifying the difficulties of causal

identification and the moderating variable of "quality of work." Section 4 organizes the

evidence on what generative AI compresses (differences in production) and what it does not

(differences in verification), and carries out the conceptual separation of domain experience

from chronological age. Section 5 is the theoretical core of the paper, presenting eight

definitions and twelve propositions together with refutation conditions. Section 6 discusses

the design of the multigenerational ecosystem: the role reconfiguration of the four players,

Working Paper | Ageless Management in the AI Era 14

the operation of dynamic role allocation, and the organizational design of generational

decorrelation. Section 7 compares internationally the institutional hurdles to employmentbased

and non-employment-based participation, treating as a proposition the risk of

degeneration into exploitation brought about by gaps in institutional protection. Section 8

designs Brain Safety in both directions: protection from cognitive load and protection from

exploitation. Section 9 presents the map of measurement and verification, detailing the

identification strategies and experimental designs for Hypotheses H1–H3. Section 10 discusses

the paper's limitations and self-criticism—the risk of new stereotyping, the misuse of

measurement, survivorship bias, and conflicts of interest—and Section 11 concludes.

Appendix A contains the experimental protocol for Hypothesis H3, and Appendix B the

details of the institutional comparison.

A reading guide is attached. Readers who wish to confirm only the skeleton of the theory

may read the definitions and propositions of Section 5 and the map of verification of Section 9

first. Practitioners interested in implementation will find Section 6's design discussion and

Section 8's Brain Safety guidelines relevant. The limitations and self-criticism of Section 10,

however, are part of this paper's claims, and it is hoped that they will be consulted alongside

the propositions whenever these are cited or used. Until it is tested, this paper's theory is a

proposal, and depending on the results of testing it may be rejected—presenting it with that

possibility left open is, on this view, the first safeguard against the term Ageless Management

turning into a device of new age stereotyping.

2. Cognitive-Scientific Foundations

This section organizes the findings of cognitive science that constitute the theoretical

premises of Ageless Management. The purpose is fourfold. First, it introduces the distinction

between fluid intelligence (Gf) and crystallized intelligence (Gc) accurately from the original

sources, and confirms that this distinction is a relative weighting on a continuum, not a

binary classification (Section 2.1). Second, it presents the empirical finding that the peak ages

of cognitive abilities differ greatly across abilities, while at the same time making explicit the

limitations of the cross-sectional designs on which that finding rests (Section 2.2). Third, it

introduces SOC theory, which frames adaptation to aging as "selective optimization with

compensation," and the "extended mind" framework, which does not confine cognition to the

inside of the skull, and positions cognitive complementation by AI within this lineage of

compensation (Section 2.3). Fourth, it organizes the controversy surrounding the concept of

cognitive reserve and the "use it or lose it" hypothesis, together with the evidence on both

sides (Section 2.4). The work of this section is the foundation for the subsequent theory

construction (Section 5), and it places priority on drawing a clear line between the parts

where the evidence is established and the parts that remain contested.

Working Paper | Ageless Management in the AI Era 15

2.1 Fluid and Crystallized Intelligence — Cattell-Horn Theory

The standard starting point for discussing age-related change in intelligence is the distinction

between fluid intelligence (Gf) and crystallized intelligence (Gc). Cattell (1963) formulated,

through a critical experiment, the theory that intelligence is not reducible to a single general

factor but is composed of two general factors of different natures. Fluid intelligence (Gf) is the

reasoning capacity for solving novel problems, referring to cognitive operations that depend

relatively little on prior knowledge — working memory, information-processing speed, and

the learning of novel procedures. Crystallized intelligence (Gc) is the knowledge and skill

accumulated through education, occupation, and life experience, manifesting as vocabulary,

general knowledge, contextual interpretation, and interpersonal judgment.

The original source that demonstrated the age differences of these two factors with crosssectional

data is Horn & Cattell (1967). That study reported a divergence pattern in which

fluid intelligence (Gf) shows a declining trend with age throughout adulthood, while

crystallized intelligence (Gc) shows a trend of maintenance or increase. This divergence is the

starting point of the entire argument of this paper. That is, age-related cognitive change is not

a uniform "decline" but a structural change in which declining components and maintained

or growing components coexist. The singular statement "cognitive function declines with age"

is a summary that paints over this structure with an average, and in the process of

summarizing it discards the information most important for management — which

components remain and which components are lost. From the standpoint of role design, the

identification of the components that are lost and the components that remain is precisely the

starting point, and the Gf/Gc distinction provides the minimal vocabulary for that purpose.

Note that the study is based on cross-sectional data, and it is not appropriate to attribute

specific onset ages of decline or effect sizes to that paper (the limitations of cross-sectional

designs are detailed in Section 2.2).

Two conceptual cautions should be made explicit here. First, Gf and Gc are statistical

constructs extracted by factor analysis, and it is rare for a real task or job to depend on only

one of them. Drafting a document requires both vocabulary (Gc) and working memory (Gf),

and negotiation requires both contextual interpretation (Gc) and immediate information

processing (Gf). Accordingly, the distinction between "Gf-type tasks / Gc-type tasks" that this

paper introduces in Section 5 (Definition 2) is a relative weighting on a continuum — on

which component the task's performance primarily depends — and not a binary

classification. Blurring this point leads to a new simplification — "just give older workers Gctype

jobs" — which is itself a variant of the age stereotyping this paper criticizes.

Second, the Gf/Gc distinction does not imply a typology of individuals in which "the young =

Gf, the old = Gc." For both factors, individual differences are no smaller than age differences,

and the distributions of age groups overlap widely. People in their 70s who retain high fluid

intelligence (Gf), and people in their 30s with rich crystallized intelligence (Gc), are entirely

ordinary. What this section treats are average tendencies of age groups, and the validity of

Working Paper | Ageless Management in the AI Era 16

using age as the basis for placing individuals is precisely what this paper examines negatively

in Section 5 (Proposition 12).

The reason this paper adopts the Gf/Gc distinction as the basis of its theory lies in its

parsimony and in the directness of its mapping onto the functional characteristics of AI.

Regarding the factor structure of intelligence, more multilayered models have been

developed after the Cattell-Horn lineage, but for this paper's purpose — constructing a

framework that identifies which cognitive components of tasks AI can cheaply substitute for

or complement and which parts remain on the human side — the two-component distinction

between the bundle of processing speed, working memory, and novel learning (Gf) and the

bundle of accumulated knowledge, contextual interpretation, and interpersonal judgment

(Gc) provides necessary and sufficient resolution. What generative AI primarily substitutes

for is Gf-type load — search, summarization, documentation, and the execution of routine

procedures — and this correspondence is the premise of Proposition 1 (the marginal value

shift) in Section 5.

An observation follows immediately from this framework. A real job is not a single task but

a bundle of tasks in which Gf-type components and Gc-type components are tied together.

And under conventional employment institutions, this bundle has been required to be

performed in its entirety by the same individual. As a consequence, when the performance of

a Gf-type component within the bundle — for example, mastering the operating procedures

of a new business system, or processing large volumes of material at high speed — becomes

difficult, exit from the job as a whole is forced, however high the level at which the rest of the

bundle can be performed. This structure, in which retained crystallized intelligence (Gc) is

lost without ever reaching the work, is what this paper formalizes in Section 5 as the

cognitive bottleneck (Definition 3). The divergence pattern confirmed in this section — the

coexistence of declining components and maintained or growing components — is the

empirical underlay of that formalization. Conversely, if the Gf/Gc divergence did not exist, or

if every ability declined with the same slope, this paper's central claim — that selective

complementation by AI changes the employability of older workers — would not hold. In that

sense, the findings of this section are the precondition on which the viability of this paper's

theory turns.

Furthermore, for the subsequent argument (Proposition 4 in Section 5), one more

distinction is introduced within the accumulation-type abilities: the distinction between

generalized crystallized intelligence — knowledge and skills such as vocabulary, reading

comprehension, and general education that are exercised across broad domains detached

from the context of acquisition — and domain-specific knowledge — facts, procedures, and

operational schemata that have meaning within a particular industry, job, or practice. Both

are products of learning and experience, and within the framework of Cattell-Horn theory

they are grouped on the same side as acquired abilities, but they differ in transferability.

Generalized Gc transfers across domains, whereas domain knowledge is bound to the domain

in which it was formed, and its value drops sharply outside that domain. What the

Working Paper | Ageless Management in the AI Era 17

vocabulary and general-knowledge scales of intelligence tests capture is mainly the former,

and what research on occupational expertise has targeted is mainly the latter. This distinction

matters for this paper. What long-term occupational experience accumulates is mainly the

latter — knowledge and operational schemata concerning the cases, failures, and tacit

constraints of the domain in question — and that is a variable distinct from vocabulary-test

scores. When Proposition 4 in Section 5 attributes the formation of experiential audit capacity

to "the interaction between domain-specific knowledge and operational schemata,

generalized crystallized intelligence, and metacognition," this conceptual distinction between

the two components is presupposed.

There is also a construct closely related to, but not identical with, crystallized intelligence

(Gc): metacognition — the capacity to monitor and control, from a bird's-eye view, one's own

cognitive states and the limits of one's knowledge. The experiential audit capacity that this

paper defines in Section 5 (Definition 4) is a capacity operationally defined by detection

performance on tasks of verifying AI output, and it is not equated with Gc alone. This paper

attributes its formation to the interaction of long-term domain experience with Gc-type

abilities and metacognition, but this is not part of the definition; it is an empirical claim to be

tested (Proposition 4). Possessing much knowledge and knowing where one's knowledge runs

out are distinct abilities, and in the context of verifying AI output, this paper holds that the

latter — the judgment of where to stop, what to doubt, and whom to consult — plays the

decisive role. This distinction is also consistent with the wisdom research discussed below

(Section 2.2), which includes "recognition of the limits of one's knowledge" among the core

dimensions of wise reasoning.

2.2 Peak Ages by Ability — There Is No Single Prime

The Gf/Gc divergence pattern was shown at finer resolution by Hartshorne & Germine (2015).

That study comprehensively analyzed two lines of evidence — cognitive-task data from

48,537 web-based participants and normative data from standardized intelligence and

memory tests — and concluded that "there is considerable heterogeneity in the peak ages of

cognitive abilities: some abilities peak around high-school graduation, some plateau in early

adulthood and begin declining in the 30s, and others do not peak until the 40s or later." The

ability-by-ability guide values are given in Table 1. Processing speed peaks at about age 18

and visual working memory at about age 25, while reading others' emotional states plateaus

from ages 40 to 60, and vocabulary peaks in the late 60s to early 70s. Even within the same

person, decades separate the point at which the first ability passes its peak from the point at

which the last ability reaches its own. Especially noteworthy is that the reading of others'

emotional states — a component of social cognition — remains on a plateau from ages 40 to

60, the middle to later stretch of a working life. This suggests that the cognitive basis of roles

requiring interpersonal judgment — negotiation, coordination, mentoring, and the contextual

evaluation of AI output on which this paper focuses — is maintained until far later than the

peak of processing speed.

Working Paper | Ageless Management in the AI Era 18

Table 1 Peak ages by cognitive ability (guide values from cross-sectional data)

Ability (task) Approximate peak age Source

Processing speed (digit-symbol

coding)

About age 18 Hartshorne &

Germine (2015)

Visual working memory About age 25 Hartshorne &

Germine (2015)

Working memory for numbers

(digit span)

Early to mid-30s Hartshorne &

Germine (2015)

Emotion reading (Reading the

Mind in the Eyes)

Plateau from ages 40 to 60 Hartshorne &

Germine (2015)

Vocabulary Late 60s to early 70s Hartshorne &

Germine (2015)

Wise reasoning (social-conflict

tasks)

Group aged 60+ scored higher than young

and middle-aged groups

Grossmann et al.

(2010)

Note: All rows are based on cross-sectional designs comparing age groups and do not directly show individual

developmental trajectories. Peak ages are guide values to be interpreted with latitude. Grossmann et al. (2010) is a

comparison across age groups rather than an estimate of peak age, and thus differs in kind from the other rows.

Figure 2 The asynchrony of peak ages across abilities. Note: Schematic. Based on Hartshorne &

Germine (2015) and related sources. Because of the cross-sectional design, note the confounding with

cohort effects.

The implication of this finding is that the question "when does cognitive function peak?" is

itself ill-posed. What Hartshorne & Germine (2015) showed is that there is no single age at

which all abilities reach their summit simultaneously — no cognitive prime — and that at

nearly every age, abilities that are still rising, abilities on a plateau, and abilities in decline

coexist. Figure 2 depicts this asynchrony schematically. This structure invalidates one-

20 30 40 50 60 70 80

Age (years)

Relative performance (schematic)

Processing speed (peak ≈ age 18) Working memory (peak ≈ age 25–30)

Emotion reading (plateau, 40s–60s) Vocabulary / crystallized knowledge (late 60s–early 70s)

Working Paper | Ageless Management in the AI Era 19

dimensional rankings along the axis of age — both "the younger, the more capable" and "the

older, the more seasoned." The processing speed of an 18-year-old and the vocabulary of a 70-

year-old are the relative heights of different abilities at different points in time, not ranks on a

single scale. Translated into the language of organizational design, members of different ages

should be described not as having "more or less of the same ability" but as having "different

ability profiles," and the heterogeneity of profiles is a resource that can be deployed

complementarily according to the cognitive demands of bundled tasks. The validity and

conditions of this translation — in particular, when mediation by AI becomes necessary — are

the subject of Sections 5 and 6.

In the domain of social judgment as well, components that improve with age have been

reported. Grossmann et al. (2010) had approximately 250 adults (sample size pending

verification against the original; three groups — young, middle-aged, and 60+) read materials

depicting social conflicts and then had their reasoning about subsequent developments blindrated

on dimensions of wisdom (recognition of others' perspectives, recognition of the limits

of one's own knowledge, emphasis on compromise and multiple resolutions, and so on). The

group aged 60 and over scored higher on wise reasoning than the young and middle-aged

groups, and this advantage held after controlling for fluid intelligence (Gf) and social class.

However, what this result measures is rated scores on a particular social-reasoning task, and

it does not support the generalization that "older people are wise." Extrapolation to

management judgment is an inference across differences in task characteristics, and this

paper treats it at the level of suggestion.

Next, the limitations of the cross-sectional designs on which these figures rest must be

made explicit. A cross-sectional study compares different age groups at a single point in time.

Thus the statement "vocabulary peaks in the late 60s to early 70s" means that the group at

that age at the time of measurement had the highest mean score; it does not directly show

that an individual's vocabulary keeps growing to that age. Cohort effects — intergenerational

differences in educational attainment and intellectual environment, the so-called Flynn effect

— inevitably contaminate age differences. Indeed, Hartshorne & Germine (2015) themselves

found that the peak of vocabulary had shifted about 15 years later than in the Wechsler

standardization data of the 1970s–80s, and cited changes in the modern environment of

intellectual stimulation as a factor. The estimated peak ages themselves move with cohorts.

Furthermore, web-based samples may carry self-selection biases with respect to education

and digital literacy.

From the standpoint of this paper, the existence of cohort effects is at once a limitation and

an implication. That age differences in peak ages and ability levels include generational

differences in educational and intellectual environments means there is no guarantee that

the values observed for today's cohorts in their 60s and 70s will apply unchanged to the

cohorts who will be in their 60s and 70s twenty years from now. Future super-seniors (aged

60 to their 90s), with longer years of education and deeper exposure to digital environments,

may enter old age with cognitive profiles different from today's same-age cohorts. Embedding

Working Paper | Ageless Management in the AI Era 20

age into institutions as a fixed indicator of ability means taking on, in addition to the

divergence between group means and individuals, a second source of error: cohort drift. This

too is one of the reasons allocation criteria should be moved from age to measured

characteristics.

The divergence between cross-sectional and longitudinal estimates has sharpened into a

public controversy over the onset age of decline. Salthouse (2009), on the basis of data

including 2,350 cross-sectional and 729 longitudinal participants, argued that "some aspects

of age-related cognitive decline begin in healthy, educated adults in their 20s and 30s." In that

study, the peaks of 12 cognitive variables lay in the range of ages 22–27, and significant

declines were detected at ages 27–42. Performance differences between ages 18 and 60 reach

about 1 standard deviation on speed tasks and 0.6–0.7 standard deviations on reasoning and

memory variables. At the same time, the same study also reported that knowledge-based

measures such as vocabulary and general knowledge rise consistently until at least age 60 —

the Gf/Gc divergence pattern is reconfirmed here as well. This paper notes, however, that the

article was an invited controversy paper, and a critical comment by Schaie, who led the

Seattle Longitudinal Study, was published alongside it in the same journal (Schaie 2009). The

rebuttal is that by longitudinal estimates the onset of decline comes after midlife, and that

cross-sectional estimates make the onset look excessively early because of cohort effects. That

is, the controversy structure itself — "decline begins in the 20s (cross-sectional)" versus "it

begins after midlife (longitudinal)" — is the current state of the evidence, and this paper does

not adopt either side as settled doctrine.

What matters for this paper's argument is that however this controversy is resolved, the

core findings of this section — the heterogeneity of peak ages across abilities, and the longterm

maintenance and growth of knowledge-based abilities — are unshaken. In both crosssectional

and longitudinal evidence, the divergence between the trajectories of processingspeed-

type abilities and knowledge-accumulation-type abilities is consistently observed. The

dispute concerns the onset timing and slope of decline, not the existence of the divergence.

In deriving managerial implications, two further reservations must be layered on. First,

average trajectories do not justify the placement of individuals. The figures above concerning

peak ages are all means of age groups; individual differences within each age group are large,

and the distributions across groups overlap widely. The difference of about 1 standard

deviation on speed tasks between ages 18 and 60 reported by Salthouse (2009) is large as a

difference of group means, but not large enough to erase the overlap of distributions.

Inference from mean differences to the treatment of individuals is the very structure of

statistical discrimination. Second, "peak ages" based on cross-sectional data must not be

translated literally into an individual's career design. The cross-sectional finding that

"vocabulary grows until the late 60s" does not guarantee that any particular individual's

vocabulary will keep growing to that age. The reason this paper argues, from Section 5

onward, for allocation based on measured characteristics rather than age is precisely to

absorb these two reservations — the divergence between group means and individuals, and

Working Paper | Ageless Management in the AI Era 21

the divergence between cross-sectional and longitudinal evidence — on the side of

institutional design. The variable of age is a coarse proxy that appears valid only when all of

these divergences are ignored.

2.3 SOC Theory and the Extended Mind — Placing AI in the Lineage of

"Compensation"

On how individuals adapt to age-related change in cognitive resources, developmental

psychology has offered a picture different from "passive acceptance of decline." SOC theory

(selective optimization with compensation), presented by Baltes & Baltes (1990), frames

successful aging as the cooperation of three processes: (1) Selection — narrowing goals and

domains of activity; (2) Optimization — concentrating remaining resources on the narrowed

domains to maintain mastery; (3) Compensation — making up for lost means with substitute

means. In a celebrated example that Baltes and colleagues used repeatedly, the elderly pianist

Rubinstein recounted that he narrowed his repertoire (selection), concentrated his practice

on it (optimization), and deliberately played the passages just before fast ones more slowly so

that the fast passages would sound faster by contrast (compensation). At the stage of the 1990

theoretical chapter, SOC theory was a presentation of a framework, and the empirical

demonstration of its effects was left to subsequent research; but in formalizing adaptation to

aging as a design problem of resource allocation and substitution of means, it is the direct

scaffold of this paper's theory construction.

The "selection" process is underwritten on the motivational side by the socioemotional

selectivity theory of Carstensen, Isaacowitz & Charles (1999). According to that theory, the

selection of social goals is determined not by chronological age itself but by the "perception of

remaining time." When time is perceived as plentiful, knowledge-acquisition goals take

priority; when time is perceived as limited, goals of emotional meaning and emotion

regulation take priority. What matters is that the theory asserts not a "decline" of motivation

with age but a "reallocation," and moreover that time perception is plastic: young people

under time constraints show the same goal shifts as older people. This finding — that the

driver is time horizon, not age — is consistent with this paper's direction of removing

chronological age from the seat of explanatory variable. As an implication for role design, for

participants whose time horizon prioritizes goals of emotional meaning and emotion

regulation, meaning-fulfilling roles such as developing successors and assuring the quality of

judgment may have high motivational fit. But this is a suggestion from theory, not a

justification for assigning roles to particular age brackets — time horizon does not correlate

perfectly with age, and it is plastic.

The means of compensation are not confined to the interior of the individual. The

"extended mind" thesis of Clark & Chalmers (1998) presented an active externalism on which

cognitive processes are not closed within the skull and external tools and records can

function as constituents of the cognitive system. In the famous thought experiment, Otto, who

has a memory impairment, uses a notebook as a substitute for memory. If the notebook is

Working Paper | Ageless Management in the AI Era 22

reliably and constantly consulted and its contents are automatically trusted, it is a bearer of

belief on a par with in-head memory (the parity principle). What must be stressed here is that

this is a philosophical argument, not empirical research. This paper does not use this

framework as grounds for the empirical proposition that "older adults' cognition can be

supplemented by external tools." It is used as the original source of a conceptual framework

that does not confine the locus of cognitive ability to the individual skull, and as justification

for a perspective that includes external resources, AI among them, within the object of design

as parts of the cognitive system. This perspective carries a practical implication for the unit of

ability assessment. What role-allocation judgment should ask is not "what can this individual

do unaided?" but "what can this individual do as a system with reliably available external

resources?" In the same way that an institution that measured a spectacle-wearer's eyesight

without the spectacles to judge fitness to drive would be irrational, measuring performance

capacity under unaided conditions in a work environment where AI complementation is

standard is a mistaking of the object of measurement. However, as discussed later (Section 4),

the divergence between aided performance and unaided ability is itself also a risk to be

managed.

At the intersection of these two lineages — the "compensation" of SOC theory and the

"externalization" of the extended mind — stands technology as prosthesis. Spectacles and

hearing aids are prostheses of the senses; notebooks and calendars are prostheses of memory.

The fields of Assistive Technology, HCI, and disability studies have accumulated decades of

research on cognitive prosthetics for cognitive impairment and aging. The idea of using

generative AI as a prosthesis for Gf-type components — search, summarization,

documentation, and the execution of novel procedures — clearly belongs to this lineage, and

this paper claims no novelty for the idea itself. Technologies of cognitive accessibility —

memory-aid systems, reminders, screen readers, voice interfaces — were researched and

implemented before the advent of AI as means of technically lifting the exclusion from

activity grounded in cognitive constraints. Likewise, the practice of inclusive design, which

treats disability and old age not as objects of special accommodation but as initial conditions

of design, also precedes this paper, and the perspective of treating the participation of socially

marginalized groups as a design problem is not this paper's invention either. This paper's

theoretical contribution does not lie there. The novelty this paper claims, as shown in Section

5, lies in bidirectional complementarity (Definition 6) — heterogeneous cognitive assets

connected across individuals through the mediation of the AI prosthesis — and in formalizing

it as an allocation problem of management resources. Whereas a prosthesis is a one-way

relation that fills an individual's deficit, what this paper treats is an organizational structure

in which the heterogeneous abilities freed by prostheses audit and complement one another.

Summarized from the standpoint of SOC theory, this paper's undertaking is as follows.

Traditionally, selection, optimization, and compensation were adaptive strategies that

individuals carried out within their own resource constraints. When AI provides

compensation for Gf-type components cheaply, the unit of selection and optimization

Working Paper | Ageless Management in the AI Era 23

becomes extensible from the individual to the organization. That is, room emerges to design

organizationally — on the basis of measured characteristics rather than the coarse proxy of

chronological age — who selects which roles and which abilities to optimize for. This is the

cognitive-scientific grounding of dynamic role allocation in Section 5 (Definition 1 and

Proposition 12).

At the same time, it should be foreshadowed that AI as compensation has a distinctive

failure mode that traditional prostheses largely lacked. Spectacles do not form false images,

but generative AI generates hallucinations, outputting plausible errors in fluent form.

Compensation of Gf-type components by AI therefore generates a new cognitive load — the

verification of output — a load that, this paper argues, requires Gc-type abilities and domain

experience. The structure in which compensation is not an unconditional solution but

demands another scarce resource, oversight, is treated in earnest in Section 4 (the evidence

on failures of AI oversight) and Section 5 (experiential audit capacity, Definition 4). What

should be confirmed at the level of this section is that this structure can be described without

contradiction inside SOC theory. The introduction of a means of compensation demands new

selection and optimization — who should attain mastery in verification — and that allocation

problem is precisely the subject of this paper.

2.4 Cognitive Reserve and the Controversy over "Use It or Lose It"

In discussions of Ageless Management, the claim that "continuing to work helps maintain

cognitive function" is often placed as a premise. This paper does not adopt that premise

without examination. This section distinguishes two related concepts — cognitive reserve and

the "use it or lose it" hypothesis — and then organizes the current state of the evidence

together with the arguments on both sides. The two are often spoken of interchangeably, but

the structure of their claims differs. Cognitive reserve is a claim of "buffering of level" — that

accumulated intellectual assets buffer the impact of brain pathology and age-related change;

use it or lose it is a claim of "alteration of slope" — that continued intellectual activity changes

the rate of decline itself. The argumentative move of supporting the latter with evidence for

the former, while leaving this distinction blurred, circulates widely in practitioner discourse.

Cognitive reserve is a concept introduced to explain individual differences in the ability to

maintain cognitive function even in the face of age-related brain change or pathology. Stern

(2002) presented a framework distinguishing passive "brain reserve," referring to hardwarelike

individual differences such as brain volume, from active "cognitive reserve," the

efficiency and flexibility of processing. Years of education, occupational attainment, and

engagement in intellectual activity are used as proxy indicators of reserve, and the

framework has permeated the epidemiology of aging and dementia (Stern 2012). However, as

Stern himself repeatedly cautions, the evidence for reserve consists of observational findings

based on proxy indicators, and the causal assertion that "cognitive reserve prevents

dementia" cannot be drawn from the current evidence. In addition, the limitations of the

proxy indicators themselves require attention. Years of education and occupational

Working Paper | Ageless Management in the AI Era 24

attainment are strongly entangled with socioeconomic status, upbringing environment, and

childhood intelligence, and even if an association between these indicators and late-life

cognitive function is observed, it is difficult to disentangle from observational data alone

whether it is an effect of accumulated intellectual load or a reflection of confounding initial

conditions.

An example of observational evidence consistent with the cognitive-reserve hypothesis is

Staff, Murray, Deary & Whalley (2004). Using 92 members of the Aberdeen 1921 birth cohort,

with intelligence-test scores at age 11 and cognitive function at age 79, together with brain

MRI in old age, the study examined whether three candidate proxies of reserve (years of

education, intracranial volume, and occupational attainment) predicted late-life cognitive

function even after controlling for childhood intelligence and age-related brain change. The

results were that education explained 5–6% of the variance in late-life memory, and

occupational attainment about 5% of memory and 6–8% of reasoning, while intracranial

volume was not a significant predictor; the authors concluded that this was consistent with

an active model in which lifetime intellectual load accumulates reserve. This paper cites this

study as observational evidence supporting the cognitive-reserve hypothesis. What must be

stressed is that this study is not a demonstration of the "use it or lose it" hypothesis. The

sample is an observational study of 92 people and cannot establish causation; what was

measured was the association between the proxy indicators of education and occupation and

late-life cognitive level, not whether continued intellectual activity changes the rate of

decline. Moreover, the variance explained is small, at 5–8%.

What made this distinction between "level" and "slope" explicit is Salthouse's (2006) critical

examination of the "use it or lose it" hypothesis. Salthouse pointed out that the existence of a

correlation between intellectual activity and cognitive performance (more active people score

higher) and the claim that intellectual activity changes the rate of age-related decline itself

are separate claims, and that it is the latter that requires testing. Examining the available

evidence on that basis, he concluded that "at present there is little scientific evidence that

differences in engagement in intellectually stimulating activities alter the rate of mental

aging." This critique weighs heavily on this paper. The popular claim that "keeping working

protects cognitive function" is sustained by an unreflective slide from the correlation between

activity and level (which is observed) to an activity-induced change in slope (which is not

established).

However, summarizing Salthouse (2006) as "use it or lose it has been refuted" is equally

mistaken. Salthouse's conclusion is insufficiency of evidence, not disproof, and he

acknowledges the correlation between activity and performance itself. Furthermore,

Schooler (2007) presented a rebuttal in the same journal, and the question remains an open

controversy on the same pages. Salthouse himself appends a practical recommendation: even

though evidence that decline can be slowed is unestablished, there is also no evidence of

harm, so people should behave as if the hypothesis were true and continue intellectually

stimulating activities. To summarize the current state of the evidence: (1) correlations

Working Paper | Ageless Management in the AI Era 25

between intellectual activity, education, and occupation and cognitive level have been

observed repeatedly; (2) causal evidence that these change the slope of decline is lacking; (3)

the controversy is ongoing.

This controversy gives direct discipline to this paper's hypothesis design. Following

Salthouse's (2006) distinction, to claim brain-health effects of Ageless Management requires

three things: (1) showing a difference in the slope of decline, not the correlation that workers

have higher cognitive levels; (2) distinguishing health selection — the healthier keep working

— from reverse causation — cognitive decline causes retirement; and (3) identifying the

mediation — through which pathway any effect runs. How far the empirical literature on

retirement and cognitive function since the so-called Mental Retirement study has answered

this identification problem is examined in Section 3, and this paper's own testing design

(Hypothesis H1 in Section 9) makes explicit, in line with this discipline, an identification

strategy using exogenous variation and a mediation analysis of cognitive engagement.

This paper does not adjudicate this controversy. Not adjudicating it is the starting point of

this paper's theoretical construction. That is, this paper does not treat "work protects brain

health" as an established premise; it examines the empirical evidence on work and brain

health with grades in Section 3, formalizes the claim as a mediation hypothesis (Proposition

7) — if an effect exists, its pathway runs not through hours of work but through the

maintenance of cognitive engagement — and reduces it to the testable Hypothesis H1 (Section

9). What this paper inherits from the cognitive-reserve framework is not a causal assertion

but a structural perspective — that accumulated intellectual assets can function as a buffer

against age-related change — and this connects to the stock (K) concept of brain capital (Brain

Capital) (Sections 5 and 8).

Finally, one more finding about the shape of the trajectory of decline should be added, as it

gives a boundary to this paper's theory. The main analytic focus of Salthouse (2009), treated

in Section 2.2, is adults aged 18–60, but the paper also refers, supplementarily, to evidence

that the magnitude of age-related decline accelerates at older ages. In a cross-sectional sample

of about 800 adults aged 61–96 from the author's laboratory, the per-year gradient of decline

exceeded the estimates for adults under 60 — a gradient about 2 times as steep on speed

variables and nearly 4 times as steep on memory variables. This is a comparison across age

groups based on cross-sectional data — the reservation of Section 2.2 about the

contamination of cohort effects applies as it stands — and it does not directly show

nonlinearity in individual trajectories. But the finding suggests that the decline of fluid

intelligence (Gf) in old age cannot be extrapolated indefinitely as "a gentle constant slope" —

that is, decline may not remain linear but may accelerate. The implications for this paper's

theory are two. First, compensation of Gf-type components by AI (Section 2.3) may have a

floor — a state in which attentional resources themselves are depleted cannot be filled by

external complementation, and Proposition 2 in Section 5 builds this floor in explicitly as a

boundary condition of complementation. Second, making this boundary explicit is not a

Working Paper | Ageless Management in the AI Era 26

justification for exclusion in old age but an honest demarcation of the theory's scope of

application. The implications of the boundary are discussed again in Section 10.

The upshot of this section can be summarized in a single point. This section has confirmed

the Gf/Gc divergence (Section 2.1), the heterogeneity of peak ages across abilities and its crosssectional

limitations (Section 2.2), the theoretical lineage of compensation (Section 2.3), and

the unresolved controversy over activity and cognitive maintenance (Section 2.4). The

conclusion running through them is that the evidence of cognitive science shows there is no

single cognitive prime at which all abilities reach their summit simultaneously. Processingspeed-

type abilities and knowledge-accumulation-type abilities trace different trajectories;

peak ages are scattered across decades depending on the ability; individual differences

within age groups and the overlap of distributions across groups are large; and crosssectional

and longitudinal estimates remain opposed on the onset timing of decline. Under

this evidentiary situation, an institution that decides role allocation and exit by the single

variable of chronological age lacks cognitive-scientific grounding. The rational alternative is

to allocate roles dynamically on the basis of measured cognitive characteristics, accumulated

domain experience, health status, and the person's own intent. This is the section's set-up for

the definition of Ageless Management formalized in Section 5 (Definition 1), and

compensation of Gf-type components by AI (Section 2.3) is examined from the next section

onward as the technical condition that makes this allocation shift feasible.

3. Work and Brain Health: The Evidence

This section examines the empirical foundation on which this paper's theory may rest — the

evidence on the relationship between work, retirement, cognitive function, and health. The

contour of the conclusion can be stated in advance. The evidence in this field is not

monolithic. While multiple estimates exist suggesting that cognitive decline accelerates after

retirement, estimates of equal methodological standing exist suggesting that retirement

instead improves health. This paper therefore does not adopt the proposition "work protects

the brain" as a premise. The purpose of this section is to survey precisely the reach and limits

of the evidence, and to identify which questions are closed and which remain open. As

discussed below, the open question is not "does work protect the brain?" but "what kind of

work could protect it?" — and this paper's theory (Section 5) is designed as an answer to this

open question.

3.1 The Evidence Since Mental Retirement

The idea of "use it or lose it" is the most widely circulated piece of common wisdom about

cognitive aging. What must be noted here is that this is the name of a hypothesis, not the

name of an established finding (recall the controversy over cognitive reserve discussed in

Section 2.4). Work has been regarded as a natural domain of application for this hypothesis:

for most people, work is the largest single activity that simultaneously supplies intellectual

Working Paper | Ageless Management in the AI Era 27

stimulation, social contact, time structure, and a sense of role. If the hypothesis holds for

work, then retirement is a systematic loss of cognitive stimulation and should hasten

cognitive decline. This subsection surveys the body of research that has tested this prediction,

attending to the strength of the identification strategies and the direction of the results. To

state it in advance: the evidence is split between directions that support the hypothesis and

directions that do not, and the pattern of that split itself carries important information for

this paper's theory.

The starting point of the research lineage that treats the relationship between work and

cognitive function with the identification strategies of economics is the "Mental Retirement"

paper of Rohwedder & Willis (2010). They compared internationally the 2004 waves of the US

HRS (about 20,000 people), the English ELSA (about 9,000), and SHARE covering 11 European

countries (1,000–3,000 per country), using immediate and delayed recall of 10 words (0–20

points) as the cognitive measure. The share of people aged 60–64 not in paid work varies

greatly by country, from about 30% in the United States and Denmark to 80–90% in France

and Austria. This difference was produced by each country's pension, tax, and disabilitybenefit

institutions, and can be regarded as determined independently of individual cognitive

ability. Using this institutional variation as an instrumental variable, they estimated a causal

effect in which early retirement lowers memory scores by about 4.7 points (5.7 points in the

specification controlling age in one-year increments). The authors summarize that "early

retirement appears to have a significant negative and quantitatively important causal effect

on the cognitive ability of people in their early 60s."

This effect size, however, cannot be taken at face value. The standard deviation of the

cognitive scores in the sample is about 3.3, so a 4.7-point drop corresponds to more than 1.4

times that. Subsequent research has not replicated this magnitude, and some estimates —

such as Coe et al. (2012), discussed below — could not confirm even an effect in the same

direction. Being a country-level comparison, it cannot fully remove country-specific

confounders such as educational systems, test administration, and cultural differences. That

is, Rohwedder & Willis (2010) was the watershed that brought the problem of causal

identification into this field, and at the same time it is a study whose effect size should be

cited only with the reservation that "later research has produced smaller estimates, and some

estimates of zero."

There is other evidence in the same direction. Bonsang, Adam & Perelman (2012), applying

instrumental variables based on Social Security eligibility ages and individual fixed effects to

the US HRS panel, reported that retirement has a significant negative effect on cognitive

function, and that the effect appears not immediately after retirement but with a delay.

Dufouil et al. (2014) analyzed linked claims and pension data on 429,803 retired selfemployed

workers in France, and reported that each additional year of age at retirement was

associated with a dementia hazard ratio of 0.968 (95% CI 0.962–0.973), equivalent to a risk

reduction of about 3.2% per year. However, the authors themselves state explicitly that

further evidence is needed to assess whether this association is causal, and that reverse

Working Paper | Ageless Management in the AI Era 28

causation — prodromal cognitive decline leading to earlier retirement — cannot be ruled out.

This paper likewise does not cite it as a causal effect.

As longitudinal evidence tracking the same individuals across the transition into

retirement, there is Xue et al. (2018) from the British civil-service cohort Whitehall II.

Following 3,433 people for up to 14 years each before and after retirement (up to 28 years in

total), verbal memory declined about 38% faster after retirement than before, after adjusting

for age-related decline. Abstract reasoning and verbal fluency, by contrast, showed no

significant change across retirement; the effect was specific to memory. A secondary finding

is also suggestive: while employed, higher employment grade was associated with slower

memory decline, but after retirement this protective association disappeared, and the rate of

decline became equivalent regardless of grade. Note that "38% faster" is a relative value; the

decline in absolute terms is gradual. Moreover, the sample is limited to white-collar civil

servants, and the endogeneity of retirement timing remains, so this result too cannot be

asserted as causal.

What about the level of systematic review? Meng, Nexø & Borg (2017) systematically

reviewed longitudinal studies on retirement and age-related cognitive decline, and noted that

only 7 studies met the inclusion criteria (1 more added in an updated search), of which 4

derived from the same cohort (the US HRS) and thus had low independence. The results

divide by domain. For fluid intelligence (Gf) the evidence is conflicting: two high-quality

studies pointed toward retirement slowing decline, while one of moderate quality pointed

toward acceleration. Only two studies addressed crystallized intelligence (Gc), yielding no

more than weak evidence that decline accelerates only for those retiring from jobs high in

interpersonal complexity. The authors' overall assessment is that no conclusion can be drawn

that retirement uniformly accelerates cognitive decline; the evidence is mixed and there are

large research gaps. It would be an error to cite this review as grounds that "retirement has

been established as bad for cognition," and equally an error to cite it as grounds that

"retirement has been established as harmless."

Randomized controlled trial (RCT) evidence that goes beyond the limits of observational

research exists not for work itself but for structured productive social engagement.

Experience Corps is a US program in which older adults serve 15 hours per week in

elementary schools providing reading support, library support, and classroom support; Fried

et al. (2004) presented its design philosophy as a social model of health promotion. Carlson et

al. (2008), in exploratory analyses of a pilot RCT (149 people), reported that the participation

group improved in executive function and memory relative to the waitlist control group, with

improvements of 44–51% among those with impaired executive function at baseline (while

the same stratum in the control group declined). Carlson et al. (2009), in a preliminary fMRI

study of 17 women aged 65 and over, reported that participation was accompanied by

improved cognitive function and significant changes in prefrontal activity patterns.

Furthermore, the imaging substudy of the Baltimore Experience Corps Trial (Carlson et al.

2015) followed 111 people (58 intervention, 53 control; mean age 67.2; predominantly African

Working Paper | Ageless Management in the AI Era 29

Americans from low-income areas) for 24 months and reported that, whereas annual brain

atrophy of 0.8–2% normally occurs after age 65, cortical and hippocampal volumes in men in

the intervention group increased by 0.7–1.6% over two years. Women in the intervention

group showed only slight increases that did not reach statistical significance, and women in

the control group declined by about 1% over 24 months. Participants with larger volume

increases also showed larger improvements on memory tests. Because these are RCTs,

"improved" can be written — but two reservations are required. First, the participants were a

restricted low-income, urban population, and generalization requires caution. Second, this is

the effect of 15 hours per week of structured volunteering, not of employed labor, and it is not

direct evidence of the effect of continued employment. This point is taken up again in Section

3.3.

The evidence above is organized in Table 2. Three notes on how to read the table. First, the

evidence grades classify the strength of identification strategies, not a ranking of the

credibility of results. Given that instrumental-variable studies of the same A− grade estimate

opposite signs, a high grade does not guarantee agreement of conclusions. Second, the

outcomes differ across studies. Memory scores, dementia registration, brain volume, and selfreported

health are not mutually substitutable measures, and effects on "cognitive function"

and effects on "health" must be read separately. Third, the restrictions of the study

populations (civil servants only, the self-employed only, low-income urban areas only, 14

Japanese municipalities only) are not footnote-level details but direct determinants of

external validity. Table 2 should be read not as a list of answers to the single question "work

and brain health," but as a map of estimates under different populations, measures, and

identification strategies.

Table 2 Principal evidence on work, retirement, cognitive function, and health

Study Design Main result

Evidence

grade

Rohwedder &

Willis (2010)

International comparison of

HRS, ELSA, SHARE.

Institutional differences in

pensions etc. as

instrumental variables

Estimated that early retirement lowers

memory scores (out of 20) by about 4.7

points. Effect size exceeds 1.4 times the

standard deviation (about 3.3); repeatedly

criticized as too large

A−

Bonsang et

al. (2012)

US HRS panel. Social

Security eligibility-age IV +

individual fixed effects

Retirement has a significant negative

effect on cognitive function. Effect

reported to appear with a delay

A−

Coe et al.

(2012)

US HRS men. Employer

early-retirement "windows"

used as IV

Concluded that the negative simple

correlation cannot be called causal. No

clear relationship between retirement

duration and cognition for white-collar

workers; for blue-collar workers, a

positive relationship instead

A−

Insler (2014) Retirement significantly positive for

health. Mediated by improved health

A−

Working Paper | Ageless Management in the AI Era 30

US HRS. Subjective

probability of continued

work used as IV

behaviors such as reduced smoking and

increased exercise

Eibich (2015) German SOEP. Regression

discontinuity using

eligibility-age

discontinuities

Retirement improves health in the long

run and reduces healthcare utilization.

Mechanisms: relief from stress, sleep,

exercise

A−

Dufouil et al.

(2014)

429,803 retired selfemployed

workers in

France. Observational study

of linked claims and

pension data

Each additional year of retirement age

associated with dementia hazard ratio

0.968 (95% CI 0.962–0.973). Authors

themselves state causality is unestablished

B

Xue et al.

(2018)

Whitehall II, 3,433 people.

Longitudinal follow-up up

to 28 years around

retirement

Verbal memory declines about 38% faster

after retirement (relative value).

Reasoning and fluency unchanged.

Protective association of employment

grade during employment disappears

after retirement

B

Meng et al.

(2017)

Systematic review of 7+1

longitudinal studies (no

meta-analysis)

Evidence mixed. Gf conflicting; for Gc only

weak evidence of accelerated decline after

retirement from jobs high in interpersonal

complexity

S

van Ours

(2022)

Narrative review of recent

causal-inference research

On average, mental health improves with

retirement, cognitive skills decline,

mortality unaffected. Effects highly

heterogeneous by occupation,

voluntariness, etc.

N

Carlson et al.

(2008)

Exploratory analysis of

Experience Corps pilot RCT

(149 people)

Participation group improved in executive

function and memory relative to controls.

44–51% improvement in the baselineimpaired

stratum

A (small)

Carlson et al.

(2015)

Baltimore Experience Corps

Trial imaging substudy (111

people, 24 months)

Cortical and hippocampal volumes in

intervention-group men increased 0.7–

1.6% over 2 years. No significant

difference in women. Restricted

population

A

Parker et al.

(2020)

Multistate life-table

estimation, UK ELSA, 15,284

people

Healthy working life expectancy at age 50

about 9 years. Men 10.9 years, women 8.3

years; regional gap about 4.5 years

B

(descriptive)

Takeuchi et

al. (2024)

JAGES, 48,221 people, 6-year

longitudinal study (Japan)

Relative to retirees, agricultural workers

show favorable associations such as

dementia OR 0.45 and mortality OR 0.68.

Healthy worker effect not removed

B

Note: Evidence grades follow the classification of the evidence notes. A = RCT or valid natural experiment /

instrumental variables (A− with identification reservations), B = longitudinal observational study (causation cannot be

claimed), S = systematic review, N = narrative review. Note that even within the same grade the signs of the estimates

do not agree. All are peer-reviewed publications.

Working Paper | Ageless Management in the AI Era 31

3.2 The Difficulty of Causal Identification — Health Selection, Reverse

Causation, and the Voluntariness of Retirement

Before interpreting the evidence in Table 2, the central identification problem of this field

must be made explicit: "does one work because one is healthy, or is one healthy because one

works?" The association "workers are healthy" in observational data is contaminated by at

least two mechanisms. The first is health selection (the healthy worker effect). Because

healthier people keep working longer, a positive association between work and health arises

even without any effect of work. The second is reverse causation. The prodromal phase of

dementia is held to extend over 10 years or more, and if prodromal decline hastens

retirement, then what looks like "cognitive decline after retirement" is in fact "retirement

caused by cognitive decline." It is for this reason that Dufouil et al. (2014), even after

reporting in sensitivity analyses that the association remained when retirements immediately

preceding onset were excluded, still did not claim causality. A third difficulty is the

voluntariness of retirement. Retirement chosen by oneself and forced retirement due to

deteriorating health or dismissal may carry different implications for subsequent health.

The problem of voluntariness deserves a little more elaboration. As van Ours (2022)

organizes it, the effects of retirement are highly heterogeneous depending on whether

retirement is voluntary or involuntary. This heterogeneity is an obstacle to identification and,

at the same time, itself a substantive finding. Voluntary retirement usually occurs in a state in

which the transition to post-retirement life has been prepared. Forced retirement occurs as

the simultaneous loss of role, income, and social contact. If the two are lumped into the single

treatment "retirement" and an average effect is estimated, which sign emerges depends on

the composition of the two within the sample. What matters from the standpoint of Ageless

Management is that uniform age-based exit compulsion, such as mandatory retirement

(teinen) systems, is precisely an apparatus that institutionally mass-produces this "forced

retirement." If the effect of retirement depends on voluntariness, the question to ask is not

whether retirement is good or bad, but who decides the timing and form of exit. This point is

taken up again in the institutional comparison of Section 7 (international differences in

mandatory retirement and pension systems).

The standard response to these problems is the quasi-experiment: using institutional

variation determined independently of individual health and cognition — pension-eligibility

ages or the offer of early-retirement incentives — as instrumental variables (IV) or regression

discontinuities (RD). But the important point is that the reality of this field is that conclusions

can flip depending on the choice of IV. Coe et al. (2012) used as an instrument the earlyretirement

"windows" that employers are legally required to offer on a non-discriminatory

basis unrelated to cognitive ability, and concluded that the association between retirement

and cognitive decline seen in simple correlations cannot be called causal. Moreover, for

white-collar workers there was no clear relationship between retirement duration and

cognition, while for blue-collar workers a positive relationship between retirement duration

Working Paper | Ageless Management in the AI Era 32

and cognition was estimated instead. This is an estimate squarely at odds with Rohwedder &

Willis (2010) and Bonsang et al. (2012), and this paper presents the two as a pair of evidence

of equal standing.

Why do methods of equal standing produce opposite answers? One reading is that different

instruments identify effects in different subpopulations. The people who decide to retire in

response to institutional variation in pension-eligibility ages and the people who respond to

an employer's early-retirement incentive are not the same population in occupation, health

status, or voluntariness of retirement. What the instrumental-variables method estimates is a

local effect among those who responded to the institutional variation in question, not a

universal "effect of retirement." If so, the disagreement among estimates can be read less as

evidence of methodological defect than as evidence that the effect of retirement itself differs

by population and context. Indeed, the occupation-specific results of Coe et al. (2012) — no

relationship for white-collar, positive for blue-collar — show that structure invisible so long

as one asks about average effects appears the moment one asks about heterogeneity. This

reading foreshadows the argument of Section 3.3.

Evidence on the opposite side has also accumulated for health outcomes other than

cognition. Insler (2014), using as an instrument respondents' previously reported subjective

probability of working past age 62/65, reported that retirement is significantly positive for

health, and that reductions in smoking and increases in exercise enabled by increased free

time mediate the positive effect. Eibich (2015), using the discontinuity in financial incentives

generated by eligibility ages in the German SOEP, estimated that retirement improves health

in the long run and reduces healthcare utilization. As mechanisms he cites relief from workrelated

stress, increased sleep, and increased frequency of exercise. In the reading of van

Ours (2022), who organizes the recent causal-inference literature, the empirical results vary

greatly across studies, but on average, mental health improves with retirement, cognitive

skills decline the longer the retirement duration, and mortality is largely unaffected. And the

effects are highly heterogeneous by individual attributes, occupation, institutions, and the

voluntariness of retirement — so much so that the author himself describes the field as prone

to "not seeing the forest for the trees." The policy implication van Ours draws is neither

uniform extension of working lives nor uniform support for early retirement, but the

expansion of individual flexibility of choice over the timing of retirement. This implication

points in the same direction as this paper's Definition 1, which rejects uniform role allocation

by chronological age and makes the person's own intent a constituent of allocation decisions.

One note is also in order on how to read the numbers. Many of the effects cited in this

section are relative values. The "38% faster decline" of Xue et al. (2018) is a relative

comparison of rates of decline, and the decline in absolute terms is gradual. The "about 3.2%

per year" of Dufouil et al. (2014) is a conversion of a hazard ratio, and given the baseline

dementia prevalence (2.65% in that cohort), the difference in absolute risk is small. Relative

values are useful for conveying the existence of an effect, but practical decision-making

requires absolute values, and this paper does not circulate relative expressions on their own.

Working Paper | Ageless Management in the AI Era 33

This paper does not paper over this evidentiary situation. The most honest summary of the

current state is as follows. Mental and physical health can improve with retirement

(especially from high-strain, physical work). For cognitive function, multiple estimates

suggest accelerated decline after retirement, but estimates of no effect or improvement exist

with methods of equal standing, and the sign changes with occupation and the voluntariness

of retirement. Therefore it cannot be written monolithically that "work is good for health."

This paper takes up this controversy as Hypothesis H1. That is, rather than presupposing a

brain-health effect of work, it presents the effect as a testable hypothesis with specified

conditions, and designs it — identification strategy included — in Section 9.

Hypothesis H1 (Brain-Health Hypothesis)

Super-seniors engaged in Gc-exercising, AI-complemented roles show a lower rate of

cognitive decline (MoCA, etc.) than same-age peers who have exited work, mediated by

the maintenance of cognitive engagement. Identification strategy: a quasi-experiment

using exogenous variation from pension-system reforms or mandatory-retirement rules

as instrumental variables, or exploiting exogeneity in reasons for working. Treatment of

health selection and reverse causation to be specified explicitly.

3.3 The Quality of Work as a Moderating Variable

So long as the body of evidence in Table 2 is read as a "contest of average-effect estimates,"

what one obtains is deadlock. But rereading it with the heterogeneity of effects as the subject,

a consistent structure comes into view. The variable that flips the sign is not whether one is

working, but what kind of work it is, and why one left. In Coe et al. (2012), cognition moved in

the direction of improvement after retirement for blue-collar workers. Meng et al. (2017)

found weak evidence of accelerated Gc decline only among those who retired from jobs high

in interpersonal complexity. In Xue et al. (2018), the protective association of employment

grade during employment disappeared after retirement. The mechanism in Eibich (2015) was

relief from stress. These are consistent with the reading that the meaning work has for

cognition depends strongly on the cognitive content and load of the job. Exit from depleting

work can improve health; exit from cognitively rich roles can be the loss of stimulation that

had been maintained.

This paper formalizes this reading as the moderating variable "quality of work." The key

construct is cognitive engagement — active involvement in roles containing Gc-type

components such as evaluation, contextual interpretation, and interpersonal judgment. If a

brain-health effect of work exists, what carries it is not the length of working hours or the

existence of an employment contract but the maintenance of this cognitive engagement —

that is this paper's theoretical wager. That is: if a brain-health effect of work exists, it is

mediated not by the length of working hours but by the maintenance of cognitive

engagement through occupying Gc-type roles. This claim is formally stated in Section 5 as

Working Paper | Ageless Management in the AI Era 34

Proposition 7 (the cognitive-engagement pathway), with a refutation condition attached. The

point here is that Proposition 7 stands as a candidate hypothesis that explains the deadlock of

Sections 3.1–3.2. The sign of the average effect of work fails to settle, on this reading, because

the category "work" mixes cognitively rich roles with depleting ones. However, this mediating

structure is itself untested, and the need for testing by mediation analysis is stated explicitly

in the refutation condition of Proposition 7.

This formalization connects to the cognitive-scientific foundations of Section 2. As seen

there, fluid intelligence (Gf) tends to decline throughout adulthood, while crystallized

intelligence (Gc) can be maintained or rise into old age. Defining cognitive engagement as

engagement with Gc-type components carries a double implication for work in old age. First,

roles that use maintained abilities permit continued engagement and can remain sources of

stimulation. Second, roles that overtax declining abilities are depletion before they are

stimulation, and become sources of exit pressure. The single piece of Gc-side evidence found

by Meng et al. (2017) — weak evidence that decline accelerates only after retirement from

jobs high in interpersonal complexity — is at least not inconsistent with this reading, since

jobs high in interpersonal complexity are the typical case of jobs weighted toward the Gc-type

components of contextual interpretation and interpersonal judgment. However, this is a posthoc

consistency check, not a test. Whether this interpretation is right is precisely what

Hypothesis H1 must ask through mediation analysis.

Seen from this standpoint, the significance of Experience Corps is more than "an RCT of

older volunteers." Experience Corps is a program that explicitly designed the quality of the

role. Within the structured time of 15 hours per week, participants took on a role — reading

support for schoolchildren — that demands interpersonal judgment and contextual

interpretation and whose outcomes accrue to others. Fried et al. (2004) called this a "social

model" of health promotion because the intervention instrument was the provision of a

meaningful social role, not an individual prescription of exercise or cognitive training. That

its RCT arms showed, albeit in restricted populations, improvements in executive function

and memory (Carlson et al. 2008) and increases in brain volume among men (Carlson et al.

2015) suggests that the designed quality of a role can act on cognition in old age. To repeat,

this is not evidence about employed labor. Extrapolation to employment requires a bridge

through the higher-order concept of "productive social engagement," and the strength of that

bridge is itself an object of testing. But for this paper's theory this distinction is, if anything,

convenient. What this paper seeks to design is not the extension of employment contracts but

the allocation of roles within a multigenerational ecosystem (Definition 7) not limited to the

boundary of employment — and Experience Corps is precisely a case of a designed, nonemployment

role.

3.4 Healthy Working Life Expectancy — Health as a Precondition

Finally, this section confirms the constraint that determines the work-health relationship

from the opposite side: the evidence on "the number of years one can work in health in the

Working Paper | Ageless Management in the AI Era 35

first place." Parker, Bucknall, Jagger & Wilkie (2020), using multistate life-table estimation on

the English ELSA (2002–2013, 15,284 people aged 50 and over, linked to NHS mortality

records), estimated healthy working life expectancy (HWLE) at age 50 — the expected

number of years spent both healthy and in work — at about 9 years. The implication of this

figure lies less in the mean than in the distribution. Men 10.9 years versus women 8.3 years.

The North East about 4.5 years shorter than the South East; 6.8 years in the most deprived

areas. Manual workers about 1 year shorter than non-manual workers. And HWLE from age

50 falls short of the years remaining until the state pension age. That is, "working in health

until pension age" is not guaranteed even on average, and the shortfall is systematically

skewed by sex, region, occupation, and income. This is a descriptive estimate, and the figures

move with the definitions of health and work, but the implication for policy and management

design is clear. Any framework that discusses the extension of working lives must build in the

fact that health is its precondition and is unequally distributed.

At the same time, it must be noted that healthy working life expectancy is not a fixed

quantity. The estimate of Parker et al. (2020) is a description that takes current working

environments, job design, and health distributions as given, and if any of these change, the

figures can change too. The fact that manual workers' HWLE is about 1 year shorter suggests

that it is the physically and cognitively depleting components of jobs that determine the

number of workable years. If so, changing the composition of the job bundle —

complementing the depleting components with technology and shifting the center of gravity

of roles toward components that use maintained abilities — becomes an intervention

hypothesis that could act on healthy working life expectancy itself. This is the public-health

restatement of the AI complementation of Gf-type components (Proposition 2) developed

from Section 4 onward. However, whether role redesign actually extends HWLE is untested,

and this paper presents it only as a design goal.

What of the Japanese evidence? From the Japan Gerontological Evaluation Study (JAGES),

Takeuchi et al. (2024) followed 48,221 people aged 65 and over in 14 municipalities for 6 years

and reported that, compared with retirees, agricultural workers showed favorable

associations: dementia odds ratio 0.45, care-need risk ratio 0.64, severe care-need risk ratio

0.65, healthy-life-expectancy-loss risk ratio 0.69, and mortality odds ratio 0.68. Nonagricultural

workers also showed similar risk reductions across all outcomes, and the neverworked

group had higher risk than retirees. The authors cite the high physical activity of

rural areas among candidate mechanisms. However, this study has no instrumental variable,

and the authors themselves note the healthy worker effect, respondent bias in selfadministered

surveys, and the restriction to a Japanese population. This result is therefore

kept as a description of association — "favorable associations between continued work and

health outcomes are observed in large Japanese longitudinal data as well" — and is not cited

as causal, since the selection whereby healthier people keep working could explain much of

the association. In addition, Japan also has a prospective study of the association between

work and care-need onset by frailty status (Fujiwara et al. 2023), and the association between

Working Paper | Ageless Management in the AI Era 36

work and late-life health has been observed repeatedly in Japanese data as well. But all of

these are observational studies, and the identification problem is of the same form as in the

Anglo-American literature.

The evidence on healthy working life expectancy imposes two disciplines on this paper.

First, Ageless Management cannot presuppose that "everyone can work indefinitely." The

years one can work in health are finite and unequally distributed along socioeconomic lines.

A design that does not render invisible, once again, those who cannot work or can no longer

work is an internal condition of this paper's theory (this survivorship-bias problem is

revisited as self-critique in Section 10). Second, since pathways exist by which continued

work itself depletes health (recall the mechanisms in Insler 2014 and Eibich 2015), an

occupational safety and health standard that protects participants' brain capital from

depreciation — Brain Safety (Definition 8) — is not an appendage of the theory but an

essential component. This point is developed in Section 8.

3.5 Conclusion of This Section — Identifying the Open Questions

The examination in this section can be summarized as follows. First, "work protects the

brain" cannot be asserted on current evidence. IV estimates and longitudinal evidence

suggesting post-retirement cognitive decline (Rohwedder & Willis 2010; Bonsang et al. 2012;

Xue et al. 2018) coexist with estimates of equal standing suggesting no effect or health

improvement (Coe et al. 2012; Insler 2014; Eibich 2015), and the conclusion of the systematic

review (Meng et al. 2017) is likewise "mixed." Second, this deadlock is nevertheless not

uninformative. The variables that divide the sign of the effect are occupation, the

voluntariness of retirement, and the cognitive content of the work — suggesting that only

"continued work conditional on voluntariness and quality of work," not "uniform extension

of working lives," can be consistent with the evidence. Third, the RCT that designed the

quality of the role (Experience Corps) showed that designed productive social engagement

can act on cognitive function and brain structure in restricted populations.

Fourth, the years one can work in health are themselves finite and unequally distributed

(Parker et al. 2020), and although favorable associations between work and health outcomes

are observed in large Japanese longitudinal data as well (Takeuchi et al. 2024), causal

evidence purged of health selection does not exist.

Accordingly, what is open is not the average-effect question "does work protect the brain?"

but the design question "what kind of work could protect it, and for whom?" This paper's

theory is constructed as an answer to this question. Section 4 examines the evidence on how

AI changes the cognitive content of work, and Section 5 presents Proposition 7, with cognitive

engagement as the mediating pathway, along with the set of propositions including dynamic

allocation to Gc-type roles. And the brain-health effect of work itself is registered not as an

assertion but as Hypothesis H1 on the map of verification in Section 9. This is as far as the

Working Paper | Ageless Management in the AI Era 37

evidence permits — and as far as the evidence permits, this section goes: that is the

conclusion of this section.

4. AI and Experience: What Is Compressed and What Is Not

The preceding sections confirmed that age-related change in cognitive abilities is not

monolithic (Section 2) and that the relationship between work and brain health depends on

the moderating variable of quality of work (Section 3). As the third pillar of this paper's

theory, this section organizes the empirical research on how generative AI is reorganizing the

economic value of "experience." There are three questions. First, what aspects of experience

does AI compress? Second, what does it not compress? Third, which abilities does dependence

on AI depreciate?

To anticipate the conclusion, the picture drawn by the evidence is as follows. For the

production of deliverables within AI's zone of competence, those with less experience gain

the most, and experience differences are compressed (Section 4.1). Outside AI's zone of

competence, by contrast, users' performance actually deteriorates, and human oversight of AI

output can fail systematically (Section 4.2). Furthermore, assisted performance gains do not

guarantee unassisted ability, and everyday dependence on AI can depreciate unaided skill

(Section 4.3). These three findings suggest the relative scarcification of the functions of

evaluation, verification, and oversight, rather than production. However, the "expertise"

measured in these empirical studies is domain experience and skill, not chronological age.

This conceptual separation is the pivot of this section (Section 4.4); it leads to the

reinterpretation of the headwind data surrounding older workers and AI (Section 4.5), and to

the empirical foundation — and the demarcation of the limits — of the definitions and

propositions presented in Section 5.

4.1 Empirical Evidence on Experience Compression

Empirical evidence on the effects of generative AI on job performance has accumulated

rapidly since 2023. Among the most credible findings is the compression of differences in

experience and ability. Brynjolfsson, Li & Raymond (2025) conducted a quasi-experiment

exploiting the staggered rollout of a generative AI conversational assistant among 5,172

agents at a customer-support firm. Average productivity, measured as resolutions per hour,

rose by 15%. The effect was not uniform, however. Workers with less experience and skill

showed improvements of roughly 30% across all productivity measures, while the most

experienced and highest-skilled workers saw only small gains in speed, with a slight decline

in conversation quality.

The degree of compression can be expressed in the language of the experience curve. AIusing

agents with two months of tenure performed on par with non-using agents with more

than six months of tenure. The authors interpret this as a shortening of the experience curve

by about four months. The mechanism has also been identified. The AI was trained on the

Working Paper | Ageless Management in the AI Era 38

conversation texts of top performers, and it presents as recommendations their tacit

behavioral patterns — clarifying questions, active listening, avoidance of escalation, and

adjustment of tone. After deployment, the communication patterns of low-skill workers

converged toward those of high-skill workers. What is being compressed here, in substance,

is the extraction and transfer of experts' tacit knowledge.

The study also reports two ancillary findings. First, even during AI system outages, workers

with longer AI experience — especially those who had followed the recommendations

faithfully — maintained higher productivity than their own pre-deployment levels. Part of AI

use can take hold as learning rather than mere dependence. This point is revisited in Section

4.3, paired with Budzyล„ et al. (2025). Second, after deployment, customer sentiment improved

and turnover declined (especially the retention of new hires). Note, however, that the study

covers routine tasks in a single firm and a single occupation, and age is not among the axes of

analysis.

Noy & Zhang (2023), in a preregistered online experiment with 453 college-educated

professionals, tested the effects on occupation-related writing tasks. The ChatGPT group took

40% less time and received quality ratings 18% higher. For this paper, the core is the change in

the distribution. In the control group, the correlation between performance on the first and

second tasks was 0.49, whereas in the treatment group it fell to 0.25. This is compression in

the sense that initial performance differences shrink by roughly half under treatment, and

participants with lower initial performance benefited more. In a follow-up survey two weeks

later (82% response rate), 33% of the treatment group had continued using the tool in their

actual work; among those with no prior experience, the figures were 26% in the treatment

group versus 9% in the control group (p=0.048). As a limitation, the tasks were one-off and

brief, and learning and long-term skill formation were not measured.

As a large-scale experiment in a near-field setting, Dell'Acqua et al. (2023) conducted a

preregistered field experiment using GPT-4 with 758 consultants at the Boston Consulting

Group (about 7% of the firm's individual-contributor level). On the 18 tasks designed to fall

within AI's capabilities ("inside the frontier"), tasks completed rose by +12.2%, speed by

+25.1%, and quality by more than 40% relative to the control group. The compression pattern

was replicated. Performers with below-average baseline performance improved by +43%

relative to their own baseline, and above-average performers by +17%. Gains are larger

toward the bottom, but the top gains as well; the summary "no effect at the top" is wrong. The

study's conceptual contribution lies in its formalization of the "jagged technological frontier":

the boundary between what AI does well and does poorly is hard to see in advance, and

discerning that boundary itself becomes a new human skill. This formalization leads directly

to the problem of oversight in the next subsection.

The same pattern appears in software development. Peng et al. (2023) randomly assigned

95 professional developers to treatment and control and reported that use of GitHub Copilot

shortened completion time on an HTTP-server implementation task to 71.17 minutes versus

Working Paper | Ageless Management in the AI Era 39

160.89 minutes (−55.8%). Heterogeneity analyses suggest larger benefits for developers with

fewer years of experience, developers with heavier coding loads, and older developers (aged

25–44). The study is an unrefereed vendor-affiliated study, with the limitations of a low

completion rate and a single task, but it replicates the compression-type result on the

experience axis while standing as one of the few data points suggesting that, on the age axis,

"the younger, the greater the gain" may not hold; it is referenced again in Section 4.4.

What is robust across these four studies is the pattern that, on tasks within AI's zone of

competence, the lower-performing and less experienced gain the most. Translated into the

language of economics, this is downward pressure on the experience premium in deliverable

production. The market value of the production-side differences that experience has

incrementally conferred — deliverable quality, working speed, procedural knowledge —

begins to decline once AI can substitute for and transfer them in units of several months.

Here the content of the compressed abilities must be examined precisely. What is being

compressed is not only Gf-type components. The behavioral patterns whose transfer

Brynjolfsson, Li & Raymond (2025) confirmed — clarifying questions, active listening,

avoidance of escalation, adjustment of tone — include Gc-type components belonging to

interpersonal judgment in the classification of Definition 2 (Section 5). That is, the mapping of

Section 2 — "what generative AI primarily substitutes for is the Gf-type load" — must be

qualified as a mapping in the context of production. In the context of production, so long as

they can be formalized as training data, even the interpersonal and contextual Gc-type

components become objects of transfer. The axis along which the boundary of compression is

drawn is therefore not the type of ability (Gf-type or Gc-type) but the distinction of function —

differences in production, or differences in verification. Proposition 1 of Section 5 (the

marginal value shift) formalizes, as the obverse of this downward pressure, the rise in the

relative marginal value of Gc-type tasks — above all evaluation- and audit-type tasks (those

that exercise the contextual interpretation and interpersonal judgment of Definition 2 in

verification). But what that "uncompressed function" is, and who possesses it, cannot be

derived from the compression evidence. That is the subject of the following two subsections.

Table 3 Compression of experience and ability differences by generative AI: key empirical studies

Study Task and context Main effects

Group benefiting

most Evidence grade

Brynjolfsson,

Li & Raymond

(2025)

Customer-support

work. 5,172 agents;

quasi-experiment

with staggered

rollout

Resolutions +15%. AIusing

agents with 2

months' tenure on par

with non-users with

over 6 months' tenure

(experience curve

shortened by about 4

months)

Workers with less

experience and

skill (about +30%).

Top tier: only small

speed gains, with a

slight decline in

quality

A (peerreviewed;

largescale

field quasiexperiment)

Noy & Zhang

(2023)

Occupation-related

writing. 453 collegeeducated

Time −40%; quality

+18%. Cross-task

performance

A (peerreviewed,

but a

Working Paper | Ageless Management in the AI Era 40

professionals;

preregistered online

experiment

correlation fell from

0.49 to 0.25

(compression of initial

differences)

Participants with

lower performance

on the first task

laboratory-style

task)

Dell'Acqua et

al. (2023)

18 consulting tasks

(inside the frontier).

758 BCG consultants;

GPT-4

Completions +12.2%;

speed +25.1%; quality

over +40%

Below-average

performers (+43%).

Above-average

performers also

gained +17%

B+ (workingpaper

preprint;

large-scale

preregistered

field

experiment)

Peng et al.

(2023)

HTTP-server

implementation task.

95 professional

developers; GitHub

Copilot; RCT

Completion time 71.17

vs. 160.89 minutes

(−55.8%)

Developers with

fewer years of

experience,

developers with

heavier coding

loads, and older

developers (aged

25–44)

B (unrefereed;

vendor-affiliated

study; single

task)

Note: All effects are relative to the control condition or to performers' own baselines. The figures for Noy & Zhang

(2023) follow the version published in Science. The figures for Dell'Acqua et al. (2023) follow the working-paper

version (an Organization Science version appears to have been published in 2025, but this paper uses the verified

working-paper figures). Evidence grades are this paper's rating (A = peer-reviewed; B = preprint or technical report; C

= grey literature; D = essay; +/− indicate relative rating within a grade).

In summarizing these results, the subject of compression must be delimited precisely. What

is compressed is the experience difference in deliverable production, under

contemporaneous assistance, within AI's zone of competence. Judging the zone of

competence, verifying the output, and internalizing the learning are each separate problems,

and, as the following two subsections show, the compression evidence does not carry over to

them. This delimitation is the empirical basis of the first half of Proposition 3 of Section 5 (the

non-compressibility of experience), namely that "contemporaneous AI assistance compresses

differences in production (deliverable quality, working speed, procedural knowledge)." The

second half — that it does not compress differences in verification (differences in experiential

audit capacity) — is a theoretical claim that at present lacks direct evidence of the same

standard, and its status is clarified in Section 4.4.

4.2 Oversight Failure

The compression story has a reverse side. Dell'Acqua et al. (2023) also ran tasks designed so

that AI would be prone to error ("outside the frontier"). On these tasks, the probability that

the AI-using group reached the correct answer was 19 percentage points lower than the

control group. It is not that quality fell by 19%; the probability of reaching the correct answer

fell by 19 points. The same population that gained over +40% in quality inside the frontier

clearly deteriorated outside the boundary. Because the boundary is hard to see in advance,

discerning how much to entrust to AI — that is, oversight of AI output — emerges as the

human-side function that separates outcomes.

Working Paper | Ageless Management in the AI Era 41

Can humans, then, perform that oversight well? Dell'Acqua's (2022–2023) solo-authored

working paper "Falling Asleep at the Wheel" offers cautionary evidence on this point. 181

professional HR recruiters each evaluated 44 résumés (7,964 in total; 5,184 after attention

checks), with the quality of the AI provided randomly assigned: a "near-perfect AI" disclosed

as roughly 99% accurate, a "high-quality AI" at 85%, a "low-quality AI" at 75%, and a no-AI

control. The central finding is that evaluators given the low-quality AI were more accurate

than those given the high-quality AI (accuracy improvement over control on a 10-point scale:

+0.314 for the low-quality AI versus +0.103 for the high-quality AI). The low-quality-AI group

examined each résumé for about 8.8–10 seconds longer and invested significantly more effort

(p<0.01). When the AI is understood to be excellent, humans do not raise their cognitive effort

— they "fall asleep at the wheel." This relationship is non-monotonic, however. The condition

disclosed as "near-perfect" performed best (+0.787), and the high-quality-AI group also

improved over the control. The accurate summary is therefore not "the better the AI, the

worse the human" but "human vigilance slackens most under an imperfect but highperforming

AI." Note that, at the time of verification, the study was a working paper not

published in a refereed journal, and this paper's figures are based on a mirrored distribution

copy (evidence grade B−). Even with this caveat, the finding that oversight effort is

endogenously determined by the perceived quality of the AI is consistent with the WP8

framework discussed below.

This finding attaches a condition to the compression evidence of Section 4.1. In

Brynjolfsson, Li & Raymond (2025), faithful adherence to AI recommendations was associated

with high performance because the tasks were routine and within AI's zone of competence.

The same adherence behavior flips to the −19-percentage-point side outside the jagged

frontier, and among evaluators whose oversight effort has slackened, it appears as degraded

detection accuracy. Whether "following the AI pays," in other words, depends on the task's

position, and the task's position is hard to see in advance. Precisely for this reason, the

function of doubting and verifying output cannot be switched off even in environments

where adherence performs well.

Human oversight failure is observed beyond individual studies, at the level of metaanalysis.

Vaccaro, Almaatouq & Malone (2024), through a systematic review and metaanalysis

of 106 experimental studies and 370 effect sizes, reported that human–AI

combinations on average performed below the best of human alone or AI alone (Hedges'

g=−0.23, 95%CI −0.39 to −0.07). The assumption that adding a human improves outcomes does

not, on average, hold.

WP8 of this series (Kadowaki 2026h) theorized this body of oversight-failure evidence

within the framework of "Human on the Loop (HOTL)." Five points of its skeleton are needed

for this section's argument. First, oversight is a consumed resource. Supervisors' attention

and cognitive effort are finite, and one cannot assume they are supplied constantly and free

of charge. Second, the solvency condition. The cognitive resources, time, and cost that can be

devoted to oversight have an upper bound, and oversight can be sustained only within that

Working Paper | Ageless Management in the AI Era 42

range. No solution exists that thickens oversight without limit. Third, alarm fatigue. The more

outputs and alerts there are to check, and the more false alarms are mixed in, the more

sluggish supervisors' responses become. Fourth, automation bias. Humans have a systematic

tendency to over-follow a system's suggestions, and the Falling Asleep at the Wheel findings

can be read as the effort-allocation version of this. Fifth, the value of oversight is determined

by independence × detection probability. That the supervisor's errors are independent of the

AI's errors (error decorrelation) is a necessary but not a sufficient condition; it becomes value

only when multiplied by the probability of actually detecting errors. In addition, WP8 points

out that there is an upper bound on the number of targets one supervisor can effectively

oversee (span), and that supervisor vigilance is harder to sustain in environments where

errors occur only rarely (the rarity effect); the Falling Asleep at the Wheel finding that

improvement in AI performance itself makes oversight harder is consistent with the latter.

WP8 further organizes, from case analyses, the point that homogeneous supervisor groups

trained in the same era as the AI share blind spots and have difficulty satisfying the

independence condition. This point becomes the starting point of this paper's Proposition 5,

which positions generational heterogeneity as an organizational source of decorrelation, and

of the design arguments of Section 6.

Here the relationship between automation bias and experience requires a theoretical

treatment of the reverse mechanism. The oversight-failure evidence tends to invite a reading

on which experience is an immunity to this failure mode. According to systematic reviews,

however, automation bias is widely observed in professional decision making, including

clinical decision support (Goddard, Roudsari & Wyatt 2012), and arises most readily on tasks

with high verification complexity and heavy cognitive load (Lyell & Coiera 2017). The

verification of contextual and practical risks — precisely the setting into which experiential

audit capacity should be deployed — falls under this condition. On the other hand, that expert

judgment can be systematically distorted by prior expectations and contextual information

has been shown by pre-AI expertise research. Dror, Charlton & Péron (2006) reported that

when five fingerprint experts with an average of 17 years of experience were re-presented

with fingerprint pairs they themselves had previously identified as matches, together with

context suggesting a case of misidentification, four of the five reversed their own past

judgments. With the limitation of being a preliminary report on a small number of cases, it is

suggestive on the point that long professional experience confers no immunity to

confirmatory contextual distortion.

Superimposing these two lines of findings yields the following theoretical possibility. When

AI output is plausibly constructed and also fits the auditor's own past successes and industry

received wisdom, confirmation bias and automation bias can compound. The experienced

auditor's mental model processes convention-conforming output fluently, and that processing

ease is readily confused with a sense of "already verified." Because verification effort has

already been curtailed by automation bias, no additional cross-checking is triggered. The

consequence is that the more experienced the auditor, the more he or she may uncritically

Working Paper | Ageless Management in the AI Era 43

accept plausible AI output that fits his or her own mental model. Ironically, the

inexperienced, who lack the conventional schemata against which to check, are relatively less

liable to this particular trap. That is, the claim of the audit value of experience operates most

strongly when the error contradicts the auditor's received wisdom, and can break down

when the error conforms to it. However, no study directly measuring the compounding of

confirmation bias and automation bias in the auditing of AI output could be confirmed within

the scope of this paper's search, and this treatment is theoretical. Proposition 4 of Section 5

writes this possibility into the body of the proposition as boundary condition (i), and

Hypothesis H3, by embedding in the task both errors that conform to the auditor's industry

received wisdom (confirmation-conforming) and errors that contradict it (deviant), tests this

boundary simultaneously with the main effect of experience (Section 9, Appendix A).

Alongside the compounding with confirmation bias, a route by which the surface

properties of the object of verification paralyze the audit is also to be anticipated from classic

findings of cognitive psychology. According to research on processing fluency, the more

subjectively easy a stimulus is to process, the more people judge its content to be true. That

manipulation of perceptual fluency alone raises truth judgments of sentences has been

shown experimentally (Reber & Schwarz 1999), and the finding that fluency broadly elevates

judgments of truth, liking, and confidence has been organized by a systematic review (Alter &

Oppenheimer 2009). Generative AI output — grammatically well-formed, stylistically

polished, and low in processing resistance — sits at this high-fluency pole. The detection

trigger of experiential audit capacity is the cognitive dissonance generated by mismatch

between accumulated schemata and the output, but high-fluency output raises the very

threshold of that dissonance, so the detection mechanism can be bypassed without any deficit

in the auditor's capacity. Fluency bias, that is, is a second paralysis route for auditing,

operating independently of the confirmation-bias route that originates in the auditor's

internal priors. Proposition 4 of Section 5 writes this into the body of the proposition as

boundary condition (iii), and Hypothesis H3, by controlling the fluency of the task documents

(high fluency versus low fluency) as a factor or covariate, tests this boundary within the same

experiment as the main test (Section 9, Appendix A).

Let this paper's position be stated explicitly here. This paper does not claim that "humans

are good at oversight." As this section has shown, the evidence points rather to the failureproneness

of human oversight. This paper's claim is confined to a single point. As WP8

showed, so long as a non-delegable residual of oversight and verification exists — a residual

that cannot be delegated to AI — the supply side of oversight, a scarce and consumed

resource, must be interrogated. As a candidate for that scarce input, this paper theorizes

experiential audit capacity (Section 5, Definition 4). Whether experiential audit capacity

actually raises detection probability, however, is untested and is the object of Hypothesis H3

(Section 9). Super-seniors' participation in auditing is likewise subject to WP8's capacity

constraints, including the solvency condition and alarm fatigue. The design of the individuallevel

version of these constraints is Brain Safety in Section 8.

Working Paper | Ageless Management in the AI Era 44

4.3 The Divergence Between Assisted and Unassisted Performance

The compression evidence (Section 4.1) measured performance under AI assistance. What

remains when the assistance is removed? The most direct answer to this question comes from

Bastani et al. (2025), a randomized controlled trial within mathematics classes covering

roughly 1,000 students (grades 9–11) at a large high school in Turkey. Students were assigned

to three groups: a control group with textbook and notes only, a GPT Base group using a raw

ChatGPT-type interface, and a GPT Tutor group using teacher-designed guardrailed prompts.

During assisted practice sessions, performance improved substantially: +48% over control for

GPT Base and +127% for GPT Tutor. On the unassisted exam with access removed, however,

the GPT Base group performed significantly worse than control by 17%, and the GPT Tutor

group was statistically indistinguishable from control. The guardrails eliminated the harm

but generated no gain. In the mechanism analyses, the messages of GPT Base students were

dominated by the "what is the answer" type, which the authors call the "crutch" effect.

Answer-copying use bypasses conceptual learning and collapses in unaided settings. A

correction to the paper has been published, but its content is solely a fix to the authors'

affiliation information, with no change to the result figures above (correction content verified

on August 21, 2026). It should be kept in mind that extrapolation from high-school

mathematics to workplace skill formation remains an analogy.

For occupational skill, the same type of concern is shown by Budzyล„ et al. (2025), a

retrospective observational study at four endoscopy centers in Poland comparing the

adenoma detection rate (ADR) of standard non-AI colonoscopies before and after the

introduction of AI-assisted colonoscopy (1,443 procedures in total: 795 before introduction,

648 after). ADR fell from 28.4% (226/795) before introduction to 22.4% (145/648) after, an

absolute difference of −6.0 percentage points (95%CI −10.5 to −1.6, p=0.0089), and AI exposure

was an independent factor in the ADR decline (odds ratio 0.69, 95%CI 0.53–0.89). Being a

retrospective observational study, it cannot establish causation and goes no further than

"suggestion." Yet it is the first clinical suggestion that everyday dependence on AI can

depreciate, on a timescale of months, the unaided skill that can be exercised without AI. The

detection skill of expert endoscopists had been considered a paradigm of the kind of

perceptual and contextual skill that AI cannot compress. Even that "uncompressible ability"

can depreciate if it is not exercised.

At first glance, this finding contradicts the outage analysis of Brynjolfsson, Li & Raymond

(2025) (Section 4.1), in which AI use takes hold as learning. The two can, however, be read

integratively as a difference in task structure and AI design. The endoscopy AI substitutes for

the detection of lesions itself; the human detection skill lies dormant unexercised and

depreciates. The customer-support AI presents suggested responses, and adoption and

execution remain with the human. Because exercise continues, learning through imitation

can occur. That is, "does AI aid learning or dissolve skill" is not an either-or rule of thumb but

depends on the design variable of which of the human's cognitive processes the AI takes over

Working Paper | Ageless Management in the AI Era 45

and which it leaves on the human side — such is this paper's reading. This treatment is itself

hypothetical; no direct comparative experiment exists.

WP5 of this series (Kadowaki 2026e), within its account of the formation of brain capital

(Brain Capital), proposed the institutional securing of "protected unassisted practice," that is,

practice opportunities from which AI assistance is deliberately removed. The guardrail-group

results of Bastani et al. (2025) and the depreciation suggestion of Budzyล„ et al. (2025) are

empirical backing for the concern behind that proposal, and at the same time demand an

extension of its scope of application. Depreciation risk is not confined to young people in the

course of learning. What must be stressed in the context of Ageless Management is that the

same logic operates on super-seniors' crystallized intelligence (Gc) and on the experiential

audit capacity that stands upon it. If audit and verification roles are designed to depend on AI

pre-screening, the auditors' own detection skills depreciate, and the source of oversight value

can be undermined. The role design of the multigenerational ecosystem must therefore build

unassisted exercise opportunities into seniors' Gc-type roles as well (Section 8). Whether

protected unassisted practice is effective for maintaining experiential audit capacity,

however, is untested, and this paper hands it over to the measurement framework of Section

9 as a design hypothesis to be tested.

It should be added that this logic of depreciation applies to this paper's resource-conversion

thesis itself. The "held but idle Gc-type abilities and experiential audit capacity" assumed by

Propositions 2 and 11 (Section 5) are, for as long as a period without exercise continues after

exit, exposed to depreciation and obsolescence by the logic of this subsection. That is, the Gc

and detection abilities of super-seniors long after exit are not a reservoir sleeping intact.

Domain-contextual knowledge becomes obsolete as institutions, technologies, and markets

change, and detection skill lacking exercise has no reason to escape the depreciation pathway

that Budzyล„ et al. (2025) suggested even for professionals in active practice. The resourceconversion

claim is therefore conditioned on time elapsed since exit. Those who have

recently exited and those long after exit cannot be treated alike, and for the latter, estimating

depreciation with years since exit as a variable — including the possibility of recovery

through re-exercise — becomes an independent empirical task (see the commentary on

Proposition 2 and Section 10.4).

The divergence between assisted and unassisted performance also has implications for the

institutions of measurement and evaluation. What an organization can observe is, in many

cases, only the deliverables produced under AI assistance. Looking only at the practicesession

performance in Bastani et al. (2025) (+48%, +127%), one cannot see the deterioration

on the unassisted exam (−17%). Likewise, looking only at the paperwork of AI-assisted audit

work, the depreciation of the auditor's own detection skill remains invisible until an event

occurs in which the AI drops out. Measuring assisted performance and unassisted ability

separately is therefore not a theoretical taste but a requirement of risk management. That the

experimental design for Hypothesis H3 (Section 9, Appendix A) incorporates both the AIassisted

and unassisted conditions as factors, and measures the detection of surface errors

Working Paper | Ageless Management in the AI Era 46

and of contextual risks separately, is intended to capture this divergence at the level of

measurement design.

4.4 Conceptual Separation: Domain Experience and Chronological Age

Before connecting this section's evidence to the debate on older workers, the conceptual

distinction on which the success or failure of this paper's entire argument turns must be

made explicit. The "experts" in the prior empirical studies are operationalized by the level of

domain experience and skill, not by chronological age. The axis in Brynjolfsson, Li &

Raymond (2025) is tenure and skill measures. Noy & Zhang (2023) stratify by first-task

performance, and Dell'Acqua et al. (2023) by baseline performance. The published abstract of

Budzyล„ et al. (2025) includes no analysis by endoscopists' age or years of experience. That is,

there is almost nothing this section's evidence can say directly about "older workers."

The inference "AI helps the less experienced; therefore the experience premium of older

workers is lost (or preserved)" is accordingly doubly short-circuited, whichever direction it

takes. First, the correlation between experience and age varies greatly across occupations and

job-change histories. Second, older novices (career changers and returners) and young

experts (early specializers) exist in large numbers in reality. For this reason, this paper treats

the experience premium, the judgment-and-oversight premium, and chronological age as

three separate variables. As one of the few data points concerning age, Peng et al. (2023)

reports larger benefits for "older developers" alongside developers with fewer years of

experience, but this "older" is a matter within the 25–44 range, not data on those aged 50 and

over. It is a fragment showing that a simple "the younger, the greater the gain" gradient on the

age axis is not self-evident, but it cannot be extrapolated to the debate on older workers.

This distinction can be made intuitive by a thought experiment. A person who moves to a

different industry at 60 is high in chronological age but shallow in experience of that domain,

and belongs to the side this section's evidence identifies as "benefiting most." Conversely, an

expert in his or her 30s who has trained in a single domain since the teens is low in

chronological age but is a holder of long-term domain experience. If what determines the

competence of oversight and verification is experience, this expert in the 30s can possess high

experiential audit capacity, and the 60-year-old career changer does not. This paper's

framework accepts this consequence. Indeed, it is precisely because it accepts this

consequence that Proposition 4 can specify as its own refutation condition that "chronological

age independently predicts detection ability even after controlling for years of domain

experience." If age itself retains independent explanatory power, this paper's reattribution

thesis is rejected.

On that basis, this paper's claim is made explicit in two stages. The first stage is the

empirically established range. Heterogeneity of outcomes under AI use has been observed

along the axis of domain experience and skill. That discerning the frontier and verifying the

output are emerging as new human-side skills is also consistent with the evidence of Sections

Working Paper | Ageless Management in the AI Era 47

4.1–4.2. However, how far expertise improves that competence in verification and oversight is

itself still an open question. A report of a randomized labeling experiment said to show that

even experts miss AI's errors has appeared in PNAS Nexus (verified by title and journal only:

Why do experts miss AI's errors? Evidence from a randomized labeling experiment. PNAS

Nexus, 5(6), pgag146. The author names and content are unverified, and the bibliographic

details are under verification. This paper therefore confines itself to a footnote-level mention

of its existence and does not include it in the reference list). In addition, as organized in

Section 4.2, there is the reverse theoretical possibility that when errors conform to the

auditor's received wisdom, experience may instead impede detection through the

compounding of confirmation bias and automation bias. The audit advantage of expertise, if

it exists, is conditioned on the content dimension of the error. The second stage is this paper's

theoretical claim, and it is untested. Section 5 first defines the ability to detect the contextual

errors, practical risks, and ethical risks contained in AI output operationally — by detection

performance on verification tasks, not by attributes or rank (Definition 4, experiential audit

capacity). The definition itself does not include what forms this capacity. The claim that

individual differences in this detection ability are formed by the interaction of long-term

domain experience with Gc-type abilities and metacognition is presented, as an empirical

claim independent of the definition, in the form of Proposition 3 (the non-compressibility of

experience) and Proposition 4 (the formation of experiential audit capacity), with refutation

conditions, and is directly tested by the 2×2 factorial design of Hypothesis H3 (Section 9,

Appendix A).

Between these two stages lies an unfilled empirical gap. No study that operationalized age

or tenure and measured performance in AI oversight or verification could be found within

the scope of this paper's search. Not only the relationship between experience and AIoversight

ability, but direct evidence on the age axis is missing, and it is this absence that

grounds this paper's research-gap claim.

This subsection's conceptual distinction faces one more anticipated objection. If the

substance of experiential audit capacity lies in long-term contextual memory — the retention

of past failure cases, organizational history, and the history of customer relationships — then,

the objection runs, once expanding context windows and retrieval augmentation make it

possible to inject past data in bulk, Gc's contextual memory would likewise be substituted by

AI, and experiential audit capacity would move over to the compressed side. The response

has two stages. First, this objection mistakes the core of experiential audit capacity for the

retained volume of recorded past data. A substantial part of the judgment material at work in

auditing is never recorded in the first place — the subtleties of interpersonal relationships,

the emotional history of an organization, and value judgments never made explicit either lose

most of their information at the moment of documentation or are never documented at all.

Moreover, even where records exist, what the auditor is discerning is the semantic gap

between the described data and living reality — the mismatch whereby a description that is

consistent on paper diverges from actual conditions on the ground. However much the

Working Paper | Ageless Management in the AI Era 48

volume of data injected into the context is expanded, what the machine is given is only the

description side, and this gap is not closed. As a limitation at a different level from this, it has

also been reported that the use of injected long contexts is itself imperfect — language model

performance degrades significantly when the relevant information sits in the middle of a long

input context (lost in the middle: Liu et al. 2024). This performance limit, however, may be

alleviated by generational model turnover, and this paper does not place it among the main

pillars of its response.

Second, the more fundamental response lies at the level of the independence of

verification. Even if the substitution of contextual memory were to succeed completely — that

is, even if the same model read the whole of its own vast context and verified its own output

— that act supplies no independence of the interpreter. As WP8 (Kadowaki 2026h) formalized,

the value of oversight is independence × detection probability, and a verifier that shares its

error distribution with the generator exhibits systematic false negatives toward the

generator's own systematic blind spots. The same model reading its own context is, in

precisely this sense, a configuration in which generation and verification share an error

distribution, and no increase in the volume of contextual memory increases the

independence of verification. The volume of contextual memory and the independence of

verification are separate problems. What the expansion of context windows can compress is

therefore the function of retention and recall of memory in the context of production, not the

function of an independent interpreter in the context of verification — the boundary line

drawn at the end of Section 4.1, between differences in production and differences in

verification, is drawn in the same place against this objection as well.

Search record: as of August 21, 2026. Via WebSearch (in English), three queries — (1) study older

workers experience advantage supervising AI oversight expertise age empirical evidence experiment;

(2) "domain expertise" "AI oversight" OR "human oversight" experiment experts better detecting AI

errors age older professionals; (3) seniority tenure moderates ability to verify AI output errors

experiment "years of experience" LLM verification randomized — together with derivative searches

during the verification of individual references, were run, but no empirical study operationalizing age

or tenure and measuring AI-oversight performance could be confirmed. Adjacent studies (the expertiseaxis

miss experiment, the heterogeneity within the 25–44 band, etc.) are noted in the main text.

Finally, the bridging assumption under which this paper speaks of age is made explicit. Age

is merely a function of the time that makes possible the accumulation of long-term domain

experience, and the oversight value is attributed not to age but to experience. Even if holders

of experiential audit capacity are expected to be relatively numerous among super-seniors

(aged 60 to their 90s), that is a distributional tendency, and it does not mean that an

individual's ability may be inferred from his or her chronological age. It is precisely for this

reason that this paper's framework is "ageless" — independent of age — and not a celebration

of seniors (Section 1.4, Section 10).

Working Paper | Ageless Management in the AI Era 49

4.5 Headwind Data on Older Workers and AI, and Their Conversion

With the concepts separated, consider the data that do currently exist on the age axis. They

are mainly disparities in use, adoption, and opportunity, and their direction is a headwind.

The most credible is the representative repeated survey of Bick, Blandin & Deming (2024)

(NBER working paper, evidence grade B+). As of late 2024, roughly 40% of Americans aged 18–

64 used generative AI, an adoption speed exceeding that of the PC and the internet.

Workplace usage rates, however, show a clear age gradient: 34.5% at ages 18–29, 34.6% at 30–

39, and 29.5% at 40–49, against 16.7% at 50–64 — those in their 50s and above stand at

roughly half the level of those under 40. Adoption rates are moving rapidly: in the survey's

updated figures (August 2025 survey), overall adoption reached 54.6% (+10 percentage points

year over year) and workplace use 37.4%. Explicit statement of the survey date is

indispensable for adoption figures.

The grey-literature surveys are also directionally consistent. The five-country survey by the

employment-support NGO Generation (France, Ireland, Spain, the UK, and the US; published

October 2024) covers 2,610 entry- and mid-level workers aged 45 and over and 1,488

employers. In hiring for AI-related roles, 90% of US recruiters said they would "consider

candidates under 35," while only 32% would "consider those over 60" (86% versus 33% in

Europe). Also, only 15% of those aged 45 and over used generative AI at work. This 90%

versus 32%, however, is a statement of recruiters' intent, not an audit experiment measuring

hiring outcomes. In the AARP survey (conducted March 2026; 1,015 US workers aged 50 and

over), familiarity with workplace AI stood at 52%, use in daily work at 23%, and experience of

AI training at only 12%, while 49% wanted to learn — an observed 37-point "gap between

willingness and training opportunity." As a policy-oriented commentary, Pizzinelli & Tavares

(2026) point out an "asymmetry of opportunity and risk": older workers are relatively more

likely to hold jobs with high AI exposure and high complementarity and can benefit from AI

in jobs requiring experience, judgment, and interpersonal skills, but because their labormarket

mobility is low and the costs of job change and retraining are high, the blow is greater

should they end up on the substituted side.

Evidence-grade note: the Generation survey (C) and the AARP survey (C) in this subsection are grey

literature not subject to peer review (an NGO-commissioned survey and a membership-organization

survey), and the figures rest on intent and self-report. Pizzinelli & Tavares (2026) is a policy commentary

(C+), and citation is confined to conceptual points. None of these are measurements of ability or

outcomes. Bick, Blandin & Deming (2024) is an NBER working paper (B+), and publication in a refereed

journal is unconfirmed.

How to read these headwind data is the fork in the road for this paper. First, the age

gradient in usage rates is evidence neither of ability differences nor of benefit differences.

The disparities are confounded with differences in opportunity, training, and job

composition, and AARP's 37-point gap suggests a shortfall on the supply side (training

opportunity), not the demand side (willingness to learn). Second, none of the surveys

Working Paper | Ageless Management in the AI Era 50

measures the quality of use, that is, supervisory use versus substitutive use. Against this

section's body of evidence, this is a decisive omission. Third, from the standpoint of the

cognitive bottleneck (Section 5, Definition 3), this headwind can be explained as the

consequence of current job design and training allocation tacitly presupposing young,

execution-type (Gf-type-component-centered) AI use. If so, the prescription is neither to keep

older workers away from AI nor to hurry their assimilation into execution-type use, but a

conversion of role design: access to Gc-type, audit-type roles and the corresponding

reallocation of training and opportunity (Section 6). For this paper, the headwind data are not

a refutation but a given that demonstrates the need for conversion.

Moreover, this section's body of evidence shows that this headwind carries a danger of selffulfillment.

By the logic of Section 4.3, abilities that are not exercised depreciate. If older

workers continue to be kept away from opportunities for AI use, from audit-type roles, and

from training, then what was initially a mere opportunity gap can, with the passage of time,

convert into a real skill gap — for the discernment of AI's zone of competence and the

verification of its output are honed only within the experience of engaging with AI.

Conversely, a design that places holders of experience into audit-type roles and gives them

opportunities for exercise puts existing Gc-type assets to work and at the same time prevents

their depreciation. Whether to leave the headwind data as a given or to convert them through

role design is an organizational choice variable. It must be noted again, however, that the

effect of this conversion itself belongs to this paper's untested propositions (Propositions 2

and 11).

The direction of role conversion itself has been articulated in anticipatory form by

practitioners. Nielsen argues that "wise winnowing" — a division of labor in which generative

AI generates large numbers of options and elders holding accumulated evaluative capacity

select the best of them — can extend the productive careers of older knowledge workers.

This, however, is an essay by a prominent practitioner (evidence grade D+) and contains no

new empirical data whatsoever on AI and older users. This paper positions it not as evidence

but as a pioneering statement of a testable hypothesis. What this paper does from Section 5

onward is to formalize this intuition into the refutable form of the conceptual apparatus of

Definitions 4–6, Propositions 3–5, and Hypotheses H2–H3.

To summarize this section. (i) The compression of experience differences is empirically

established, but its subject is differences in deliverable production within AI's zone of

competence. (ii) Oversight is a scarce, consumed resource and can fail systematically.

Experience is not an immunity to this failure mode and, for convention-conforming errors,

can even become an impediment. This paper does not presuppose human skill at oversight; it

interrogates the supply side of oversight. (iii) Assisted performance does not guarantee

unassisted ability, and abilities that are not exercised — including Gc and experts' detection

skills — can depreciate. (iv) All of the above evidence lies on the experience-and-skill axis,

and direct evidence on the age axis is missing. To fill this gap, the next section presents the

Working Paper | Ageless Management in the AI Era 51

group of definitions centered on experiential audit capacity and the group of propositions

equipped with refutation conditions.

5. Theory: Definitions and Propositions

5.1 The Architecture of the Theory

Building on the evidence organized in the preceding sections, this section constructs the

theory of Ageless Management as eight definitions and twelve propositions. This paper is a

conceptual and measurement-proposal paper, and most of its propositions are untested. The

discipline of this section therefore rests on two commitments. First, every proposition carries

a refutation condition, so that the proposition itself states what observations would reject the

theory. Second, the commentary attached to each proposition, together with Table 7 (the map

of propositions), makes explicit which evidence from which section supports each

proposition and where this paper's untested claims begin. Not allowing established findings

to be conflated with this paper's theoretical wagers is the design policy of this section.

The core of the theory consists of the three theoretical operations announced in the

introduction (Section 1.4). The first operation is the reattribution from age to experience. This

paper does not adopt the popular argument that attributes the oversight value of older

workers to "age" — the argument that seniors are suited to auditing because they are

experienced. The oversight value is attributed to experiential audit capacity (Definition 4),

operationally defined by detection performance, and Proposition 4 attributes its formation to

the interaction between long-term domain experience and Gc-type abilities and

metacognition. Age is merely a function of the time that makes accumulation possible. It is

precisely this reattribution that makes this paper's framework "ageless" rather than a

celebration of seniors. As confirmed in Section 4.4, what prior empirical studies measured as

"expertise" was domain experience and skill, not chronological age, and this reattribution is

consistent with the very structure of the evidence. Definition 4 and Propositions 3 and 4 carry

this operation.

The second operation is to formalize generational heterogeneity as a source of "error

decorrelation." WP8 (Kadowaki 2026h) formalized the value of AI oversight as independence

× detection probability and generalized the value of independent oversight channels to "error

decorrelation." This paper positions the property that the overlap of judgmental blind spots is

small among individuals who have experienced different historical environments,

technological generations, and failure cases (Definition 5) as the organizational source of this

decorrelation. The starting point is the observation, drawn from WP8's case analyses, that

homogeneous supervisor groups sharing the same contemporaneous training and

information environment as the AI are prone to shared blind spots. This extension to the

generational axis, however, is a new claim of this paper that WP8 itself did not make, and it is

untested. Definition 5 and Proposition 5 carry this operation.

Working Paper | Ageless Management in the AI Era 52

The third operation is the conversion of social problems into resources. The three social

problems of population aging, constrained youth participation, and the exclusion of

marginalized groups are converted, conditional on AI-mediated complementarity (Definition

6) and Brain Safety (Definition 8), into untapped sources of brain capital (Brain Capital) =

stock (K) × utilization rate (u). What must be emphasized is the conditional side. This paper

does not claim unconditional conversion. The condition-dependence of employment effects

shown in Section 3, the risks of oversight failure and skill depreciation shown in Section 4,

and the empirical baseline that the average effect of age diversity is near zero (detailed in the

commentary on Proposition 6) all forbid the simple claim that "participation generates value."

Definitions 6–8 and Propositions 6–11 carry this operation.

The correspondence between evidence and propositions can be sketched in advance. The

cognitive-science foundations of Section 2 — the divergence of Gf and Gc, the heterogeneity of

peak ages across abilities, and the overlap of distributions across age groups — are the

empirical underpinning of Definitions 2 and 3, and they supply the premises of Proposition 2

and the grounds of Proposition 12. The evidence on work and brain health in Section 3 shows

the deadlock of average effects and the emergence of "quality of work" as a moderating

variable, motivating Proposition 7. The empirical work on AI and experience in Section 4

gives direct evidence for the first half of Proposition 1 and the first half (the compression

side) of Proposition 3. Up to this point is the territory supported by independent empirical

research. By contrast, the existence claim (a) and removal effect (b) of Proposition 2, the

second half of Proposition 3, the reattribution of Proposition 4, the generational decorrelation

of Proposition 5, the AI-mediated moderation of Proposition 6, the bidirectional capital

formation of Proposition 8, the transferability of Proposition 10, the resource conversion of

Proposition 11, and the allocation comparison of Proposition 12 are untested theoretical

propositions newly asserted by this paper, and the means of testing them are Hypotheses H1–

H3 in Section 9 (Table 7). This paper does not prove that the theory of this section is correct.

What this section does is fix the theory in a form in which its correctness can be adjudicated.

5.2 Definitions

The following eight definitions are held fixed throughout this paper. Immediately after each

definition, a note records why the definition was chosen and what it excludes.

Definition 1 (Ageless Management)

A management regime that removes the variable of chronological age from decisions on

role allocation, evaluation, participation, and exit, and that dynamically allocates roles

— within an ecosystem not limited to the boundary of employment — on the basis of

measured cognitive characteristics, accumulated domain experience, health status, and

the person's own intent.

Working Paper | Ageless Management in the AI Era 53

The core of this definition lies in pairing the removal of age with the specification of

replacement variables. As confirmed in Section 2, the peak ages of cognitive abilities are

scattered across decades depending on the ability, and the distributions of age groups overlap

widely. Under this evidentiary situation, chronological age is no more than a coarse proxy for

cognitive characteristics, and if the proxy is to be removed, the variables on which decisions

rest must be explicitly substituted. Of the four variables specified by Definition 1, cognitive

characteristics and domain experience are on the ability side, while health status and the

person's own intent are on the constraint-and-preference side; the explicit inclusion of the

latter two is meant to avoid presupposing that one can work or wants to work (a structural

response to the survivorship-bias critique in Section 10).

It is worth making explicit what this definition excludes. Extending the mandatory

retirement age (teinen) or raising the age ceiling of continued employment is an operation

that moves the threshold while retaining the variable of age, and is not Ageless Management.

Preferential treatment of a particular age bracket, such as "promoting the active engagement

of seniors," also falls outside the definition, because it uses age as an allocation variable. On

the other hand, because the definition demands measurement-based allocation, it

internalizes the danger of measurement misuse — the danger that the measurement of

cognitive characteristics turns into a new apparatus of selection and discrimination. This

danger is a cost the definition accepts, and it is addressed explicitly in the refutation

condition of Proposition 12 and the self-critique of Section 10.

Definition 2 (Gf-type tasks / Gc-type tasks)

Tasks whose performance depends primarily on fluid intelligence (processing speed,

working memory, and the learning of novel procedures) are called Gf-type tasks; tasks

whose performance depends primarily on crystallized intelligence (accumulated

knowledge, contextual interpretation, and interpersonal judgment) and on

metacognition are called Gc-type tasks. This classification is a relative weighting on a

continuum, not a binary.

This definition transfers the factor distinction of Cattell-Horn theory (Section 2.1) from

persons to the attributes of tasks. There are two reasons for the transfer. The first is the

mapping onto the functional characteristics of AI. What generative AI primarily substitutes

for is Gf-type cognitive load — search, summarization, documentation, and the execution of

routine procedures (Section 4) — and describing tasks in terms of Gf/Gc allows the domain AI

can substitute for and the functions remaining on the human side to be discussed in the same

vocabulary. The second is the avoidance of typing individuals. Definition 2 classifies tasks, not

people. The rereading "older people = Gc-type personnel" is a variant of the age stereotyping

this paper criticizes, and it is explicitly excluded, together with the definition's continuum

clause — not a binary. Real jobs are bundles of Gf-type and Gc-type components, and the

structure of this bundle is precisely the subject of the next definition, Definition 3.

Working Paper | Ageless Management in the AI Era 54

Definition 3 (Cognitive bottleneck)

Given that a job is a bundle of Gf-type and Gc-type components, the state in which

difficulty in performing the Gf-type components forces exit from the job as a whole, so

that the Gc-type abilities the person holds never reach the work.

Definition 3 is a concept for describing the job exit of older workers not as a wholesale loss of

ability but as a consequence of an institutional structure — how jobs are bundled. Given the

Gf/Gc divergence confirmed in Section 2 — the coexistence of declining components and

maintained or growing components — difficulty in performing the Gf-type components of a

bundle can force exit no matter how well the remainder of the bundle could be performed.

What is lost then is not only the individual's income. The organization, too, is relinquishing

an asset — the Gc-type abilities held — while leaving it idle. Definition 3 is a description of a

"state"; the claim that this state exists as an independent exit pathway is not the definition but

Proposition 2(a). Nor does this definition restrict the cause of the bottleneck to aging.

Circumstances that constrain the performance of Gf-type components — illness, disability,

career interruption for childcare or caregiving — exist regardless of age, and Definition 3

treats them as the same structure. It is this commonality of structure that allows Ageless

Management to be discussed in continuity with the participation of socially marginalized

groups (Proposition 11).

Definition 4 (Experiential Audit Capacity)

The capacity to detect contextual errors and practical and ethical risks contained in AI

output. This paper defines it operationally, not by attributes or position, but by detection

performance on verification tasks. (Note: what forms this capacity is not part of the

definition. The formation mechanism is the empirical claim of Proposition 4.)

Definition 4 is the central concept of this paper and embodies two constructional choices. The

first is purification into an operational definition. The definition specifies the detection

capacity solely by a measurable quantity — detection performance on verification tasks —

and does not build into the definition what forms the capacity: domain experience, Gc-type

abilities, metacognition, or chronological age. This division of labor is a structural choice to

protect the falsifiability of the theory. If the formation mechanism (the interaction of

experience × Gc × metacognition) or age-independence were written into the definition itself,

Proposition 4 would become analytically true the moment the definition was adopted, and its

empirical content would be emptied. Moreover, if a result emerged in which chronological

age predicted detection capacity even after controlling for years of experience, an escape into

the definition — "what was measured was not the experiential audit capacity of Definition 4"

— would become available as a way of evading refutation. By purifying the definition into an

operational one and making the formation mechanism the empirical claim of Proposition 4,

Working Paper | Ageless Management in the AI Era 55

this escape route is sealed, and the variance decomposition of Hypothesis H3 — the

interaction term of experience × Gc × metacognition, and the partial effect of age after

controlling for experience — becomes, as it stands, the test of Proposition 4.

The second is the restriction of the detection targets. The definition restricts the targets of

detection to contextual errors and practical and ethical risks, and does not include surfacelevel

errors (errors of form, procedure, and document quality). The empirical work in Section

4 shows that differences in production are compressed by AI, and if the definition included

detection power in the domain being compressed, the concept would fail to capture the value

specific to experience. This restriction leads directly to the design of the dependent variables

of Hypothesis H3 (the separation of detection rates for surface-level errors and for contextual

risks). Note that nothing can be derived from the definition about whether this operationally

defined detection capacity is systematically higher among holders of long-term domain

experience, or whether it is independent of chronological age. Those are the empirical claims

of Propositions 3 and 4, and they are untested. Moreover, following the logic of Section 4.3,

experiential audit capacity is itself an asset that can depreciate if not exercised, and its

protection is a design object of Brain Safety in Section 8.

Definition 5 (Generational Decorrelation)

The property that, among individuals who have experienced different historical

environments, technological generations, and failure cases, the correlation of the error

distributions of judgment (blind spots) is lower than within a single generation.

Definition 5 defines generational heterogeneity not as values or attitudes but as a statistical

property of error distributions. This choice excludes two things. The first is generational

theory — categorical claims about generational traits of the form "Generation Z values X."

Effects specific to generational categories find little support in peer-reviewed research, and

this paper does not ask generational labels for explanatory power. What Definition 5 refers to

is the history of experience — exposure to different technological environments and different

failure cases — and the trace it leaves on the correlation structure of errors. The second is the

presumption of normative implications. Decorrelation is in itself neither good nor bad; it

acquires value in the context of oversight only by way of WP8's formalization — oversight

value = independence × detection probability. By placing the definition on a measurable

quantity, error correlation, Proposition 5 can have a direct refutation procedure: the

comparison of cross-generational and within-generational correlations of detection errors.

Working Paper | Ageless Management in the AI Era 56

Definition 6 (AI-Mediated Complementarity)

The state in which, through AI substituting for or complementing the Gf-type

components of a task bundle, heterogeneous cognitive assets that conventionally had to

be combined within a single individual (Gf-type execution capacity, Gc-type audit

capacity, and problem perception as a person directly affected) become connectable and

exchangeable across individuals. This connection is not a one-way prosthesis: it is

bidirectional in that the benefit of complementation (substitution for Gf-type

components) and the supply of auditing (provision of Gc-type verification) flow mutually

among participants (bidirectional complementarity).

Definition 6 carries this paper's claim to novelty (Section 1.4). The point is that AI is

positioned not as an agent but as a medium. As confirmed in Section 2.3, the idea of using AI

as a prosthesis for Gf-type components itself belongs to the lineage of assistive-technology

and prosthetics research and is not this paper's invention. A prosthesis is a one-way relation

that fills an individual's deficit. What Definition 6 captures is the structure beyond it: the state

in which the prosthesis loosens the conventional requirement that all components of a job

bundle be performed by the same individual, so that heterogeneous cognitive assets become

connectable across individuals. Of the three assets listed, "problem perception as a person

directly affected" refers to the problem-finding capacity that those with the experience of

illness, disability, or exclusion possess precisely because of that experience; it is the term that

allows the participation of socially marginalized groups to be treated as the connection of

assets rather than as accommodation (Proposition 11). The stipulation of bidirectionality at

the end anchors the introduction's (Section 1.4) core term "bidirectional complementarity" to

this definition, and is distinguished from the "bidirectional capital formation" of Proposition 8

— a claim at a different level, that brain capital increases across all participating strata. Note

also that Definition 6 is a definition of a state; it is distinguished from the claim that the state

produces outcomes (Proposition 6), and the latter is untested.

Definition 7 (Multigenerational ecosystem)

An organizational form that, not limited to the firm's own boundary of employment,

connects the cognitive assets of participants ranging from youth to super-seniors and

including socially marginalized groups, through multiple contractual forms such as

employment, outsourced engagement, advisory roles, PBL-based educational

partnerships, and NPO partnerships.

That Definition 7 lifts the restriction to the employment boundary is demanded by both

practice and theory. In practice, the participation of super-seniors (aged 60 to their 90s) and of

youth still in school more often takes the form of outsourced engagement, advisory roles, or

educational partnership than of full-time employment. In theory, the Experience Corps

Working Paper | Ageless Management in the AI Era 57

evidence confirmed in Section 3 comes from designed, non-employment roles, and if what

operates is the quality of the role, then the contractual form is not essential. This choice,

however, carries a heavy price. Outside the employment boundary lies a vacuum of labor-law

protection (Section 7), and an ecosystem built on Definition 7 can, if designed badly,

degenerate into unpaid labor and exploitation. This danger is built into the theory as

Proposition 9, and it is answered by pairing Definition 8's Brain Safety with it. Definitions 7

and 8 must not be used apart.

Definition 8 (Brain Safety)

A health-and-safety standard that protects the brain capital of ecosystem participants

from depreciation, comprising both directions: (i) protection from cognitive load and

brain fatigue, and (ii) protection from unpaid-labor conversion and exploitation arising

from asymmetries of bargaining power.

The distinctive feature of Definition 8 is that it binds protection from cognitive load (i) and

protection from exploitation (ii) into a single standard. (i) is the individual-level version of the

solvency condition WP8 imposed on supervisors — the cognitive resources that can be

devoted to oversight have an upper bound — and follows from the fact that oversight- and

audit-type roles are inherently roles of high cognitive load. The audit participation of superseniors

is subject to WP8's capacity constraints (including alarm fatigue and span

constraints), and Brain Safety redesigns these constraints as an individual-level health-andsafety

standard (Section 8). (ii) responds to Definition 7's lifting of the employment boundary.

Turning people into unpaid advisors with "purpose" or "social contribution" as a substitute

for compensation, and converting youth PBL into labor beyond its educational purpose, are

both expropriations of brain capital arising from asymmetries of bargaining power, and they

must be prohibited within the same system of standards as (i). A welfare-benefit reading of

Brain Safety, which regards protection as a discretionary benefit, is what this definition

excludes.

5.3 Propositions

The following twelve propositions constitute the entirety of this paper's theory. Each

proposition carries a refutation condition, and the commentary immediately following shows

its derivation and evidentiary status.

Working Paper | Ageless Management in the AI Era 58

Proposition 1 (Marginal value shift)

The diffusion of generative AI lowers the marginal cost of Gf-type tasks and raises the

relative marginal value of Gc-type tasks (contextual interpretation and interpersonal

judgment in the sense of Definition 2, and above all evaluative and audit-type tasks).

This rise is conditional on Gf-type production and Gc-type verification being

complementary in production, and on verification actually detecting errors

(Propositions 3 and 4).

Refutation condition: Systematic evidence that, in labor markets after AI adoption, demand for

or compensation premiums on evaluative, audit, and contextual-judgment tasks do not rise, or

that the relative compensation of Gf-type tasks rises persistently. The observation window is, as

a guideline, occupational wage and task-demand data covering roughly ten years from the fullscale

diffusion of generative AI.

The derivation of Proposition 1 is an application, to cognitive tasks, of the standard economic

logic that a fall in the price of a substitute raises the relative value of complements. The first

half — the fall in the marginal cost of Gf-type tasks — is directly supported by the empirical

work in Section 4.1. In customer support, writing, consulting, and software development

alike, the production of deliverables within AI's competence range was substantially

accelerated and leveled, with the largest gains accruing to the less experienced (Brynjolfsson,

Li & Raymond 2025; Noy & Zhang 2023; Dell'Acqua et al. 2023). This is evidence that the cost of

procuring the performance of Gf-type components from the market is falling rapidly. The

parenthetical in the proposition text refers to Definition 2's exemplification of the Gc type

(contextual interpretation and interpersonal judgment), with the evaluative and audit-type

tasks central to this paper's concern attached as a restriction. It does not introduce an

extension not appearing in Definition 2.

The second half — the rise in the relative marginal value of Gc-type tasks — is a theoretical

consequence of the first half, but the derivation involves an assumption that must be made

explicit: the assumption that Gf-type production and Gc-type verification are complementary

in the production function. The complementarity is assumed, not derived, and since

generative AI is itself absorbing summarization, evaluation, and verification functions

(Section 10.5), the possibility that the two turn into substitutes is not empty. That is why the

second sentence of the proposition text states this assumption as a condition of the

proposition. There is, however, a structural limit to the progress of this substitution. As WP8

(Kadowaki 2026h) laid out, AI self-verification shares its error distribution with the generator

and therefore does not constitute an independent verification channel — configurations

using an LLM as verifier show systematic false negatives for the types of error the generator

itself is prone to, and as long as generation and verification derive from the same training

distribution, their blind spots are correlated. Accordingly, improvements in AI's verification

capacity can erode the marginal value of Gc-type verification, but the erosion remains inside

Working Paper | Ageless Management in the AI Era 59

the correlated blind spots, and the niche of detecting systematic blind spots by verifiers who

do not share the generator's error distribution — decorrelated verification — remains. It is to

this niche that the rise in the marginal value of Gc-type verification asserted by Proposition 1

is ultimately anchored. This residual, however, is not a fixed quantity. The AI-side diversity

required by condition (iii) of Proposition 5 — cross-verification by models of different

architectures and developers — is an operation that lowers the error correlation of AI

verification on the machine side, and its progress works to narrow further the residual niche

of decorrelation supplied by humans. How much human verification value remains is an

open empirical question that depends on the speed of this machine-side decorrelation.

Furthermore, the very evidence this paper cited in Section 4.2 shows that adding verification

does not always generate value: human-AI combinations underperform the best of either

alone on average (the g=−0.23 of Vaccaro, Almaatouq & Malone 2024). If so, what Proposition

1 can assert unconditionally extends only to an increase in the demand for (necessity of)

verification, and the rise in marginal value is a claim conditional on verification actually

detecting errors — the validity of Propositions 3 and 4, that is, the effectiveness of detection.

Even in interpreting the refutation condition, the possibility cannot be excluded that

spending on verification labor increases while detecting no errors, and this conditionality

connects Proposition 1 to Propositions 3 and 4. Note that Proposition 1 is a claim about

relative value and does not imply a rise in the absolute wages of those engaged in Gc-type

tasks. Moreover, "the value of Gc-type tasks rises" and "who can perform them" are separate

questions, the latter being the subject of Propositions 3 and 4.

The "demand for and compensation premiums on evaluative and audit-type tasks" that the

refutation condition of Proposition 1 takes as its observable can be given a concrete

generative pathway from the economics of information. The collapse in production costs

brought by generative AI means the mass circulation of unverified artifacts, widening the

informational asymmetry of quality as seen by buyers — clients, readers, regulators. Akerlof

(1970) formalized, as the analysis of the market for lemons, that in markets where quality is

hard to discern, adverse selection can arise in which inferior goods drive out good ones. A

market flooded with unverified AI-generated artifacts approaches the conditions for this

lemons market. In that situation, certification of having passed verification by experiential

audit (HOTL) can function in the structure of Spence's (1973) signaling. Because the incentive

to bear the cost of verification and obtain certification is skewed toward suppliers of quality

that withstands verification, audit certification can support a separating equilibrium as a

signal of quality, and a trust premium on certified artifacts is realized as the market value of

verification labor — a pathway that gives a generating mechanism to the compensation

premium the refutation condition of Proposition 1 takes as its observable. This paper presents

this as a theoretical pathway, however, and does not assert that the premium will arise.

Whether the credibility of certification itself is maintained (the gaming of certification — a

problem of the same form as Proposition 12), and whether the premium exceeds the cost of

verification, are both open empirical questions, to be adjudicated within the observation

Working Paper | Ageless Management in the AI Era 60

window of the refutation condition. Note that the root of demand in this pathway is anchored

not in the performance limits of current-generation models but in the structure of

verification independence — demand for verification that does not share the generator's

error distribution (Section 10.5) — and does not disappear with the turnover of model

generations.

Proposition 2 (Removal of the cognitive bottleneck)

(a) The Gf-component bottleneck (Definition 3) exists as an independent job-exit

pathway alongside institutional compulsion such as mandatory retirement, health

constraints, and demand-side age discrimination. (b) AI's complementation of Gf-type

components removes this bottleneck and lets the Gc-type abilities a person holds reach

the work. The share of pathway (a) in total exits is not quantified in this paper and is

treated as an open empirical question. The validity of (b) has two boundary conditions.

First, unassisted cognitive exercise and the periodic insertion of unmediated audit must

be maintained — constant dependence on AI without them can, through the skilldepreciation

pathway of Section 4.3, erode over the medium to long term the very

foundations of the Gc-type abilities and metacognition that the complementation is

meant to serve. Second, complementation holds only within the range in which

standard cognitive screening does not fall below the threshold of mild cognitive

impairment — the decline of Gf can accelerate nonlinearly in old age, and this model

cannot complement a state in which attentional resources themselves are exhausted.

This lower bound is a boundary on allocation to experiential-audit roles, not a

constraint on participation in the ecosystem in general.

Refutation condition: For (a): systematic evidence that, in studies decomposing reasons for

exit, difficulty in performing Gf-type components is not observed as an independent reason for

exit. For (b): experimental or quasi-experimental evidence that, even after tools complementing

Gf-type components are provided, the job performance and job continuation of older workers

holding Gc-type abilities do not improve.

Proposition 2 is split into the existence claim (a) and the removal claim (b), each with its own

refutation condition. The reason for the split should be made explicit. This paper squarely

acknowledges that the broad pathways of job exit for older workers are institutional

compulsion — mandatory retirement and ceilings on continued employment (Section 7) —

health constraints (the HWLE evidence of Section 3), and age discrimination on the labordemand

side. Institutional compulsion operates regardless of an individual's Gf level, and

66.0% of Japanese aged 65 and over are retirees (Section 10.4). Accordingly, (a) does not claim

that the bottleneck is the main cause of exit or that a "substantial share" is attributable to it.

What it claims is only that Gf-bottleneck-induced exit exists as an independent pathway

alongside these known pathways, and the quantification of its share of total exits is, as the

proposition text states, an open empirical question. The test design for (a) is the

Working Paper | Ageless Management in the AI Era 61

decomposition of exit reasons. In surveys of leavers and panel data, exit reasons are

decomposed into institutional compulsion, health, discrimination, and difficulty of job

performance, and the question is whether difficulty in performing Gf-type components

(difficulty adapting to new procedures and new technologies, etc.) is observed as an

independent exit reason not reducible to the others. If it is not observed, (a) is rejected. The

reality of the Gf/Gc divergence (Section 2) shows the structural possibility of (a) arising, but

the existence of the divergence is not evidence that the divergence causes exit, and (a) is an

untested empirical claim.

Figure 3 The cognitive bottleneck (Definition 3), its removal by AI, and the formation mechanism of

experiential audit capacity (Definition 4) (Proposition 4). Note: schematic. On the left, difficulty in

performing the Gf-type components of a job bundle forces exit from the whole bundle, and the Gc-type

abilities held never reach the work (Definition 3). On the right, AI substitutes for or complements the Gftype

components, the Gc-type abilities reach the work, and experiential audit capacity (Definition 4) is

connected to audit-type roles. The formation mechanism through the interaction of long-term domain

experience with Gc-type abilities and metacognition is the claim of Proposition 4. Together with the

removal effect (Proposition 2(b)) and the verification power of experiential audit capacity (Proposition

3), the formation mechanism is an untested theoretical claim and an object of testing by Hypothesis H3

and related studies.

Claim (b) (that complementation by AI removes the bottleneck) is the transition depicted in

Figure 3, and at present it lacks direct experimental or quasi-experimental evidence.

Research of the form the refutation condition for (b) demands — providing Gfcomplementing

tools to older workers and measuring changes in job performance and

continuation — is not found within the scope of this paper's search. On the contrary, as seen

in Section 4.5, current data show a headwind — AI usage rates among older groups are

roughly half those of younger groups — which this paper reinterpreted as a consequence of

job design and the allocation of training; but the validity of that reinterpretation is itself part

of the test of (b). Furthermore, removal of the bottleneck is a necessary condition for reach,

(a) Status quo: the Gf-type component gates exit from the whole job

Job (task bundle)

Gf-type component Gc-type component

Impairment → job exit

Held Gc never reaches

the work (lies idle)

(b) Ageless Management: AI complements the Gf-type component

Gf-type component Gc-type component

AI (complement)Deployed as experiential audit capacity

Formation of experiential audit capacity (Def. 4) — Prop. 4: a function of accumulated experience, not age (untested claim)

Long-term domain experience× Gc-type capability (contextual knowledge, interpersonal × Metjaucdoggmneitniot)n

Detection of contextual errors and practical/ethical risks in AI output (tested via Props. 3–5 and H3)

Working Paper | Ageless Management in the AI Era 62

not a sufficient one. Whether the Gc-type abilities that reach the work produce value depends

on Propositions 3 and 4, and the institutional pathways of participation depend on Section 7.

In addition, the "Gc-type abilities a person holds" in (b) requires a temporal qualification.

Following the logic of Section 4.3, unexercised abilities depreciate and knowledge of the

domain context becomes obsolete. The reservoir image — that unutilized Gc-type abilities are

preserved intact in those long past exit — is not permitted by this paper's depreciation logic.

The removal effect of (b), and the resource-conversion claims that presuppose it (Propositions

8 and 11), are therefore conditional on the time elapsed since exit. The eligibility criterion of

Hypothesis H3 (no more than five years away from practice) is the experimental-design

reflection of this condition, and it means at the same time that extrapolation to super-seniors

long after exit requires a separate estimation of a depreciation function with years since

leaving work as a continuous variable (Section 9).

The two boundary conditions the proposition text attaches to (b) are each demanded by the

logic of other parts of this paper. The first boundary condition (maintenance of unassisted

cognitive exercise and the periodic insertion of unmediated audit) is a requirement of

consistency with Section 4.3. The evidence of Section 4.3 showed that constant dependence on

AI can erode performance under unassisted conditions. If, while accepting this skilldepreciation

logic, (b) recommended constant AI complementation unconditionally, this

paper would fall into self-contradiction — dependence on the complementation would, over

the medium to long term, undermine the very foundations of the Gc-type abilities and

metacognition that the complementation is supposed to make reachable. The boundary

condition converts this contradiction into a design requirement: the removal effect of (b)

persists only when unassisted cognitive exercise and unmediated audit of raw output are

periodically inserted into the design of the role, cutting off the depreciation pathway. This

requirement is given operational form in the Brain Safety design of Section 8 (the dual-track

auditing of Section 8.1). The second boundary condition (the lower bound of cognitive

screening) is a requirement to distinguish the object of complementation from its foundation.

What AI complements is the Gf-type components of the task bundle, not the attentional

resources that the exercise of Gc-type abilities itself demands. As confirmed in Section 2.4, the

decline of Gf can accelerate nonlinearly in old age, and in a state where standard cognitive

screening falls below the threshold of mild cognitive impairment and attentional resources

themselves are exhausted, the foundation for the Gc-type exercise to be complemented is lost.

Making this neurological lower bound explicit does not justify the exclusion of any particular

group; it is the theory's honest demarcation, acknowledging that the complementation model

has a physical limit — as the end of the proposition text states, the lower bound is a boundary

on allocation to experiential-audit roles and does not constrain participation in the ecosystem

in general. The implications and misuse risks of this boundary are treated self-critically in

Section 10.4.

Working Paper | Ageless Management in the AI Era 63

Proposition 3 (Non-compressibility of experience)

Contemporaneous assistance by AI compresses differences in production (deliverable

quality, working speed, procedural knowledge) but does not compress differences in

verification (differences in the experiential audit capacity of Definition 4).

Refutation condition: Experimental evidence that, under AI-assisted conditions, the difference

in contextual-error detection rates between holders of long-term domain experience and nonholders

disappears or reverses. Or longitudinal evidence showing that the acquisition of

verification capacity accelerates in AI-use environments to the point of substantially

substituting for the effect of accumulated experience.

The boundary of Proposition 3 is placed on the axis "differences in production vs. differences

in verification." The choice of this axis rests on scrutiny of what the compression evidence

actually compressed. The substance of the compression shown by Brynjolfsson, Li &

Raymond (2025) was the transfer of tacit behavioral patterns extracted from top performers'

interactions — clarifying questions, listening, adjustment of tone — which, in the vocabulary

of Definition 2, includes interpersonal judgment, that is, Gc-type components. The withinfrontier

compression of Dell'Acqua et al. (2023) is likewise a context-dependent output: the

quality of consulting deliverables. In other words, what AI compresses is not limited to Gftype

components; insofar as they are used in production, differences in Gc-type components

are compressed as well. The boundary of compression versus non-compression must

therefore be drawn not between kinds of ability — Gf versus Gc — but between functions:

production (making the deliverable) and verification (detecting its errors). That is why

Proposition 3 is formulated on this axis rather than as "surface versus context," and the

evidentiary status is asymmetric, as follows. The first half — compression of differences in

production — is supported by multiple independent empirical studies in Section 4.1 (quasiexperiments,

preregistered experiments, and field experiments) and is the most heavily

evidenced part of this paper's propositions. The second half — differences in verification are

not compressed — is a theoretical claim of this paper that at present has no direct evidence of

the same standard. The indirect support is limited to the following: the compression studies

all measured the production of deliverables within AI's competence range and did not

measure verification capacity; and the deterioration of AI users outside the frontier

(Dell'Acqua et al. 2023) suggests the persistence of a human-side function of discerning the

boundary of that competence range.

The qualifier "contemporaneous assistance" in the proposition text, and the second

sentence of the refutation condition (longitudinal evidence), are also demanded by the

implications of the compression evidence. Brynjolfsson's compression is a shortening of the

experience curve — a compression of acquisition time — and the possibility that the same

mechanism operates on the acquisition of verification capacity — a pathway in which AI

turns failure cases and regulatory context into teaching material and accelerates the

Working Paper | Ageless Management in the AI Era 64

formation of junior auditors' capacity — cannot be excluded a priori. If this longitudinal

compression pathway is real, then even if experience differences persist in the

contemporaneous 2×2 experiment (H3), the effect of accumulated experience would be

substituted over time, and Proposition 3 would effectively fail. By writing this pathway —

which a static experiment cannot observe in principle — into the refutation condition, and

placing the corresponding longitudinal measurement in the implementation roadmap of

Section 9, the refutation condition was made to cover the full range of the proposition's claim.

Furthermore, the existence of reports unfavorable to the second half must be mentioned. A

randomized-experiment report said to show that even experts miss AI errors has appeared in

PNAS Nexus (Section 4.4; as its content and figures are unverified, mention is limited to its

existence), and whether expertise improves verification performance is itself still an open

question. Precisely for this reason, the second half of Proposition 3 is tested directly by the

2×2 factorial design of Hypothesis H3 (experience level × presence of AI assistance). H3's

prediction is that in the detection of surface-level errors the experience difference shrinks

with AI assistance (consistent with the first half), while in the detection of contextual and

practical risks and in the quality of proposed corrections, the main effect of experience

persists and does not shrink under AI assistance. If this prediction fails — if, as the refutation

condition states, the experience difference in contextual-error detection rates disappears or

reverses — Proposition 3 is rejected and this paper's theory loses its core (Section 5.5).

Working Paper | Ageless Management in the AI Era 65

Proposition 4 (Formation of experiential audit capacity)

Experiential audit capacity is formed not by chronological age but by the interaction

between domain-specific knowledge and operational schemata in a particular field (the

product of long-term domain experience) and generalized crystallized intelligence

(vocabulary, reading comprehension, general knowledge) together with metacognition.

The two are conceptually distinct — generalized Gc transfers across domains, whereas

domain knowledge is bound to its domain. This formation has three boundary

conditions. (i) When AI output conforms to the auditor's own past successes and

industry received wisdom, the synergy of confirmation bias and automation bias means

that experience can instead impede detection. (ii) In domains where the speed of

technological change is high and the half-life of domain knowledge is short, the audit

effectiveness of accumulated experience declines and can turn negative. (iii) When the

fluency and stylistic polish of AI output are high, the effect of processing fluency raises

the threshold of the auditor's cognitive sense of incongruity, and the detection

performance of experiential audit capacity can decline — the very trigger of detection,

the "sense of incongruity" itself, becomes less likely to fire in the face of fluent output.

Refutation condition: Systematic evidence that chronological age independently predicts

detection capacity even after controlling for years of domain experience, or that years of

experience lose predictive power after such control. For the boundary conditions: evidence that

experienced auditors' detection rates do not fall below non-experienced auditors' even on

confirmation-conforming errors empties (i); evidence that the audit effectiveness of experience

is preserved even in high-change-speed domains empties (ii); and evidence that detection

performance does not decline when output fluency is manipulated empties (iii) (each is a

refutation of a boundary and works in the direction of strengthening the main proposition).

Proposition 4 puts the first theoretical operation (the reattribution from age to experience)

into refutable form. Note that the refutation condition is bidirectional. If chronological age

has independent predictive power even after controlling for years of experience, the

reattribution is wrong and there is something in age itself — this paper's framework loses its

entitlement to call itself "ageless." Conversely, if years of experience lose predictive power

after control, the paper's reattribution of detection performance (Definition 4) to experience

loses support, and the proposition collapses in a different way. Because Definition 4 is

purified into an operational definition (Section 5.2), either result rejects Proposition 4 directly,

with no retreat into the definition. Proposition 4 survives only when a specific variance

decomposition is observed: experience predicts, and age does not.

The proposition text explicitly distinguishes generalized crystallized intelligence from

domain-specific knowledge in order to make the interaction claim withstand the criticism of

multicollinearity. Holders of long-term domain experience tend to have high levels of

generalized Gc as well, and if the two were measured as a single "experience-and-knowledge"

construct, the interaction term would become inseparable from the main effects, and

Working Paper | Ageless Management in the AI Era 66

Proposition 4 would degenerate into an untestable paraphrase. The substance of the

distinction lies in transferability. Generalized Gc (vocabulary, reading comprehension,

general knowledge) transfers across domains, whereas domain knowledge and operational

schemata are bound to their domain. From this distinction a testable divergence of

predictions is obtained. A non-experienced person high only in generalized Gc (a welleducated

outsider) should be inferior to a long-term experienced person in detecting

contextual risks in the domain, and a person with domain knowledge but low generalized Gc

and metacognition should likewise be inferior in detection — because Proposition 4's claim is

an interaction, not an addition. Corresponding to this prediction, Hypothesis H3 measures

generalized Gc and domain knowledge with separate standardized indicators and explicitly

estimates the interaction term (Appendix A).

The theoretical grounds of the boundary conditions are as follows. Boundary condition (i) is

the proposition-level reflection of the synergy of confirmation bias and automation bias

organized in Sections 4.2 and 4.4. Experience supplies the prior distribution in verification.

When AI output conforms to the auditor's own past successes and industry received wisdom,

the experience-derived prior works in the direction of endorsing the output's validity and

aligns with overtrust in AI output (automation bias) — at that point experience turns from a

resource for detection into a cause of missed detection. Boundary condition (ii) is a

consequence of the half-life of knowledge. The audit effectiveness of domain knowledge holds

only insofar as the correspondence between accumulated schemata and current practice,

technology, and regulation is preserved; in domains of rapid technological change, the

schemata systematically mispredict the current risk structure, so the audit effectiveness of

experience declines and can turn negative. Boundary condition (iii) is the reflection, in the

audit context, of the cognitive-psychology findings on processing fluency. When a stimulus is

subjectively easy to process, people judge its content to be more truthful — it has been shown

experimentally that manipulating perceptual fluency alone raises judgments of a sentence's

truth (Reber & Schwarz 1999), and the finding that fluency broadly elevates judgments of

truth, liking, and confidence has been organized in a systematic review (Alter &

Oppenheimer 2009). The output of generative AI sits precisely at this high-fluency pole —

grammatically well-formed, stylistically smooth, low in processing resistance. The trigger of

detection for experiential audit capacity is the cognitive sense of incongruity generated by a

mismatch between accumulated schemata and the output, but high-fluency output raises the

very threshold of this sense of incongruity, so the detection mechanism can be bypassed

without any deficit in the auditor's ability. That is, (iii) is a pathway in which the surface

properties of the object of verification operate, independently of (i), which originates in the

auditor's internal prior distribution. Note that the structure of the refutation condition is

asymmetric. Refutation of the main body (independent predictive power of age, or loss of

predictive power of experience) rejects the proposition, whereas refutation of the boundaries

— experienced auditors not inferior even on confirmation-conforming errors, effectiveness

preserved even in high-change-speed domains, no decline in detection when fluency is

Working Paper | Ageless Management in the AI Era 67

manipulated — empties the boundary conditions and leaves the main proposition standing in

a stronger form. Hypothesis H3 constructs the embedded errors in both confirmationconforming

and deviating types and controls the fluency of the task documents as a factor or

covariate, precisely in order to test the main body and boundary conditions (i) and (iii)

simultaneously in a single experiment (Appendix A).

The evidentiary status is untested. As confirmed in Section 4.4, no study operationalizing

age or tenure and measuring performance in AI oversight and verification was found within

the scope of this paper's search, and this absence is the basis of the paper's research-gap

claim. Because Hypothesis H3 manipulates experience level (long-term domain experience

holders versus juniors) as a factor, the complete test of Proposition 4 — a variance

decomposition varying experience and age independently — requires an extension that

includes older non-experienced participants (career changers, returnees) and young longterm-

experienced participants in the sample, and this is reflected in the participant

requirements of the experimental protocol in Appendix A. Note that Proposition 4 is a claim

about the formation factors of experiential audit capacity, not a claim of super-senior

superiority. There may be a distributional tendency for holders of long-term domain

experience to be relatively numerous among super-seniors, but that is an implication

mediated by a bridging assumption, not the content of the proposition (Section 4.4).

Proposition 5 (Generational decorrelation)

Under the following three conditions, a generationally heterogeneous supervisor group

has less overlap in what it misses in AI output than a same-generation supervisor group,

a wider collective detection set, and detections that reach decision-making. (i) Channel

condition: auditors access the raw output under verification in an unmediated way. This

need not cover the full volume; raw audit of a randomly sampled portion suffices (dualtrack

auditing — Section 8.1). (ii) Organizational condition: power gradients are

flattened so that decorrelated observations reach decision-making without suppression

or anticipatory deference (corresponding to the failure mode of "silence" that Belonging

in BCM 2026e guards against). (iii) Model condition: the AI output under audit is not

monopolized by a single foundation model, and cross-verification by models of different

architectures and developers is used in parallel — human-side generational

heterogeneity cannot override a situation in which the systematic blind spots of a single

model dominate all output. Absent these, the decorrelation supplied by generational

heterogeneity can be lost at the channel, organizational, or model level.

Refutation condition: Empirical evidence that, under a design satisfying the three conditions,

the cross-generational correlation of detection errors is equal to or higher than the withingeneration

correlation, or that the adoption rate of decorrelated observations into decisionmaking

does not differ significantly from the homogeneous-group case.

Working Paper | Ageless Management in the AI Era 68

The derivation of Proposition 5 is most accurately presented as a connection to WP8

(Kadowaki 2026h). WP8, facing squarely the systematic failures of human oversight —

automation bias, alarm fatigue, and the average inferiority of human-AI combinations

(Vaccaro, Almaatouq & Malone 2024) — formalized the value of an oversight channel as

independence × detection probability and generalized the substance of independence to

"error decorrelation." If multiple oversight channels share the same blind spots, adding

channels does not widen the detection set. And WP8 organized, from case analyses, the

observation that homogeneous supervisor groups sharing the same contemporaneous

training and information environment as the AI are precisely the ones prone to falling into

this state of shared blind spots. Proposition 5 is a supply-side response to this decorrelation

condition. If the correlation of error distributions is low among individuals who have

experienced different historical environments, technological generations, and failure cases

(Definition 5), then a generationally heterogeneous supervisor group should be able to supply

decorrelation organizationally.

 

 

The three conditions in the proposition text serve to build into the proposition itself the

three levels of pathways by which the supply of decorrelation can be lost. First, the reason for

(i), the channel condition. What generational heterogeneity supplies is the decorrelation of

error distributions inside individual auditors. But if the route by which auditors reach the

object of verification — the information channel — is shared, individual-level decorrelation is

canceled at the channel level. When all auditors view the output through a summary

produced by the same AI, the context dropped by that summary is equally unseen by all

auditors. Pre-screening and prioritization based on the AI's displayed confidence pushes the

places where the AI errs confidently — the typical hallucination — equally outside all

auditors' attention. That is, the heterogeneous blind-spot structures that different historical

environments have given individuals are homogenized at the entrance of the audit process

by passage through a single channel, and WP8's independence condition — oversight value =

independence × detection probability — is broken at the level of the channel. On the other

hand, unmediated audit of the full volume exceeds WP8's solvency condition — the upper

bound on cognitive resources that can be devoted to oversight — and collides head-on with

Brain Safety's demand for load reduction. That the proposition text specifies "this need not

cover the full volume; raw audit of a randomly sampled portion suffices" is the resolution of

this collision, and it is designed in Section 8.1 as dual-track auditing, combining raw audit —

original-output, unsummarized, unscreened — of a randomly sampled portion with AIsummary-

assisted screening of the remainder. Random sampling is essential — partial audit

filtered by AI reintroduces channel sharing through the filtering itself. The measurement of

detection-overlap rates with the audit channel (AI-mediated versus unmediated) as a factor is

detailed in Section 9.3.

(ii) The organizational condition is demanded by the distinction between statistical

decorrelation and organizational adoption. Even if generational heterogeneity actually

lowers the correlation of error distributions — statistical decorrelation obtains — oversight

Working Paper | Ageless Management in the AI Era 69

value is zero unless the detections reach decision-making. What blocks their reach is the

organization's power gradient. Youth, non-employment participants, and super-seniors alike

tend to sit downstream of the organization's power gradient, and observations that conflict

with the judgment of the majority or of superiors — decorrelated observations are precisely

such observations — can vanish short of decision-making, through suppression (explicit

dismissal) or anticipatory deference (voluntary silence). This is the reappearance, in the audit

context, of the failure mode of "silence" that BCM (Kadowaki 2026e) formalized as the

Belonging of the 3Bs. That the refutation condition of Proposition 5 lists, as an independent

refutation quantity, not only the correlation of detection errors but "the adoption rate of

decorrelated observations into decision-making" follows from this distinction — if statistical

decorrelation obtains but the adoption rate does not differ from the homogeneous group, the

oversight value claimed by Proposition 5 has not been realized.

(iii) The model condition is demanded by the limits of the level at which human-side

decorrelation can operate. What generational heterogeneity manipulates is the error

distribution on the human side. But when the AI output under audit all derives from a single

foundation model, the systematic blind spots of that model — errors rooted in the training

distribution and architecture, appearing in correlated fashion across all output — dominate

every audit task as a common factor on the output side. However heterogeneous the blindspot

structures the human auditors bring, if the object of verification is monopolized by a

single error-generating source, the model-level correlation cannot be overridden by them —

an error for which no one is given any detection cue is not detected even if the generation is

changed. Accordingly, the parallel use of cross-verification by models of different

architectures and developers is a precondition for human-side decorrelation to function. As

stated in the commentary on Proposition 1, this AI-side diversity is at the same time an

operation that narrows, from the machine side, the residual niche of human verification

value — the two stand in a relation of both substitution and precondition, and where the

equilibrium lies is an empirical question. Yet wherever this equilibrium eventually settles, the

structure itself — that the detection of systematic blind spots requires decorrelated verifiers

— does not disappear with capability gains, and Proposition 5 persists as a framework for

allocating the sources of that decorrelation — humans with heterogeneous experience and

models of different lineages (Section 10.5).

This extension to the generational axis, however, is not WP8's own claim but a new claim

made by this paper, and it is untested. To blur this point would be nothing other than the

operation of making untested claims look established through a chain of self-citations within

the series, and this paper explicitly forbids it (Section 10). What can be inherited from WP8 is

the framework — that decorrelation is part of the necessary conditions of oversight value —

and no more; whether generational heterogeneity actually lowers error correlation is an

empirical question to be tested by the quantity Definition 5 specifies in measurable form (the

comparison of cross-generational and within-generational correlations of detection errors).

As related indirect evidence, a mock-jury experiment manipulating racial diversity reported

Working Paper | Ageless Management in the AI Era 70

that diverse groups exchanged a wider range of information and that majority members' own

factual errors decreased (Sommers 2006), but this is laboratory evidence on racial diversity,

and replication with age and generation has not been confirmed. No peer-reviewed empirical

study directly showing that age diversity improves error detection or red-team performance

was found within the scope of this paper's survey.

Moreover, WP8's formalization teaches that even if Proposition 5 were true, it alone would

not guarantee oversight value. Independence is a necessary condition, not a sufficient one;

value arises only when detection probability is multiplied in. That is, Proposition 5 (the

supply of decorrelation) and Propositions 3 and 4 (the reality and attribution of detection

capacity) stand in a multiplicative relation, and only when both hold is the design of a

multigenerational supervisor group (Section 6.3) justified. This is why the detection-error

overlap analysis is built into the test design of Hypothesis H3, and why extension to oversight

experiments manipulating generational composition is needed in future research.

Proposition 6 (AI-mediated diversity effect)

Given the known conditions of task complexity and an inclusive climate, AI-mediated

complementarity (Definition 6) is an additional moderator of the relation between age

and experience diversity and organizational outcomes, and in its presence the relation

moves in a more positive direction. Outcomes here include not only the quality of ideas

but the avoidance of excess risk and the reduction of rework, and are evaluated as net

benefit after deducting the cost of decision-making time required for verification.

Refutation condition: Experimental or quasi-experimental evidence that, in comparisons

manipulating the presence of AI mediation while controlling task and climate conditions, no

difference arises in the relation between diversity and outcomes. Furthermore, if an observed

interaction derives solely from the deterioration of homogeneous teams through AI overtrust

and is not accompanied by an improvement in the absolute outcomes of mixed teams, this

proposition is not regarded as supported (decomposition of simple main effects is required —

H2).

The starting point of Proposition 6 is the frontal acceptance of the fact that the empirical

baseline for age diversity is severe. The average relation between age diversity and team

outcomes shown by meta-analyses lies near zero, from r=−.06 (Joshi & Roh 2009) to r=.014

(Wallrich et al. 2024), and Schneid et al. (2016) conclude no significant relation (the sole

exception being turnover). The unconditional claim that "multigenerational means more

value" cannot be supported empirically. At the same time, it has also been consistently

observed that the effect is condition-dependent. Age diversity has positive effects on

productivity only in firms engaged in creative tasks (Backes-Gellner & Veen 2013); the relation

between diversity and outcomes becomes more positive for tasks of high complexity that

depend on creative divergence (Wallrich et al. 2024); and a climate of low age discrimination,

positive valuation of diversity, and supportive leadership are conditions of success (Wegge et

Working Paper | Ageless Management in the AI Era 71

al. 2012). At the macro level, too, the relation between age diversity and productivity is humpshaped,

suggesting the existence of an optimum (Zélity 2023).

Proposition 6 is a theoretical claim that places an additional moderator — AI-mediated

complementarity — on top of this evidence structure in which "conditions decide everything,"

and its formulation takes the form of comparative statics. That is, Proposition 6 does not

make the absolute-level claim that the effect of diversity "becomes positive" under AI

mediation. What it claims is only a change of relation: in a comparison holding fixed the

known necessary conditions of task complexity and an inclusive climate, the presence of AImediated

complementarity moves the relation between diversity and outcomes in a more

positive direction. In this formulation, the proposition contradicts neither the empirical work

that detected positive effects in particular subpopulations with pre-AI data (Backes-Gellner &

Veen 2013) nor the known set of moderators (Wegge et al. 2012 and others). The theoretical

mechanism is as follows. The main part of the cost of diversity arises from communication

costs among heterogeneous members and from each member's need to fill their own weak

components on their own. AI-mediated complementarity (Definition 6) lowers both —

complementation of Gf-type components loosens the interdependence of weaknesses and

lowers the cost of connecting heterogeneous cognitive assets — so the region where the

benefits of diversity (complementarity of perspectives and blind spots) exceed its costs should

expand. However, the existing moderator studies (task complexity, climate) did not measure

AI mediation. Proposition 6 has a form that can explain the mixture of existing evidence after

the fact, but post hoc explanation is not verification. Until a comparison manipulating the

presence of AI mediation while controlling task and climate conditions is carried out — the

mixed-team experiment of Hypothesis H2 is the first step — Proposition 6 remains an

untested theoretical proposition.

The proposition text defines outcomes as net benefit in order to put the benefits and costs

of decorrelated auditing on the same ledger. The verification supplied by multigenerational

composition is not free — processing and coordinating heterogeneous observations delays

decision-making, and verification labor consumes cognitive resources (a consequence of

WP8's solvency condition). If outcomes were measured by idea quality alone, these costs

would go unrecorded while only benefits were observed, and Proposition 6 would be biased

toward irrefutability. Conversely, the principal benefits of auditing — the avoidance of excess

risk and the reduction of rework — are hard to see in immediate evaluation of outputs, and

looking only at costs would make diversity appear permanently inferior. The proposition text

therefore defines outcomes as net benefit, including the avoidance of excess risk and the

reduction of rework and deducting the cost of decision-making time required for verification,

and correspondingly Hypothesis H2 requires that time-to-decision and verification effort be

recorded as cost variables (Section 9). This net-benefit framework connects to the transactioncost

discussion of Section 6.5. In the analysis of H2, moreover, detecting an interaction is not

enough. Since human-AI combinations can on average fall below the best single agent

(Vaccaro, Almaatouq & Malone 2024), an interaction can also arise because the AI-mediated

Working Paper | Ageless Management in the AI Era 72

condition lowers the performance of homogeneous teams. To distinguish this from support

for Proposition 6, a decomposition of simple main effects is needed — the direction of the

diversity effect in each condition cell and the identification of the source of the interaction's

sign — and this requirement is reflected in the analysis plan for H2 in Section 9. The second

sentence of the revised refutation condition writes this requirement — excluding from

support for the proposition a spurious interaction deriving solely from the deterioration of

homogeneous teams through AI overtrust — into the proposition itself.

Proposition 7 (Cognitive engagement pathway)

If a brain-health effect of work exists, it is mediated not by the length of working hours

but by the maintenance of cognitive engagement through occupation in Gc-type roles.

Refutation condition: Evidence that, in mediation analysis, working hours themselves still

predict the maintenance of cognitive function after controlling for cognitive engagement, or

that the mediation pathway is rejected.

Note that Proposition 7 is a conditional. This paper does not presuppose that "work protects

brain health." As confirmed in Section 3, the empirical literature on retirement and cognitive

function is deadlocked between IV estimates showing negative effects (Rohwedder & Willis

2010) and estimates of comparable standing showing no effect or improvement, and the

verdict of systematic reviews is likewise mixed (Meng et al. 2017). Proposition 7 is an

explanatory hypothesis for this deadlock. If the category "work" mixes cognitively rich roles

with depleting ones, it is no surprise that the sign of the average effect fails to settle, and the

variable carrying the effect should be not the presence or duration of work but active

occupation in Gc-type components — evaluation, contextual interpretation, interpersonal

judgment — that is, cognitive engagement.

The evidentiary status is suggestive. That the variables separating the sign of the effect are

occupation type, the voluntariness of retirement, and the cognitive content of the work; that

accelerated decline on the Gc side was observed only for retirement from jobs of high

interpersonal complexity (Meng et al. 2017); and that RCTs of a non-employment program

with designed role quality showed effects on cognitive function and brain structure in limited

populations (Experience Corps: Fried et al. 2004; Carlson et al. 2008) are all consistent with

the mediation structure. But consistency is not verification. No study has identified the

mediation pathway itself, and Proposition 7 is the direct test target of Hypothesis H1 — an

identification strategy using exogenous variation together with mediation analysis of

cognitive engagement. If Proposition 7 is rejected, the brain-health benefit claim (part of

Proposition 8) loses its basis, but the economic benefit claims (Propositions 1–5) stand

independently (Section 5.5).

Working Paper | Ageless Management in the AI Era 73

Proposition 8 (Bidirectional capital formation)

A multigenerational ecosystem that satisfies ex-ante verifiable conditions — the

connection modes of Definition 7 and the satisfaction of Definition 8, plus the existence

of a developmental pathway in which youth progressively experience Gc-type roles in

small, low-risk projects with full delegation and acceptance of responsibility for failure

(micro-ownership) (cognitive apprenticeship — simulated audit without responsibility

does not qualify), operated as a checklist that a third party can adjudicate prior to the

observation of outcomes — increases brain capital = K (stock) × u (utilization rate)

bidirectionally. It appears as restrained depreciation (maintenance) of K among superseniors,

early formation of K among youth and marginalized groups, and a rise in u for

the organization.

Refutation condition: Evidence that, under a design adjudicated as satisfying the conditions

prior to the observation of outcomes, the brain-capital indicators of any participating stratum

systematically deteriorate ex post. Retracting the adjudication of condition satisfaction

retroactively from ex-post outcomes is not admitted as a defense of this proposition.

Proposition 8 extends the brain capital = K × u framework inherited from WP5's Brain Capital

Management (Kadowaki 2026e) from the single organization to the multigenerational

ecosystem. "Bidirectional" means the claim that the benefit appears not as a grant to a

particular stratum but as capital formation across all participating strata. Among superseniors,

the continued exercise of Gc-type roles restrains the depreciation of K (the pathway

of Proposition 7, and the reverse side of Section 4.3 — abilities not exercised depreciate).

Among youth and marginalized groups, connection with holders of experience promotes the

early formation of K. For the organization, the connection of previously idle Gc-type assets

raises u. Through this three-part structure, Proposition 8 stands not as an argument for

supporting seniors nor for developing the young, but as a theory of capital formation.

The dimensionality and structure of K should be made explicit here. K is not a single

number. This paper conceptualizes K as a construct measured in three dimensions: (1) clinical

cognitive function scores (standardized tests such as MoCA), (2) structural and functional

brain indicators, and (3) standardized indicators of domain knowledge. This inherits the

measurement framework of BCM (Kadowaki 2026e): brain capital is not a metaphor but a

construct whose measurement procedures can be specified dimension by dimension. The

integration of the three dimensions, however, is not an additive composite score. The

structure of K is a hierarchical function in which the foundational cognitive function Kbase

measured by (1) and (2) exceeding a threshold θ serves as a gate (precondition), on top of

which the domain knowledge and operational schemata Kdomain of (3) enter multiplicatively.

K = 1[Kbase ≥ θ] × f(Kbase, Kdomain)

Working Paper | Ageless Management in the AI Era 74

The reason for adopting the hierarchical structure lies in the absurdity of an additive sum. If

the dimensions were composited additively, the substitution of offsetting a decline in clinical

cognitive function scores with abundant domain knowledge to hold K constant would be

formally permitted. But in a state where the foundation of attentional resources is lost, no

amount of domain knowledge reaches audit exercise — the audit value of domain knowledge

manifests only on top of foundational cognitive function and is not a substitute for it. The gate

term 1[Kbase ≥ θ] expresses this non-substitutability at the level of functional form. And this

threshold θ is the same gate as the cognitive-screening lower bound specified by the second

boundary condition of Proposition 2(b) — that standard cognitive screening not fall below the

threshold of mild cognitive impairment. That is, the lower bound Proposition 2 placed as a

boundary on allocation to experiential-audit roles and the gate placed by the measurement

structure of K in Proposition 8 are two manifestations of the same theoretical fact — the

distinction between the object of complementation and its foundation — and this coincidence

secures the internal consistency of the theory. This structure refines the content of

Proposition 8's claim. The restrained depreciation of K among super-seniors should be

observed mainly as a restrained rate of decline of Kbase, and the early formation of K among

youth and marginalized groups mainly as the accumulation of Kdomain and the

developmental formation of Kbase; the "systematic deterioration of brain-capital indicators"

in the refutation condition is adjudicated by dimension-specific measurement in line with

this hierarchical structure. The details of measurement are placed in the organization-level

indicators of Section 9.3.

The proposition text adds cognitive apprenticeship to the conditions as a response to a

diachronic risk. If the AI-mediated division of labor were optimized only statically, it could

converge on a configuration that fixedly assigns Gf-type prototyping to youth and Gc-type

auditing to holders of experience. But this division of labor deprives youth of the opportunity

to accumulate experience — the pathway of Gc formation that consists of judging, failing, and

bearing the consequences. If Proposition 4 is correct, experiential audit capacity is the

product of long-term domain experience, so a division of labor lacking a youth Gc-formation

pathway destroys the very future supply of experiential audit capacity — a diachronic tradeoff

between present efficiency and the future supply of experiential audit capacity. Cognitive

apprenticeship — a developmental pathway in which youth progressively experience Gc-type

roles in small, low-risk projects (Section 6.2) — is the condition that resolves this trade-off at

the level of an ex-ante verifiable checklist. That the proposition text explicitly requires microownership

of this pathway — full delegation and acceptance of responsibility for failure —

and excludes simulated audit without responsibility, has a reason grounded in the formation

mechanism of metacognition. The core of the experiential audit capacity claimed by

Proposition 4 is the metacognition of knowing where one's own judgment can err, and it is

calibrated only under the condition that one bears the consequences of one's judgments

oneself. In simulated audit whose consequences are attributed to others, the cost of error

does not return to the auditor, so the mapping between judgment and consequence — the

Working Paper | Ageless Management in the AI Era 75

feedback loop necessary for calibrating metacognition — does not close. However much

simulated audit without responsibility is repeated, what is formed is knowledge of the formal

procedures of auditing, not genuine metacognition. The requirement of micro-ownership is

therefore not an additional desideratum of the developmental pathway but a constitutive

condition for cognitive apprenticeship to function as a Gc-formation pathway. An ecosystem

lacking it degenerates into a one-way design that maintains the K of super-seniors while

sacrificing the future K formation of youth. That the refutation condition rejects the whole

proposition upon the deterioration of "any participating stratum" is meant to refuse to

condone this one-way degeneration under the name of bidirectionality.

The evidentiary status differs by stratum. That intergenerational knowledge transfer

carries motivational benefits for both sender and receiver (Burmeister, Wang & Hirschi 2020),

that reverse mentoring can yield skill development on both sides (Kaše, Saksida & Miheliฤ

2019), and the qualitative description that intergenerational learning is bidirectional (Gerpott,

Lehmann-Willenbrock & Voelpel 2017) are peer-reviewed findings consistent with the

structure of bidirectionality. What these measured, however, was motivation, retention, and

skill development, not brain-capital indicators themselves. Verification on brain-capital

indicators is entrusted to Hypotheses H1 and H2 and the organization-level indicators of

Section 9, and in that sense Proposition 8 is an untested proposition with suggestive evidence.

The refutation condition cites the deterioration of "any participating stratum" because

bidirectionality is precisely the content of the proposition. A design that maintains the K of

super-seniors at the cost of youths' learning, or the reverse, rejects Proposition 8 even if other

strata benefit.

The proposition text requires that the conditions be operated as an "ex-ante verifiable

checklist," and the refutation condition explicitly forbids retroactive denial of the conditions,

in order to seal off a refutation-evasion structure peculiar to conditional propositions. Since

Definition 8 is itself the standard protecting brain capital from depreciation, if the

adjudication of condition satisfaction were made dependent on ex-post outcomes, every case

in which brain-capital indicators deteriorated could be reclassified retroactively as "Brain

Safety was not satisfied," and no failure would ever wound the proposition — an

immunization of the no-true-Scotsman type, under which one can keep saying "it was not

true Ageless Management." The only way to sever this structure is to make the adjudication of

condition satisfaction independent of outcomes. That is, the connection modes of Definition 7

and each requirement of Definition 8 (Table 9 in Section 8 gives their skeleton) are operated

in checklist form such that a third party can adjudicate satisfaction before seeing results, and

the adjudication procedure is written into the proposition itself: if deterioration is observed

ex post under a design adjudicated ex ante as "conditions satisfied," the proposition is

rejected. The second sentence of the refutation condition is the consequence of this

procedure, making explicit that retroactive reclassification from ex-post outcomes is not

admitted as a defense of the proposition.

Working Paper | Ageless Management in the AI Era 76

Proposition 9 (The protection vacuum of non-employment forms)

Current labor and social-security institutions are designed with the employment

relationship as the principal unit of protection and restraint, and the non-employment

participation on which Ageless Management depends (outsourced engagement, advisory

roles, PBL, NPO partnerships) falls into a vacuum of institutional protection. Ageless

Management that leaves this vacuum unaddressed can degenerate into exploitation.

Refutation condition: That in jurisdictions granting non-employment participants protections

equivalent to employment (accident compensation, adequacy of remuneration, correction of

bargaining power), cases of exploitation and unpaid-labor conversion attributable to the

protection vacuum are not systematically observed — in which case this proposition becomes

empty in that jurisdiction.

Proposition 9 is the proposition by which the theory internalizes the price of Definition 7's

lifting of the employment boundary. Protective devices such as working-hour regulation,

minimum wages, accident compensation, and dismissal regulation are designed with the

employment relationship as their unit, and outsourced engagement, advisory roles,

educational partnerships, and NPO partnerships lie outside many of them. Under this

structure, there is an ever-present danger that, in the name of Ageless Management, levels of

load and unpaid work that would be illegal in employment are legitimated as "flexible

participation." The evidentiary status of Proposition 9 is at the level of institutional analysis

rather than experimental verification, and the actual state of each jurisdiction's institutions

— Japan's freelance-protection legislation, labor protections for youth, and the institutional

environment of older-age employment — is described in Section 7 on the basis of verified

primary legal sources.

The refutation condition of Proposition 9 occupies a singular position within this theory,

because the state in which it is satisfied — the existence of jurisdictions granting nonemployment

participants protections equivalent to employment — is a desirable state for this

paper. Proposition 9 is not a claim of permanent truth but a pointer to a structural feature of

current institutions, and it is a proposition that positively hopes to be rendered "empty" by

institutional reform. In this respect Proposition 9 functions as the normative node that calls

for the institutional comparison of Section 7 and the Brain Safety design of Section 8. What

the pursuit of Proposition 8 in disregard of Proposition 9 — capital formation without

protection — degenerates into is stated explicitly at the end of the proposition.

Working Paper | Ageless Management in the AI Era 77

Proposition 10 (Transferability of the design principles of protection)

Labor protection for youth and the protection of super-seniors from exploitation and

cognitive overload are not identical in their grounds — for youth there exist principles

with no counterpart for older strata: consideration for developmental stage and the

priority of education. At the level of design principles, however, the two share a common

structure as responses to asymmetries of bargaining power and to exit costs (priority of

health, ceilings on load, adequacy of compensation), and this common part is

transferable across strata.

Refutation condition: Legal-theoretical or empirical evidence showing, for any of the design

principles held to be common structure, that it fails to function as protection when applied to

one stratum, or that it is structurally incompatible with the protection of the other stratum.

Proposition 10 does not claim identity in the grounds of protection. As the standard legal

doctrine of the ILO conventions (C138/C182) and the minor-protection provisions of the Labor

Standards Act shows (Section 7.3), at the core of youth protection lie principles with no

counterpart in the protection of older strata — consideration for developmental immaturity

and the connection to compulsory education — and the proposition text acknowledges this

asymmetry squarely within the proposition. What is claimed is weaker, but substantive as

design theory: in the residual after removing stratum-specific principles, the protections of

both strata share the character of responses to a common structure — asymmetry of

bargaining power and exit costs — and the design principles derived from it (priority of

health, ceilings on load, adequacy of compensation) are transferable across strata — the

formulation of common structure plus stratum-specific principles. Youth are prone to

accepting unfavorable terms through lack of experience, information, and alternative

opportunities; super-seniors through the exit costs of scarce reemployment opportunities and

role loss. Within the limits of this common structure, Brain Safety (Definition 8) can be

constructed as a hierarchical design that constitutes the common part as a single system of

standards and stacks stratum-specific principles (such as the primacy of education for youth)

on top. This is the theoretical basis of the design argument of Section 8.

The evidentiary status is untested, and the character of the test also differs from other

propositions. The validity of Proposition 10 is adjudicated by legal-theoretical examination

and comparative institutional analysis rather than by experiment. The refutation condition is

placed not on the interpretive question of the sameness or difference of grounds but on an

operable criterion: failure of transfer. If any of the design principles held to be common

structure is shown to fail to function as protection when applied to one stratum, or to be

structurally incompatible with the protection of the other, the transferability claim is rejected

and Brain Safety retreats entirely to stratum-specific design. This paper does not exclude that

possibility. Even if Proposition 10 falls, Definition 8 itself remains maintainable in stratified

form, and there is no propagation to the core of the theory (Section 5.5).

Working Paper | Ageless Management in the AI Era 78

Proposition 11 (Conversion of social problems into resources)

The three social problems of population aging, constrained youth participation, and the

exclusion of marginalized groups are converted into untapped sources of brain capital,

conditional on AI-mediated complementarity (Definition 6) and on Brain Safety

(Definition 8) operated in an ex-ante verifiable form. The conversion is not automatic.

Refutation condition: Systematic evidence that, even under a design adjudicated as satisfying

the conditions prior to the observation of outcomes, the participation of the strata in question

does not contribute to the organization's brain-capital and performance indicators. As with

Proposition 8, retroactive denial of the conditions ex post is not admitted as a defense.

Proposition 11 is the third theoretical operation (the conversion of social problems into

resources) itself, and it is the integrative proposition of the upstream propositions. The

conversion pathway runs through Proposition 2 (reach through removal of the bottleneck),

Propositions 3–5 (the audit value of the abilities that reach), Proposition 6 (diversity effects

under AI mediation), and Proposition 8 (bidirectional capital formation), and is conditioned

by Propositions 9 and 10 (the design of protection). Proposition 11 is therefore less an

independent claim than an integrative consequence that holds only under a conjunction of

conditions — that both AI-mediated complementarity and Brain Safety are satisfied — and it

weakens in tandem if any upstream proposition falls. The final sentence, "The conversion is

not automatic," is at once a summary of this conditionality and an explicit break with the

discourse that unconditionally relabels population aging an "asset" — an optimism that is

merely the flip side of the accommodation-and-compensation paradigm this paper rejected in

Section 1.

The reason the proposition text attaches the qualifier "operated in an ex-ante verifiable

form" to the Brain Safety condition, and the refutation condition forbids retroactive denial of

the conditions, is identical to the sealing of immunization described in the commentary on

Proposition 8. In a conditional proposition, if the adjudication of condition satisfaction

depends on ex-post outcomes, every failure case can be reclassified as "the conditions were

not satisfied," and the proposition becomes effectively irrefutable. As an integrative

proposition, Proposition 11 sits in the position most prone to this structure — every case in

which conversion failed to occur could be blamed on unsatisfied conditions. Proposition 11

therefore adopts the same adjudication procedure as Proposition 8. The satisfaction of

Definitions 6 and 8 is operated as a checklist that a third party can adjudicate prior to the

observation of outcomes (Table 9 in Section 8), and if no contribution is observed under a

design adjudicated ex ante as satisfying the conditions, the proposition is rejected.

Reclassification tracing back from ex-post outcomes — "it was not true condition satisfaction"

— is not admitted as a defense.

The evidentiary status is untested. For each of the three social problems' strata, no study

exists showing that participation under a design satisfying the conditions contributes to the

Working Paper | Ageless Management in the AI Era 79

organization's brain-capital and performance indicators. What this paper can offer is the

evidentiary status of each proposition composing the pathway (Table 7) and an order of

verification: after testing the pathway through H1–H3, integrative evaluation through the

implementation roadmap of Section 9 (pilot → case study → longitudinal). Note that the

benefit of the participation of marginalized groups, including problem perception as a person

directly affected (Definition 6), has the thinnest peer-reviewed backing of the three strata, and

— including connection to prior practice in inclusive design — it is a task for future research.

Proposition 12 (Superiority of dynamic role allocation)

Fixed role allocation based on chronological age is inferior to dynamic role allocation

based on measured characteristics (Definition 1). The grounds are three: (i) intraindividual

variation in cognitive characteristics, (ii) the magnitude of the overlap of

distributions across age groups, and (iii) the malleability and trainability of

characteristics. Measurement-based allocation, however, is structurally exposed to

gaming under Goodhart's law. The operational requirements of Definition 1 therefore

include the aperiodic replacement of measurement tasks; ex-post verification

(backtesting) by unmediated audit of real work logs rather than one-off test scores; and

real-time blind in-situ verification without advance notice (unannounced in-situ

verification) — the latter constituting a double barrier against not only adaptation to the

tests but the gaming of "ways of keeping logs that backtesting cannot detect" itself.

Refutation condition: Comparative evidence that age-fixed allocation achieves outcomes and

welfare equal to or better than dynamic allocation, or evidence that the costs of measuring

characteristics, mismeasurement, and gaming exceed the gains of dynamic allocation.

Proposition 12 is the proposition that justifies the operation of Definition 1, and its three

grounds are all supported by the evidence of Section 2. (i) Cognitive characteristics vary over

time even within the same individual, and each ability traces a different trajectory (the

asynchrony of peak ages: Hartshorne & Germine 2015). (ii) The distributions of age groups

overlap widely, and differences in group means are too coarse to use for predicting

individuals (the effect sizes of Salthouse 2009 are likewise not large enough to erase the

overlap of distributions; Section 2.2). (iii) Abilities are malleable through compensation and

training, supported by the lineage of SOC theory (Section 2.3) and the intervention evidence

of reducing age discrimination through manager training (Wegge et al. 2012). Age-fixed

allocation appears rational only when all three points are ignored.

That the grounds are empirically established, however, is distinct from the superiority of

dynamic allocation being empirically established. No evidence exists comparing the outcomes

and welfare of age-fixed and dynamic allocation, and the superiority claim itself is untested.

Moreover, the second limb of the refutation condition — the possibility that the costs of

measurement, mismeasurement, and gaming exceed the gains — is a serious implementation

issue. Measuring cognitive characteristics involves cost and error; if measurement is tied

Working Paper | Ageless Management in the AI Era 80

directly to treatment, incentives to manipulate scores arise, and the measurement apparatus

itself can turn into a new apparatus of selection and discrimination (Section 10.3). By building

this possibility into the refutation condition, Proposition 12 presents the superiority of

dynamic allocation as an empirical claim conditional on the quality of the measurement

regime. Identifying the design conditions of a measurement regime under which the

superiority holds is a task for Section 9 and future research.

The proposition text writes the operational requirements into the proposition itself because

gaming is not a peripheral implementation problem for dynamic allocation but a structural

vulnerability. Under a regime in which measurement is tied directly to treatment, the

indicator becomes a target of optimization, and a targeted indicator loses its correspondence

with what it measures — when a measure becomes a target, it ceases to be a good measure, in

the formulation of Goodhart's law (Strathern 1997). Whereas age-fixed allocation is nearly

immune to gaming precisely because chronological age cannot be manipulated, allocation

based on measured characteristics is structurally exposed to this law in the form of score

inflation through coached preparation and overfitting to the measurement tasks. The

proposition text specifies three barriers as operational requirements of Definition 1. First, the

aperiodic replacement of measurement tasks — the more fixed and predictable the tasks, the

more investment in overfitting pays off, so the very irregularity of replacement lowers the

expected return on overfitting. Second, ex-post verification by unmediated audit of real work

logs (backtesting) rather than one-off test scores — a test result at a single point can be

prepared for, but the cost of continuously falsifying a running record of the quality of

judgment in real work is high, and unmediated audit (Section 8.1) bypasses window-dressing

at the stage of summarizing and filtering the logs. Third, real-time blind in-situ verification

without advance notice (unannounced in-situ verification). This third barrier is necessary

because the second barrier can itself become a new target of optimization. What backtesting

verifies is the logs that were kept, and as long as the manner of keeping logs — what to

record, what not to record, at what granularity — is under the control of the person being

allocated, a deeper level of gaming becomes possible: adaptation toward "ways of keeping

logs that backtesting cannot detect." Unannounced in-situ verification bypasses the mediating

layer of logs altogether by observing real-time judgment itself, blind and without advance

notice, rather than via records. The second and third barriers thus constitute a double barrier

against two levels of gaming: adaptation to tests (addressed by the first and second) and

adaptation of log formation itself (addressed by the third). With these three requirements, the

superiority claim of Proposition 12 is explicitly conditioned on the existence of a gamingresistant

measurement regime. The residual risks of gaming and measurement misuse that

remain nonetheless are treated self-critically in Section 10.3.

The operation of unannounced in-situ verification carries a principle of separating gaming

deterrence from surveillance pressure on individuals. What the third barrier primarily

calibrates is system-level measurement accuracy — a test of the divergence between

backtesting results and real-time observation — not the condemnation of individuals. The

Working Paper | Ageless Management in the AI Era 81

results of unannounced verification are therefore used, in the first instance, in anonymized

and aggregated form for calibrating the system. Their use for individual treatment (such as

changes in role allocation) is limited to cases that pass through aggregation over long

windows and due process including disclosure to the person and the opportunity to respond;

no demotion or revocation of authority is carried out on the basis of a single unannounced

result. The deterrent effect on gaming is achieved by making known the fact that the system

is calibrated, and does not require diverting individual observations to individual

surveillance. Unannounced verification lacking this separation degenerates into precisely the

surveillance pressure that Section 8's Brain Safety must exclude, and destroys, upstream of

measurement, the flattening of power gradients that supports the reach of decorrelated

observations to decision-making (the organizational condition of Proposition 5). Operational

details are placed in Section 8.1.

The prior literature to which Proposition 12 should be connected is made explicit: the

economics of statistical discrimination (Phelps 1972; Aigner & Cain 1977). The conditions

under which allocation based on a coarse proxy such as age can be rational — when

observing an individual's true productivity is costly and measurement carries error, using

group statistics as a proxy can be second-best optimal under informational constraints — are

precisely the subject this literature formalized. Within this framework, Proposition 12 is

positioned as the claim that "falling measurement costs and rising measurement accuracy

undermine the rationality conditions of proxy use." In an environment where measuring

cognitive characteristics is costly and error-prone, reliance on the proxy variable of age can

persist by the logic of Aigner & Cain, and the second limb of Proposition 12's refutation

condition — evidence that the costs of measuring characteristics, mismeasurement, and

gaming exceed the gains of dynamic allocation — takes over, as its refutation condition,

exactly the rationality conditions of proxy variables that this literature formalized. That is,

Proposition 12 does not deny the theory of statistical discrimination; it asserts the superiority

of dynamic allocation as a function of the theory's conditions, and whether the superiority

holds is an empirical question depending on the state of measurement technology and the

measurement regime.

5.4 Table 7: The Map of Propositions

Table 7 lists the content of the twelve propositions, the evidence they rest on, their

evidentiary status, and the means of testing them. Evidentiary status is graded in three

categories: "established" means directly supported by independent peer-reviewed empirical

research; "suggestive" means indirect or consistent evidence exists but direct testing is

lacking; "untested" means a theoretical claim of this paper lacking direct evidence. Composite

propositions are graded part by part.

Table 7 The map of propositions: evidentiary status and means of testing

Proposition Summary of content Evidence relied on

Evidentiary

status Means of testing

Working Paper | Ageless Management in the AI Era 82

1 Marginal

value shift

AI lowers the marginal

cost of Gf-type tasks and

raises the relative

marginal value of Gc-type

tasks (conditional on

complementarity and the

effectiveness of detection

= Propositions 3 and 4)

Sections 4.1

(compression

evidence) and 4.2

(oversight failures)

Cost-decline

side

established /

value-rise

side

suggestive

(conditional)

Future research

with labor-market

data (observation

window: roughly

10 years after

diffusion)

2(a) Existence

of the

bottleneck

The Gf bottleneck exists

as an independent exit

pathway alongside

mandatory retirement,

health, and

discrimination

(quantification of its

share is an open

empirical question)

Section 2 (Gf/Gc

divergence =

structural

possibility)

Untested Exit-reason

decomposition

studies (future

research)

2(b) Removal of

the bottleneck

AI complementation of

Gf-type components

removes the bottleneck

and lets Gc-type abilities

reach the work

(boundary conditions:

maintenance of

unassisted exercise and

unmediated audit /

cognitive-screening lower

bound; also conditional

on time elapsed since

exit)

Sections 4.3 (skill

depreciation) and

4.5

(reinterpretation of

headwind data)

Untested Experiments and

quasi-experiments

providing Gfcomplementing

tools (future

research)

3 Noncompressibility

of experience

Contemporaneous AI

assistance compresses

differences in production

but not differences in

verification

Section 4.1

(production side);

Sections 4.2–4.4

(indirect evidence

on the verification

side)

Production

side

established /

verification

side untested

H3 (2×2 factorial

design) +

measurement of

the longitudinal

compression

pathway (Section 9)

4 Formation of

experiential

audit capacity

Experiential audit

capacity is formed not by

chronological age but by

the interaction of domain

knowledge and

operational schemata

with generalized Gc and

metacognition (boundary

conditions: confirmation

bias (i), the half-life of

knowledge (ii),

processing fluency (iii))

Sections 4.2 and 4.4

(the bias synergy;

absence of age-axis

empirical work =

research gap);

processing-fluency

findings (Reber &

Schwarz 1999;

Alter &

Oppenheimer 2009)

Untested H3 (separate

measurement of

generalized Gc and

domain

knowledge;

confirmationconforming/

deviating tasks;

fluency control) +

extension varying

experience and age

independently

(Appendix A,

future research)

Working Paper | Ageless Management in the AI Era 83

5 Generational

decorrelation

Under three conditions —

channel (raw audit of

random samples),

organization (flattened

power gradients and

reach of observations to

decision-making), and

model (AI-side crossverification)

generationally

heterogeneous

supervisor groups have

less overlap in misses,

wider detection sets, and

detections that reach

decision-making

WP8 framework

(inherited); BCM

2026e (Belonging);

Sommers (2006)

(indirect)

Untested (the

extension to

the

generational

axis is this

paper's new

claim)

H3 overlap analysis

+ oversight

experiments

measuring audit

channel and

adoption rates

(Section 9.3, future

research)

6 AI-mediated

diversity effect

Given task and climate

conditions, AI-mediated

complementarity is an

additional moderator of

the diversity-outcome

relation, moving it in a

more positive direction

(outcomes evaluated as

net benefit after

deducting decision-time

costs)

Section 5.3 (nearzero

meta-analytic

baseline and

condition

dependence)

Baseline and

condition

dependence

established /

AI-mediated

moderation

untested

H2 (mixed-team

experiment,

including

decomposition of

simple main effects

and recording of

cost variables)

7 Cognitive

engagement

pathway

If a brain-health effect of

work exists, cognitive

engagement mediates it

Section 3

(heterogeneity of

effects; Experience

Corps)

Suggestive H1 (mediation

analysis +

identification

through exogenous

variation)

8 Bidirectional

capital

formation

Ecosystems satisfying exante

verifiable conditions

(satisfaction of

Definitions 7 and 8 +

checklist operation of a

cognitive-apprenticeship

developmental pathway

with micro-ownership;

simulated audit without

responsibility does not

qualify) increase brain

capital K × u across all

participating strata (K is a

hierarchical function in

which domain

knowledge enters

multiplicatively above a

gate of foundational

cognitive function)

Section 5.3

(knowledgetransfer

and

reverse-mentoring

evidence); WP5

(inherited)

Motivation

and skill side

suggestive /

brain-capital

indicators

untested

H1 and H2 +

organization-level

indicators (Section

9.3)

Working Paper | Ageless Management in the AI Era 84

9 Protection

vacuum of nonemployment

forms

Non-employment

participation falls into a

vacuum of institutional

protection and, if left

unaddressed, can

degenerate into

exploitation

Section 7

(institutional

comparison based

on primary legal

sources)

Suggestive

(institutional

analysis)

Comparative legal

research and crossjurisdiction

comparison (future

research)

10

Transferability

of the design

principles of

protection

The grounds of

protection are stratumspecific

(the

developmental principle

for youth), but commonstructure

design

principles (priority of

health, ceilings on load,

adequacy of

compensation) are

transferable across strata

Sections 7 and 8

(analysis of

common structure

and stratumspecific

principles)

Untested Legal-theoretical

examination and

institutional

comparison (future

research)

11 Conversion

of social

problems into

resources

The three social

problems convert into

sources of brain capital

conditional on Definition

6 and Definition 8

operated in ex-ante

verifiable form

Integration of

Propositions 2–10

(depends on

upstream

propositions)

Untested H1–H3 +

implementation

roadmap (pilot →

case study →

longitudinal)

12 Superiority

of dynamic role

allocation

Age-fixed allocation is

inferior to measurementbased

dynamic allocation

(operational

requirements: as barriers

against Goodhart's law,

aperiodic replacement of

measurement tasks,

backtesting of real work

logs, and the double

barrier of unannounced

in-situ verification

without advance notice)

Section 2 (grounds

(i)–(iii)); Strathern

(1997)

Grounds

established /

allocation

comparison

untested

Comparative

experiments and

field studies of

allocation schemes

(future research)

 

 

 

Note: The grading of evidentiary status is this paper's own. "Established" means that independent peer-reviewed

empirical research exists for the part in question, not that the proposition as a whole has been established. The

content of H1–H3 is detailed in Section 9, and the experimental protocol of H3 in Appendix A.

5.5 Relations among the Propositions and the Structure of Refutation

The twelve propositions are not parallel; they form a structure with dependencies. Making

that structure explicit amounts to specifying which parts of the theory collapse when which

proposition falls — that is, to writing the refutation condition of the theory as a whole. The

structure can be organized into five layers. The core is Propositions 3 and 4 (the reality and

attribution of experiential audit capacity); their premise connection is Propositions 1 and 2

Working Paper | Ageless Management in the AI Era 85

(value shift and reach); the extension to organizational design is Propositions 5 and 6

(decorrelation and the mediated effect); the extension to welfare and capital formation is

Propositions 7 and 8; the institutional conditions are Propositions 9 and 10; and then the

integration is Proposition 11 and the operation is Proposition 12.

The core of the theory is Propositions 3 and 4. If Proposition 3 falls — if the experience

difference in contextual-error detection disappears under AI assistance — experiential audit

capacity (Definition 4) is definable as a concept but economically worthless. If AI compresses

even differences in verification capacity, the reason to seek the supply side of oversight in

holders of experience disappears, Proposition 5's decorrelation argument loses its supply-side

meaning, and Proposition 11's resource conversion loses its core pathway. Proposition 1 itself

survives, but the raised marginal value is no longer attributed to holders of experience, and

this paper's theory dissolves into a generic "theory of oversight in the age of AI." The fall of

Proposition 4 is graver still. If chronological age independently predicts detection capacity

even after controlling for years of experience, then the reattribution from age to experience

(the first theoretical operation) is wrong, and the very name of the theory, "ageless," fails to

hold. The theory then degenerates into one of the age-based arguments this paper rejected in

Section 1 — celebration of seniors or exclusion of seniors. Conversely, if years of experience

lose predictive power after control, the reattribution that assigns the variance of detection

performance to experience loses support, and even if Proposition 3 held, its bearers could not

be identified. It is because of this structure that the testing of Propositions 3 and 4 (H3) is

placed at the top priority of this paper's verification plan (Section 9).

The premise-connecting Propositions 1 and 2 are the conditions for the core to have value.

If Proposition 1 falls — if the relative marginal value of Gc-type tasks does not rise — then

even if experiential audit capacity is real, it is a capacity that does not become scarce, and

Ageless Management may remain an ethical imperative but does not stand as a claim of

managerial rationality. If Proposition 2(b) falls — if AI's Gf-type complementation does not

improve the job performance and continuation of older workers — Gc-type abilities remain

held without reaching the work, and the capital formation of Proposition 8 lacks supply. The

rejection of Proposition 2, however, does not directly wound the core (Propositions 3 and 4):

even if the pathway of reach is blocked, the audit value of those who have reached is an

independent question. In that case, the theory survives in reduced form, shrinking from a

"theory of expanded participation" to a "theory of role reallocation among those already

participating."

The organizational-design Propositions 5 and 6 are the layer that extends the theory from

the individual to the group level. Even if Proposition 5 falls — even if generational

heterogeneity does not supply error decorrelation — individual-level experiential audit

capacity (Propositions 3 and 4) is unscathed, and supervisors may be selected not by

generational composition but by measuring individual detection capacity. What is lost is the

oversight rationale for "multigenerational" as an organizing principle, and the organizational

design of Section 6 (Section 6.3) requires substantial revision. If Proposition 6 falls, the

Working Paper | Ageless Management in the AI Era 86

performance effect of diversity returns to the meta-analytic baseline (near zero), and the

outcome-side justification of the multigenerational ecosystem is confined to the oversight

value of Proposition 5 and the knowledge transfer and capital formation of Proposition 8. If

Propositions 7 and 8 fall, the brain-health and human-capital benefit claims disappear, but the

economic core stands independently. Independence in the reverse direction can be claimed

only more narrowly. Proposition 7 is a claim about the mediation structure of the brainhealth

effect of work and can be tested independently of any of Propositions 1–6. But

Proposition 8 is not so. The restrained depreciation of K among super-seniors depends on the

continued exercise of Gc-type roles (the pathway of Proposition 7), and reaching those roles

depends on the bottleneck removal of Proposition 2. In a world where Proposition 2 has

fallen, the capital formation of Proposition 8 lacks supply (previous paragraph). Therefore, if

Propositions 1–6 all fall, what remains unscathed is Proposition 7 alone, and the theory is

then lost as a "theory of managerial rationality," shrinking to one hypothesis in health and

welfare research on the quality of work and cognitive engagement. The economic strand (1–

6) and the welfare strand (7 and 8) are not two independent pillars; what can stand

independently is confined to the single support of Proposition 7.

The institutional Propositions 9 and 10 and the operational Proposition 12 govern not the

truth of the theory but its implementability. The refutation of Proposition 9 means, as noted

above, the realization of a desirable institutional state, and does not wound the theory. The

refutation of Proposition 10 pushes Brain Safety back to stratified design, but the need for

protection itself is unmoved. The refutation of Proposition 12 is grave. If the gains of dynamic

allocation are eaten up by measurement costs, mismeasurement, and gaming, Definition 1 is

unimplementable, and this paper's theory can maintain the critique that "age criteria lack

cognitive-scientific grounds" (Section 2) but fails to present an alternative. Finally, since

Proposition 11 is the integration of the upstream propositions, it is hard to refute in isolation;

instead, the rejection of any upstream proposition propagates to it in tandem. In sum, the

single vital point of this theory is Propositions 3 and 4, and Hypothesis H3 is designed to

strike this vital point directly (Section 9, Appendix A). For the proponent of a theory to

identify its most fragile point and direct the first test at it — that is the minimum

responsibility owed by a section that lines up twelve untested propositions.

6. Designing the Multigenerational Ecosystem

The definitions and propositions of the preceding section are, in themselves, abstract

theoretical apparatus. This section translates them into the language of organizational design:

how to configure the initial placement of the roles of the four players constituting the

multigenerational ecosystem of Definition 7 (Section 6.1); how to keep that placement in

motion without letting it degenerate into fixation by age (Section 6.2, operationalizing

Proposition 12); how to design the generational composition of the supervisor group (Section

6.3, operationalizing Proposition 5); how to condition the connections with NPOs and

Working Paper | Ageless Management in the AI Era 87

educational institutions outside the boundary of employment (Section 6.4); and where this

design does — and does not — connect to Future Value Theory (FVT) (Section 6.5).

The character of this section should be delimited in advance. It is a normative design

argument, and much of its basis consists of the untested propositions of Section 5. The parts

that are empirically established (the evidence on compression, the baseline of the diversity

meta-analyses, the evidence on conditional moderators) and the parts that this paper asserts

theoretically (AI-mediated complementarity, the oversight value of generational

decorrelation) are distinguished explicitly, subsection by subsection. Not allowing the design

argument to be read as an "established prescription" is the descriptive discipline of this

section.

6.1 Reconfiguring the Roles of the Four Players

Table 5 organizes, for the four players constituting the multigenerational ecosystem, the

bottlenecks that conventional job design has imposed, the content of complementation by AI,

and the default roles in the ecosystem. The meaning of the word "default" should be fixed

first. The role placement of Table 5 is a design of initial values derived from the distribution

of cognitive assets statistically expected to be relatively more prevalent in each group

(Section 2); it is not a rule that determines an individual's role from age. As confirmed in

Section 2, the distributions of cognitive characteristics across age groups overlap

substantially, and every age group contains a substantial number of individuals whose

characteristic profiles correspond to other rows of the table. The distinction between default

design and fixed roles is a principle running through this entire section, and its operation is

the dynamic role allocation of the next subsection.

Table 5 Reconfiguration of the four players' roles in the multigenerational ecosystem (default

design)

Player

Conventional

bottleneck

Complementation

by AI

Default role (initial

placement) Design cautions

Super-seniors

(aged 60 to

their 90s)

Difficulty in

performing the

Gf-type

components of

the job bundle

(processing

speed, working

memory,

learning of novel

procedures)

forces exit from

the job as a

whole (Definition

3); barriers to

digital operation

Substitution for and

complementation of

Gf-type components

through search,

aggregation,

summarization,

documentation, and

navigation of

operating procedures

Experiential

auditing: detection

of contextual errors

and practical and

ethical risks in AI

output (Definition

4), contextual

evaluation,

matching against

failure cases

Auditing is a

consumed resource

subject to capacity

constraints (WP8;

Section 8). Empirical

demonstration of

detection probability

is the object of

testing under H3.

Fixation into the

auditor role is itself

a new form of

stereotyping (Section

10). Securing

opportunities for

Working Paper | Ageless Management in the AI Era 88

unassisted exercise

(Sections 4.3 and 8)

Youth (under

20)

Lack of practical

experience and

domain

knowledge; long

apprentice

periods in

onboarding

Immediate supply of

past cases and

procedural

knowledge;

execution support for

generation and

analysis (experience

compression; Section

4.1)

Rapid prototyping:

iteration of trial and

error, bringing in a

feel for novel

technological

environments

Primacy of

educational purpose

and precedence of

schooling

(Proposition 10;

Section 7). Protection

of unassisted

learning. AI use

without guardrails

can harm

independent

learning (Bastani et

al. 2025)

Socially

marginalized

groups (via

NPO and

similar

partnerships)

Restricted

employment

opportunities,

accessibility

barriers, lack of

practical training

opportunities

Support for job

performance through

task subdivision and

multimodal work

support by voice,

image, and other

channels (evidence

on effectiveness in

Section 6.4)

Supplying problem

perception as those

directly affected:

surfacing

overlooked

problems and latent

needs

AI does not advance

inclusion

unconditionally

(Omri et al. 2025;

Section 6.4). Bias in

AI hiring tools (El

Morr et al. 2024).

Confinement to the

"insider-perspective"

role trivializes

participation

The midcareer

generation

(20s–50s)

Depletion of

cognitive

resources under

the double

burden of

execution and

management

Relief of load

through automation

of routine

management tasks

and reporting

Orchestration:

integrating

multigenerational

insight and AI

output, crystallizing

them into business

decisions

The integrator is

itself subject to

automation bias and

the constraints of

oversight span

(WP8). Risk of

depreciation of one's

own skills through

AI dependence in

execution (Section

4.3)

Note: This table reconstructs the role scheme of the original draft on the basis of the definitions of Section 5 and the

empirical qualifications of Section 4. Roles are defaults (initial placements), not fixed. The role placement in each row

is a design initial value based on distributional tendencies of age groups; an individual's placement is determined by

measured characteristics, experience, condition, and volition (Section 6.2). The "design cautions" column was added to

make explicit that the propositions each role relies on are untested and that role fixation is a risk.

What Table 5 inherits from the original draft's scheme is the four-player configuration;

what it changes is three points. First, the role of the super-seniors was changed from

expressions of status such as "meta-auditor and arbiter of value" to the functional expression

of experiential auditing. The ground of audit value is neither age nor status but the

interaction of accumulated domain experience with Gc-type abilities and metacognition

(Definition 4; Proposition 4), and its detection probability remains unestablished (H3).

Working Paper | Ageless Management in the AI Era 89

Second, a "design cautions" column was attached to each role. As seen in Section 4, the

benefits of AI complementation are not unconditional, oversight is failure-prone, and

dependence on assistance can depreciate skills. Role design deserves the name of design only

when it prices in these headwinds. Third, for the AI complementation of socially

marginalized groups, the original draft's assertion that "participation becomes possible" was

weakened into a conditional formulation. The grounds are presented in Section 6.4.

On the default role of the youth, one change in the environment must be added.

Conventionally, the organizational entry of the young presupposed a pathway that began

with routine entry-level tasks and accumulated domain knowledge through apprenticeship.

But precisely those routine entry-level tasks stand at the front line of substitution by AI, and it

has been pointed out that youth may be vulnerable to the effects of AI adoption alongside, or

even more than, older groups (Ayalon 2026 — a non-systematic review, cited as a prior

example of problem framing). This asymmetry is an exogenous given for the design argument

of this section. If the conventional entry pathway of apprenticeship thins out, the pathway

connecting youths' cognitive assets (adaptation to novel technological environments, less

constrained ideation) to the ecosystem will cease to exist unless deliberately designed. The

youth row of Table 5 is, in this sense, a proposal for the reconstruction of an entry pathway,

not an endorsement of the status quo. Its institutional mode (PBL-type; primacy of

educational purpose) is treated in Section 6.4 and Section 7.

The theoretical claim of this paper is that the interaction of the four players becomes

connectable only through AI-mediated complementarity (Definition 6). Conventionally,

contextual verification capacity, execution capacity, and problem perception were required to

be combined within the same individual. That is why the "mid-career worker with both

experience and execution capacity" was made the core of the organization, while holders of

experience only, or of perception only, were marginalized. When AI substitutes for the Gftype

components of task bundles, this combination requirement loosens, and a division of

labor that connects heterogeneous cognitive assets across individuals becomes technically

possible. As Definition 6 makes explicit, this connection is not a one-way prosthesis: it is

bidirectional in that the benefits of complementation (substitution of Gf-type components)

and the supply of auditing (provision of Gc-type verification) flow mutually among

participants (bidirectional complementarity). That said, "becoming possible" and "producing

results" are separate propositions, and the latter hangs on the test of Proposition 6 (Section

6.5).

The context-partition protocol for hybrid domains. The orchestration of the mid-career

generation explicitly includes a design task: the assignment of audit targets. Boundary

condition (ii) of Proposition 4 warned that in domains where technological change is fast and

the half-life of domain knowledge is short, the audit effectiveness of accumulated experience

declines and can turn negative. But this boundary condition cannot be operated through the

crude rule "do not use experiential auditing on fast-changing projects," because real projects

belong neither to purely fast domains nor to purely slow ones: they are hybrid domains.

Working Paper | Ageless Management in the AI Era 90

Within a single business proposal or deliverable, slowly changing components — legal and

governance matters, the configuration of relationships and interests, industry practice — and

fast-changing components — the specifications of frontier technologies, tool environments,

the capability range of models — are inseparably mixed. The unit of assignment is therefore

the component, not the project. The orchestrator decomposes the output under audit into

components with different speeds of change, assigns experiential auditing to the components

with long half-lives, assigns the components with short half-lives to auditing by younger and

mid-career members and to diversity on the AI side — cross-validation by different

foundation models, as the model condition of Proposition 5 requires — and evaluates each

component separately. This separated evaluation prevents erroneous auditing of experience

on the fast components (overconfidence in a previous generation's specifications) from

contaminating the detection value on the slow components, and vice versa. This protocol is

thus an operational response to boundary condition (ii) of Proposition 4: rather than leaving

the boundary condition as a limit of the theory, it converts it into a design variable, the

assignment of audit targets. Decomposition into components is itself, however, work that

requires judgment, and erroneous decomposition can become a new pathway of misses. The

effort decomposition requires and the risk of misdecomposition are entered as verification

costs in the net-benefit ledger of Section 6.5.

6.2 Dynamic Role Allocation — Operationalizing Proposition 12

Proposition 12 asserts that fixed role allocation based on chronological age is inferior to

dynamic role allocation based on measured characteristics (Definition 1). Its grounds were

three: (i) the intraindividual variability of cognitive characteristics, (ii) the magnitude of

distributional overlap across age groups, and (iii) the malleability and trainability of

characteristics. This section develops the proposition into operating rules for allocation.

The core of the operation is to allocate roles not by age but by four variables. First,

measured cognitive characteristics: describe the job bundle by the relative weights of its Gftype

and Gc-type components (Definition 2) and match it against the individual's

characteristic profile. Second, accumulated experience: the allocation criterion for audit-type

roles is the length and quality of domain experience, not chronological age (Section 4.4;

Proposition 4). Third, condition: cognitive characteristics show intraindividual variation

within the day and the week, and the same individual has states suited to auditing and states

unsuited to it. Allocation is not a qualification fixed by a single measurement but an

assignment updated according to state. Fourth, the person's own volition: unless roles are

something the person can re-choose, measurement-based allocation turns into an apparatus

of selection (Section 10).

If this operation functions properly, the correspondence between the rows of Table 5 and

individuals becomes fluid. A person who moved to a different industry at 60 is on the

inexperienced side in that domain and enters through the roles of prototyping and learning.

A practitioner in their 30s who has trained in a single domain since their teens can hold high

Working Paper | Ageless Management in the AI Era 91

experiential audit capacity (the thought experiment of Section 4.4). Participants with large

fluctuations in condition have their audit tasks reassigned to lighter-load times and formats.

Only when Table 5 is operated in this way is the default design distinguished from fixed roles.

A word on the unit of allocation. The unit of dynamic role allocation is not the "job" but the

task or session carved out of the job bundle. That the same person handles audit tasks in their

own domain in the morning and, on another day, takes the prototyping side in an adjacent

domain where their experience is shallow, is if anything a natural consequence, given that

Definition 2 defines Gf-type and Gc-type as relative weights on a continuum. Nor is the

direction of transition one-way. In place of the seniority-based single-track path of "execution

→ management → audit," the ecosystem permits multidirectional transitions, including

returns from auditing to execution, concurrent holding of execution and auditing, and

increases and decreases in the volume of participation. In particular, a design in which those

placed in audit roles are severed completely from execution undermines the very basis of

experiential audit capacity, by the skill-depreciation logic seen in Section 4.3. Transition

possibility is a requirement of fairness and, at the same time, a requirement of capacity

maintenance (Section 8).

Multi-role concurrency (identity buffer). Since the unit of allocation is the task or

session, one further operating principle can be derived: changes of allocation are operated

not as "demotion" from a single role but as reweighting of a portfolio premised on the

simultaneous holding of multiple roles. That is, every participant always holds concurrently

more than one of the roles of auditing, learning, and execution, and allocation changes driven

by measurement, condition, and volition are implemented as adjustments of the weights

within that role portfolio — lowering the weight of auditing while raising the weights of

learning and execution, and so on. This principle is needed as a response to an

organizational-psychology risk. Under a regime in which a participant's occupational identity

is unified into a single role — "I am the auditor" — a measurement-based change of allocation

is experienced by the person as the stripping of that single identity: being "removed from the

auditor role." A change experienced as role deprivation readily invites loss of self-efficacy, the

stigma of perceived decline from those around, and quiet disengagement in which the person

formally continues to participate while withdrawing inner engagement. This is nothing other

than damage to the pathway Proposition 7 identified — if a brain-health effect of work exists,

what mediates it is cognitive engagement. If dynamic role allocation is operated in a way that

destroys participants' engagement, it raises the precision of allocation while undermining the

very thing allocation is meant to protect — cognitive engagement and, beyond it, brain

capital. Multi-role concurrency is, in this sense, not a device of efficiency but a design for

protecting engagement, and it is included among the operating requirements of dynamic role

allocation as a buffer (identity buffer) that absorbs the psychological cost of allocation

changes.

The bidirectional feedback protocol. The operation of dynamic role allocation also

includes the design of how information flows between roles. This paper places the destination

Working Paper | Ageless Management in the AI Era 92

of experiential-audit feedback, as a rule, on the AI side rather than on persons. That is, the

contextual errors and risks detected by auditing are reflected, as the basic form, not as direct

instruction or correction of younger individuals but as indirect feedback into the revision of

the AI's prompts, evaluation criteria, and verification procedures. There are three reasons.

First, the object of auditing is AI output, not persons (Definition 4). An operation that diverts

detection results into the evaluation of individuals degrades auditing into over-surveillance

and destroys the psychological safety of those audited. Second, if direct instruction from

seniors to juniors were made the basic form, the division of labor in Table 5 would harden

into a one-way relation of authority between "the generation that audits" and "the generation

that is audited." A reverse-seniority relation in which junior execution constantly passes

through senior approval is a ready source of intergenerational friction. Third, reflection into

the system side prevents the recurrence of detected errors through design improvement

rather than individual attention, which is also consistent with the suppression of alarm

fatigue (Section 8).

This protocol is not complete in one direction. Its counterpart is reverse mentoring from

juniors to seniors — the operation of AI tools, trends in new models and technological

environments, sharing of the information environment of younger generations. That reverse

mentoring can yield skill development for both younger mentors and older learners has peerreviewed

evidence (Kaše et al. 2019; Section 6.5). The pair of indirect auditing and reverse

mentoring implements, at the level of information flows, the bidirectional complementarity

(Definition 6) in which the benefits of complementation and the supply of auditing flow

mutually, and is a preventive design against the audit regime sliding into a structure of

intergenerational surveillance and conflict.

The cognitive apprenticeship loop — micro-ownership. The cognitive apprenticeship

that Proposition 8 requires as an ex-ante verifiable condition — a development pathway in

which youth experience Gc-type roles in stages, in small, low-risk projects, with full

delegation and acceptance of responsibility for failure (micro-ownership) — is implemented

within the operation of this section as a development track. The core of the implementation is

the combination of the smallness of the risk and the realness of the responsibility. That is,

small, low-risk projects whose erroneous consequences are limited for the organization —

improvements to internal processes, pilots within a restricted customer scope, small-budget

initiatives — are carved out as units; within that scope, full authority over judgment,

including auditing as Human on the Loop (HOTL), is delegated to the junior, and the

consequences of failure — rework, explanation to stakeholders, recovery — are borne by the

person. Verification exercises with known embedded errors (Section 8.1) are useful as a

calibration device at induction, but they do not themselves constitute the development

pathway. As Proposition 8 expressly excludes, simulated auditing without responsibility does

not cultivate genuine metacognition — where the consequences of judgment do not come

back to oneself, the metacognitive calibration loop that maps judgments to consequences

does not close, and what forms is limited to knowledge of the formal procedures of auditing

Working Paper | Ageless Management in the AI Era 93

(Section 5, commentary on Proposition 8). The development track is completed by

progressively raising the scale and risk level of the delegated projects, together with feedback

on detection performance and the consequences of failure. This implementation is needed

because a division of labor that fixes Table 5's initial placement (juniors = prototyping,

seniors = auditing) for single-period efficiency carries a diachronic risk. If Proposition 4 is

correct, experiential audit capacity is the product of long-term domain experience, and a

fixed division of labor deprives youth of the pathway of judging, failing, and bearing the

consequences through which Gc forms, destroying the future supply of experiential audit

capacity itself — a diachronic trade-off between present efficiency and future supply (Section

5; Proposition 8). The cognitive apprenticeship loop is an operational response to this

diachronic risk, and it makes explicit that dynamic role allocation is not only single-period

matching of person and place but also a pathway of capacity formation.

Responsibility layering. The cognitive apprenticeship loop requires an explicit

demarcation concerning where responsibility lies. Even in a micro-ownership domain fully

delegated to a junior, senior experiential auditing — involvement as HOTL — can exist; but if

the mode of that involvement is left unbounded, the delegation is hollowed out. If senior

detection is operated as a veto over the junior's decisions, final decision authority reverts in

effect to the senior, and the core of micro-ownership — the experience of judgment that

carries responsibility — is lost. Conversely, if senior involvement is construed as shared

responsibility, a moral hazard arises on the junior's side — "the senior saw it, so it is not my

responsibility" — and the mapping of judgment to consequence — the metacognitive

calibration loop — again fails to close. Further, over-inhibition, in which a junior constantly

conscious of the senior's gaze optimizes for failure avoidance, damages the very purpose of

this track, trial and error. This paper therefore layers responsibility. Senior HOTL in the

junior's micro-ownership domain is confined not to a veto that halts decisions but to Socratic

coaching that prompts the junior's own reconsideration through questions — "What happens

if this premise fails?" "Who will need an explanation?" — while final decision authority and

the attributed responsibility for its consequences remain with the junior. Whether to adopt

the senior's observations is also the junior's decision. This demarcation prevents inhibition

because there is no veto, and prevents moral hazard because responsibility is not shared.

This demarcation of responsibility among participants is, moreover, not in conflict with the

external attribution of responsibility treated in Section 7 — the contractual cap on liability

that precludes shifting responsibility onto auditors and attributes final risk to the business

entity. The former is a demarcation of development and evaluation inside the ecosystem; the

latter is the allocation of legal risk in relations with third parties — a problem of a different

layer. What micro-ownership's "acceptance of responsibility for failure" means is the internal

acceptance of consequences — rework, explanation to stakeholders, recovery — not

unlimited external liability for damages.

The greatest design risk is that dynamic role allocation reproduces age stereotypes. When

Table 5's initial values — "seniors audit, juniors prototype" — slide in operation into the rule

Working Paper | Ageless Management in the AI Era 94

"audit because senior, prototype because young," Ageless Management becomes ageless in

name only. This paper specifies four design principles against this slide. First, do not close the

entrance by age: define the placement requirements for audit roles by measured experience

and characteristics, and use age neither as a requirement nor as reference information.

Second, do not impose accountability for deviation from the default: an operation that

demands justification only from deviators (juniors who audit, seniors who prototype) is de

facto fixation. Third, guarantee the transparency of measurement and a channel of appeal:

measurement-based allocation carries risks of measurement error, gaming, and

discriminatory diversion (the refutation condition of Proposition 12; Section 10). Fourth,

prepare the climate conditions. One of the few robust findings of age-diversity research is

that effects are conditioned on climate. Wegge et al. (2012), from data on more than 745 teams

and 8,848 persons across three industries, identified as conditions for age diversity to

function, in addition to high task complexity, low age stereotyping and age discrimination

and positive appraisal of diversity. Li et al.'s (2021) survey of 3,888 establishments likewise

reports that, absent age-inclusive management, the effects of age diversity are nonsignificant

or negative. Dynamic role allocation is an operation that functions only on top of these

climate conditions.

A direction on the practice of measurement should also be indicated. The measurement

dynamic role allocation requires is not a one-shot qualifying examination for selection but

continuous, low-burden monitoring in support of placement. The candidates are a

combination of structured description of domain experience (domains, years, types of failure

cases handled), standardized measurement of Gc-type abilities and metacognition, work

samples (performance on actual audit and prototyping tasks), and the person's stated

preferences; their indicator construction at the organizational level is treated in Section 9.

What must be stressed here is the restriction on the use of measurement. The same

measurement, used for placement support, becomes the foundation of dynamic role

allocation; used for selection and exclusion, it becomes an apparatus of selection more

precise than the age criterion that Definition 1 rejected. This danger is confronted head-on in

Section 10 as an internal critique of this paper's theory.

Dynamic role allocation, moreover, requires measurement costs. Measuring characteristics,

re-measuring, and maintaining role descriptions cost money; measurement that is too coarse

produces misallocation, and measurement that is too precise produces surveillance. This is

why the refutation condition of Proposition 12 explicitly includes the case in which "the costs

of measurement, mismeasurement, and gaming exceed the gains of dynamic allocation." The

superiority of dynamic role allocation is a theoretical claim, open to comparative testing

against age-fixed allocation (Section 9).

Working Paper | Ageless Management in the AI Era 95

6.3 Organizational Design for Generational Decorrelation — Operationalizing

Proposition 5

As seen in Section 4, WP8 (Kadowaki 2026h) formalized the value of oversight as

independence × detection probability and identified as a necessary condition that the

supervisor's errors be independent of the AI's errors (error decorrelation). WP8 further

organized, from case analyses, the observation that a homogeneous supervisor group trained

contemporaneously with the AI shares blind spots and has difficulty satisfying the

independence condition. Proposition 5 positioned generational heterogeneity as a supply-side

response to this decorrelation condition: the claim that individuals who have experienced

different historical environments, technological generations, and failure cases have a smaller

overlap of judgmental blind spots than within a single generation (Definition 5). Proposition 5

does not, however, make this claim unconditionally. Generational heterogeneity manifests as

an expansion of the group's detection set, and that detection reaches decision-making, only

under three conditions: (i) that auditors have unmediated access to the original output under

verification (the channel condition); (ii) that power gradients are flattened, so that

decorrelated observations reach decision-making without suppression or self-censorship (the

organizational condition); and (iii) that the AI output under audit not be monopolized by a

single foundation model (the model condition) (Section 5). What these conditions demand of

organizational design is developed in turn in the latter part of this section.

Translated into organizational design, this proposition becomes a problem of supervisorgroup

composition. That is, when designing an audit regime for AI output, the design variable

is not only the capacity of the individual auditors but the correlation of errors among

auditors. Isomorphic to the way a financial portfolio is designed not only on the expected

returns of the individual assets but on the correlations among assets, the supervisor group

should be designed not only for individual detection capacity but so that the overlap of misses

is small — this is the design idea of this section. Generational composition is one axis of this

portfolio. However many auditors are lined up who were trained in the same technological

environment and share the same failure cases, the shared blind spots remain. A formation

that combines different technological generations — for example, a generation that

experienced the domain's failures in pre-AI manual work and a generation trained together

with AI — in theory widens the detection set.

Working Paper | Ageless Management in the AI Era 96

Figure 4 The structure of the multigenerational ecosystem and generational decorrelation. The

cognitive assets of the four players are connected through AI-mediated complementarity (Definition 6),

and roles are allocated dynamically on the basis of measured characteristics rather than age

(Proposition 12). The generational heterogeneity of the supervisor group is designed as the

organizational source of error decorrelation (Definition 5). Note: this figure is a schematic of the

theoretical structure; the effectiveness of each connection hangs on the tests of Propositions 5 and 6 and

Hypotheses H2 and H3.

Three constraints on this design idea must nevertheless be made explicit. First, WP8's

capacity constraint binds here as well. Oversight is a consumed resource; adding supervisors

creates problems of coordination costs and the distribution of alarms; there is an upper

bound on the number of objects one supervisor can effectively oversee (span); and the

solvency condition operates on the cognitive resources, time, and money an organization can

devote to oversight. Diversification for decorrelation is not free, and the size of the oversight

portfolio is finite. The monotone prescription "the more diverse auditors, the better" does not

hold under the solvency condition and alarm fatigue. Second, decorrelation is a necessary

condition, not a sufficient one. As WP8 made explicit, the value of oversight arises only when

independence is multiplied by detection probability. Whether a generationally heterogeneous

auditor group in fact has less overlap of misses, and whether each auditor can detect errors

at all, are both untested, open to the refutation condition of Proposition 5 (that, under a

design satisfying the three conditions, the intergenerational correlation of detection errors be

equal to or greater than the within-generation correlation, or that the adoption rate of

decorrelated observations into decision-making not differ significantly from the

homogeneous-group case) and to the tests of Hypotheses H2 and H3 (Section 9). Third, to turn

generational decorrelation into a rule about the age composition of supervisors — quotas

such as "so many members aged 60 or over on the audit committee" — is an operation that

AI-mediated complementarity

(Definition 6)

Super-seniors (aged 60–90s)

Experiential audit, context evaluation

Receives Gf support, supplies Gc

Youth (under 20)

Rapid prototyping

AI compresses the experience barrier

Marginalized participants

Lived-experience problem perception

Cognitive accessibility

Mid-career generation (20s–50s)

Orchestration

Integration and decision-making

Dynamic role allocation (Prop. 12): roles assigned by measured traits, not age

Generational decorrelation (Def. 5): experience of different eras and technology generations → overseer pools with less-overlapping blind spots (supply-side answer to WP8)

Working Paper | Ageless Management in the AI Era 97

proxies the heterogeneity of experience, which is the ground of decorrelation, by the

heterogeneity of age, contrary to the conceptual separation of Section 4.4. The design variable

is strictly the heterogeneity of experience, training environments, and failure cases; age

composition is only its coarse observable proxy.

Independence of the audit channel. In addition to the three constraints above, the

conditional clause of Proposition 5 — that auditors access the original output under

verification without mediation — imposes an independent requirement on portfolio design.

What generational decorrelation supplies is heterogeneity of the error distributions among

auditors; but that heterogeneity manifests as an expansion of the detection set only when

each auditor reads the original output itself with their own experience and their own frame

of judgment. When all auditors share, as their information channel, summarization and prescreening

by the same AI, the heterogeneity is homogenized at the entrance of the audit

process. The context that the AI summary dropped is equally unseen by all auditors.

Prioritization based on confidence pushes the places where the AI errs with confidence — the

typical hallucination — equally outside the attention of all auditors. Then, however

generationally heterogeneous the auditor group, misses re-correlate at the level of the

channel, and the first factor of oversight value — independence × detection probability —

collapses. The financial analogy from the opening of this section applies directly: however

heterogeneous the individual assets, if the price information of all assets passes through a

single source, the diversification effect of the portfolio is lost. The audit regime of the

ecosystem must therefore always include, within the audit portfolio, an unmediated channel

— a slot in which at least some auditors access the original output directly, with neither

summarization nor pre-screening. It need not be the full volume: as Proposition 5 makes

explicit, raw audit of a randomly sampled portion suffices. This requirement, however, stands

in direct tension with the designs Section 8 considers for containing cognitive load — AI

summarization, dialogic auditing, prioritization. The management of this trade-off between

load reduction and independence, and its resolution as dual-track auditing, is treated in

Section 8.1.

Flattening the power gradient. The organizational condition (ii) of Proposition 5 derives

from the distinction between statistical decorrelation and organizational adoption. Even if

generational heterogeneity actually lowers the correlation of error distributions, oversight

value is not realized unless the detection reaches decision-making (Section 5). What blocks

arrival is the organization's power gradient. The auditors of the multigenerational ecosystem

— non-employment super-seniors, youth participating through educational partnerships,

participants via NPOs — are all likely to sit downstream of the power gradient. And

observations that conflict with the judgment of the majority or of superiors — decorrelated

observations are precisely such observations — can vanish before decision-making, through

explicit dismissal (suppression) or voluntary withholding (self-censorship). This is the

reproduction, in the audit context, of the "silence" that BCM (Kadowaki 2026e) formalized as

the failure mode of Belonging. The design requirement comes down to not making the arrival

Working Paper | Ageless Management in the AI Era 98

of observations depend on individual courage. The candidate devices are the

institutionalization of the recording of detections and of escalation channels, a duty to

respond to audit opinions (recording the reasons when dismissing), and an operation that

removes the speaker's age, contract form, and status as variables from the evaluation of an

observation; these are incorporated as the "voice channel" in Table 9 of Section 8. Isomorphic

to the way the formation of the audit portfolio is a device protecting the independence of the

channel, these are devices protecting the independence of adoption.

The model condition. Condition (iii) of Proposition 5 is a design variable on the side of the

audited object, not of the auditors. When the AI outputs under audit all derive from a single

foundation model, that model's systematic blind spots — errors rooted in the training

distribution and the architecture and appearing correlated across all outputs — become a

common factor dominating all audit tasks, and human generational heterogeneity cannot

override them (Section 5). The organizational design requirement is to use in parallel, at least

for outputs with grave consequences, cross-validation by models of different architectures

and developers. The method for incorporating this model factor into the measurement of

detection-overlap rates is treated in Section 9.

The bulwark against cognitive hold-up. Flattening the power gradient has a reverse side.

A regime in which observations are not suppressed simultaneously requires a bulwark

against the incentive to oversupply observations. Experiential auditing has no established

market price (Section 8.2), and under the information asymmetry in which the

commissioning side can hardly verify directly the validity of a flagged risk, auditors — above

all non-employment seniors whose contract continuation depends on the impression of the

"usefulness" of their observations — face a structural incentive to perpetuate their own status

and contracts by continuing to flag fictitious or inflated risks. This is a hold-up inherent in the

agency relation of auditing, and it requires no malice — excess caution is, in one's own

introspection, indistinguishable from conscientiousness. Three bulwarks can be specified.

First, exclude solo auditing by a single senior: a solitary auditor becomes, in effect, the sole

evaluator of their own observations, and the information asymmetry is maximized. Second,

make independent double-checking by multiple heterogeneous seniors the basic form: the

multiple placement of heterogeneous auditors that Proposition 5 required for the

decorrelation of misses performs here a second function as the decorrelation of over-flagging

— the likelihood that independent, heterogeneous auditors flag the same fictitious risk is low

(an application of Proposition 5). Third, periodic blind evaluation of audit validity: an

operation that blindly mixes in known errors and unproblematic outputs and periodically

measures detection rates and over-flagging rates is the operational version of the quantity

that Hypothesis H3 measures as its dependent variable (iv), the over-rejection rate (Section 9),

and it re-anchors the evaluation of auditors on the validity, not the quantity, of observations.

An audit regime lacking these bulwarks can, in the name of decorrelation, multiply

verification costs without limit. This cost ledger connects to the net-benefit framework of

Section 6.5.

Working Paper | Ageless Management in the AI Era 99

It should also be made explicit that generation is not the only axis of the portfolio. The

heterogeneities that can produce error decorrelation include domain heterogeneity (failure

cases from different industries), functional heterogeneity (legal, technical, front-line),

training-environment heterogeneity (pre-AI manual training versus AI-native training), and

heterogeneity of lived experience (Section 6.4). Generational heterogeneity can be understood

as bundling, along the time axis, the heterogeneity of training environments and of failure

cases among these. The design of the oversight portfolio is therefore not an optimization

problem of generational composition but a multi-axis composition problem that minimizes

the overlap of blind spots, in which generation merely provides one easily observable axis.

Also, though a macro-level finding, Zélity's (2023) finding of an optimal level (a hump shape)

in the relation between age diversity and productivity is suggestive for oversight portfolios as

well. Since the benefits of heterogeneity are eaten by coordination costs and capacity

constraints, the size of the portfolio and the degree of diversification have an optimum, and

the design goal is "optimization," not "maximization." The location of this optimum cannot be

derived from theory and is left to organization-specific measurement (Section 9).

The existing empirical evidence on the value of diversity in error detection should be

confirmed honestly. Direct peer-reviewed evidence that age diversity improves error

detection or audit performance is not found within the scope of this paper's search. What

exists is indirect evidence. Sommers (2006), in a mock-jury experiment, reported that racially

diverse groups exchanged a wider range of information than homogeneous groups and that

factual errors by the majority members themselves decreased — a finding showing that

diversity can reduce errors not only through the pathway of "bringing in different

information" but also through "making the majority's information processing more careful";

but this is an experiment on racial diversity, and replication with age is not confirmed. Also,

Börsch-Supan & Weiss (2016), from unique data connecting production-process errors on

assembly lines with worker attributes, reported that no evidence was obtained supporting

the conventional view that productivity declines with age; and by the summary of the

National Academies (2022), older workers make slightly more minor errors but grave errors

are rare, and experience offset the productivity decline. But this is an effect of experience (or

of age), not an effect of diversity, and the two must not be conflated. Proposition 5 is

consistent with these pieces of indirect evidence, but it is an untested theoretical proposition

not supported by them.

6.4 Modes of Connection with NPOs and Educational Institutions

The multigenerational ecosystem of Definition 7 is not limited to the firm's own boundary of

employment. Youth participation typically passes through PBL-type educational partnerships

and contest-type schemes; the participation of socially marginalized groups through NPOs

and intermediary support organizations; and the participation of super-seniors through

outsourcing and advisory contracts. The institutional conditions of these modes of connection

(labor law, social security, freelance protection) are treated in Section 7. This section confirms

Working Paper | Ageless Management in the AI Era 100

the empirical evidence on the effectiveness of connection. To state the conclusion first: the

evidence converges on a single point — connection is possible, but not unconditional.

First, connection with educational institutions. PBL-type partnership is a mode that

connects youths' ideation and prototyping capacity to the ecosystem within the frame of

educational purpose, and it is also institutionally required as a form of participation not

based on an employment contract (Section 7; Proposition 10). But the finding of Bastani et al.

(2025), seen in Section 4, gives a direct warning here: use of generative AI without guardrails

greatly raised assisted performance while harming independent learning outcomes. A PBL

design with primacy of educational purpose is therefore not only incompatible with a design

whose main purpose is the firm's acquisition of deliverables; it must also include the design

of AI use (guardrails; secured unassisted practice) as a requirement on the educational side.

Next, connection with NPOs and intermediary support organizations. As organizational

forms bearing the work integration of people who face barriers to employment, peerreviewed

research has accumulated, centered on work-integration social enterprises (WISEs),

and there exist syntheses of the knowledge on work integration of people with brain injury,

mental illness, and intellectual disability (Kirsh et al. 2009) as well as reviews of evaluation

frameworks. This field, however, is dominated by case studies and qualitative research, and

controlled effectiveness studies are few. From the standpoint of ecosystem design, what can

be expected of NPO partnership extends to the existence of participation pathways and the

accumulation of operational know-how — not to a quantitative guarantee of outcomes.

On assistive technology (AT), held to be the technical basis of connection, a systematic

review exists. Marinaci et al. (2023) systematically reviewed 41 publications from 2017

onward and concluded that assistive technologies show effectiveness for overcoming

accessibility barriers, improving job performance, independence, and expanding career

opportunities. The review itself, however, states as limitations its geographic skew and the

lack of research on the emotional and sociocultural dimensions of AT use in the workplace.

As a further important qualification, the evidence that AT and AI-based cognitive-accessibility

technologies (voice UIs, summarization, read-aloud) support job performance differs in level

from the evidence that they improve labor-market outcomes — job acquisition, retention, and

wages. Peer-reviewed evidence directly measuring the latter is not established within the

scope of this paper's search. The cautionary view — "too much promise, yet too little

substance" (Smith & Smith 2021) — remains valid.

And the currently most precise empirical study testing the relation between AI and the

employment of people with disabilities on an international panel does not permit optimism.

Omri et al. (2025), using data from 27 high-technology advanced countries over 2006–2022,

estimated with a moderated mediation analysis the effect of AI adoption (proxied by the

number of industrial robots installed) on unemployment among people with disabilities. The

direct effect was an increase in unemployment, and the authors' initial hypothesis (that AI

reduces unemployment) was rejected. The only significant unemployment-reducing pathway

Working Paper | Ageless Management in the AI Era 101

was the indirect effect through higher education (coefficient −0.0322, 95%CI [−0.0550,

−0.0151]); the indirect effect via basic and secondary education was nonsignificant. Moreover,

a counterintuitive moderation was estimated whereby the higher the quality of governance,

the more the unemployment-increasing effect of AI is amplified; the simple expectation that

"good institutions cancel out AI's harms" is not supported either. Even allowing for the study's

limitations (the proxy variable is pre-generative-AI robot adoption; causal identification is

weak), the implication to draw is clear. AI does not advance inclusion unconditionally.

Inclusion is a possibility that opens only when educational investment, accessible design, and

the conditional design of modes of connection are all in place — which is the very reason

Proposition 11 states expressly that "the conversion is not automatic."

These pieces of evidence bring into view the design elements demanded in common of the

modes of connection with NPOs and educational institutions. First, the translation and

correction function of the intermediary organization. NPOs and educational institutions bear,

between participants and firms, functions that individuals cannot bear themselves: securing

the cognitive accessibility of tasks, negotiating conditions, supervising educational purpose.

Designing the mode of connection is, in substance, designing this intermediary function.

Second, the explicit statement of conditions. Since inclusion through AI depends on the

conditions of education, design, and institutions, the agreement of connection must state

explicitly the design of AI use (guardrails; unassisted opportunities), compensation and

attribution, and the ceiling load of participation (Table 9 in Section 8). Third, the calibration

of outcome expectations. Measuring the initial outcomes of connection by employment

outcomes exceeds what the current state of the evidence confirms (the level of support for job

performance) and readily invites withdrawal through disappointment. Starting measurement

at the level of job performance, continuation of participation, and brain-capital indicators is

what is consistent with the evidence (Section 9).

IP defense (appropriability). For connection beyond the boundary of employment, the

price from the firm's standpoint must also be stated. As noted in Section 1, the theoretical

backbone of this paper lies in the lineage of the resource-based view (Barney 1991) and

dynamic capabilities (Teece, Pisano & Shuen 1997), and experiential audit capacity can be a

source of competitive advantage because it is rare and hard to imitate — the immobility of

the resource. Yet the ecosystem of Definition 7 works precisely in the direction of weakening

this immobility. Under the open modes of connection — outsourcing, advisory, PBL, NPO

partnership — the tacit knowledge transferred to participants in the course of auditing — the

business's criteria of judgment, interpretations of failure cases, organization-specific context

— can leak beyond the boundary of employment through participants' departure or

simultaneous participation in other firms (knowledge spillover). An organization that

depends on outside participants for the source of its competitive advantage structurally faces

the question of whether it can appropriate the returns from that source (appropriability).

Contractual instruments such as NDAs are necessary, but the leakage of tacit knowledge

cannot be fully captured by contract. This paper therefore specifies, in addition to contract,

Working Paper | Ageless Management in the AI Era 102

two architectural controls as complementary governance. First, the modularization and

encapsulation of interpretive context: partition the contextual information necessary for

auditing into engagement-level modules, and disclose to each participant only the scope

necessary for their assigned component — the context-partition protocol of Section 6.2 can

share the same partition with this appropriability control. Second, federated management of

access rights to the RAG layer: compartmentalize the read and write permissions to the

accumulation layer of interpretive context (the RAG layer) treated in Section 8.1 by

participant and by engagement, so that no single participant can reach the full picture of the

organization's interpretive context. These controls are an operation that re-seats

organization-specific contextual integration — the orchestration that connects the

participants' knowledge, rather than the knowledge of the individual participants — in the

seat of inimitability, a design for reconciling open participation with the defense of

appropriability. Excessive compartmentalization, however, thins the very supply of context

necessary for auditing and lowers detection probability. The adjustment of this trade-off

between the defense of appropriation and the effectiveness of auditing is likewise entered in

the net-benefit ledger of Section 6.5.

Two related points should be added. First, bias in AI hiring tools. The systematic scoping

review of El Morr et al. (2024), while indicating that AI as assistive technology can enhance

the lived experience of people with disabilities, pointed out that all 34 of the reviewed AI

model-building studies neither measured nor addressed bias, and that AI hiring tools

perpetuate discrimination against people with disabilities. A selection AI placed at the

entrance of the ecosystem can close off the mode of connection itself. Second, the relation

between inclusion and firm performance. A corporate survey exists correlating leading firms

in disability inclusion with high performance (Accenture 2018), but it is grey literature

without peer review, a correlation on a self-selected sample, and it cannot exclude reverse

causality (high-performing firms can afford to invest in inclusion). The rollout of

neurodiversity employment programs (SAP, Microsoft, and others) is a fact (Krzeminska et al.

2019), but peer-reviewed independent effectiveness studies are limited. This paper does not

use these as grounds for the proposition that "inclusion pays."

Evidence-grade note: Accenture (2018) is a corporate survey (grey literature), and its figures are not cited

in the text. The WISE literature is centered on case and qualitative research, and individual effect sizes

are not cited in this paper. Omri et al. (2025) is a peer-reviewed empirical study, but it is a panel

mediation analysis with weak causal identification, and its data predate the diffusion of generative AI

(through 2022). The statements of this section on the inclusion effects of generative AI stand inside these

qualifications.

6.5 The Connection to FVT — As a Conditional Pathway

The original draft argued that organizations with diversity of thought are more likely to

produce discontinuous innovation, and that this leads to firm value in the sense of Future

Working Paper | Ageless Management in the AI Era 103

Value Theory (FVT; Kadowaki 2026a). This section revises that claim into a conditional form,

in the light of the empirical baseline.

First, the baseline does not support optimism. The conclusions of the meta-analyses on the

relation between age diversity and team outcomes agree near zero. Joshi & Roh (2009), in a

meta-analysis of 39 studies and 8,757 teams, estimated the direct effect of age diversity on

performance at r = −.06 (k = 21, 95%CI [−.09, −.04]). Schneid et al. (2016), in a meta-analysis of

74 studies, reported that the overall relation between age diversity and team outcomes was

nonsignificant and that the only significant relation was increased turnover (r = .11),

themselves positioning this as a refutation of arguments emphasizing age diversity. In the

most recent and largest registered-report meta-analysis, Wallrich et al. (2024) (615 reports,

2,638 effect sizes), the overall effect of demographic diversity is r = .014, effectively zero, and

age diversity alone is nonsignificant. The unconditional claim that "being multigenerational

raises value" is incompatible with the totality of the existing evidence. This paper does not

make that claim.

On the other hand, the same body of evidence also identifies the conditions under which

effects appear. Backes-Gellner & Veen (2013) showed, from large-scale German linked data,

that age diversity has a positive effect on firm productivity when, and only when, the firm is

engaged in creative rather than routine tasks. In the moderator analysis of Wallrich et al.

(2024) as well, the diversity–performance relation is more positive for tasks that are high in

complexity and whose outcomes depend on creative divergence. The climate conditions of

Wegge et al. (2012) were stated in Section 6.2. At the macro level, moreover, Zélity (2023)

shows that the relation between age diversity and GDP per capita is hump-shaped, that is,

that an optimal level exists (a country-level aggregate relation, not directly transferable to

intra-firm formation). Taken together, the claim that can be written is this: age diversity can

contribute to outcomes when the task is complex and creative, when there is a climate of low

age discrimination in which diversity is positively appraised, and when age-inclusive

management accompanies it. And diversity is not monotonically good; it has an optimum.

This paper's theoretical wager is to add AI-mediated complementarity to this list of

conditions. Proposition 6 claims that, taking as given the known conditions of task complexity

and inclusive climate, AI-mediated complementarity (Definition 6) is an additional moderator

of the relation between age and experience diversity and organizational outcomes, and that

in its presence the relation moves in a more positive direction. What Proposition 6 claims is

not the absolute level of the relation (that it will necessarily be positive) but a comparative

static: the difference in the relation with and without AI mediation. The near-zero baseline of

the existing meta-analyses was measured almost entirely in environments lacking AI

mediation, and on this paper's reading it is compatible with Proposition 6. Against the

standard explanation that the benefits of diversity (the complementarity of perspectives and

knowledge) are eaten by coordination costs (communication, friction), if AI lowers the costs

of the Gf-type components and of translation and brokerage, the balance of benefits and costs

can move — this is the theoretical ground of Proposition 6. It is, however, untested, and it will

Working Paper | Ageless Management in the AI Era 104

be adjudicated only by comparisons that manipulate the presence of AI mediation (the

refutation condition of Proposition 6; Hypothesis H2). As for the diversity theorem of Hong &

Page (2004), often invoked in this context, its qualifications must be stated. The theorem is a

result of a computational model under specific conditions, and its mathematical generality

has been criticized (Thompson 2014). Moreover, the "diversity" of the theorem is cognitive

and functional diversity, and the empirical evidence that age diversity proxies for it is weak.

The pathway age → heterogeneity of experience and technological environment →

(conditionally) cognitive complementarity is a hypothesis in this paper as well.

On how outcomes are to be measured, the framework of net benefit and transaction costs is

made explicit here. Multigenerational auditing and verification are not free of charge. The

processing and coordination of heterogeneous observations delay decision-making;

independent double-checking (Section 6.3) increases verification effort; and the portfolio

formation for decorrelation itself carries coordination costs. These are the transaction costs

of multigenerational formation. On the other hand, the principal benefits of auditing — the

avoidance of collapse through excessive risk, the reduction of rework — appear only weakly

in the immediate evaluation of outputs. If this asymmetric manifestation of benefits and costs

is left unaddressed, the evaluation of multigenerational formation can be manipulated

arbitrarily through the choice of measure. It is to foreclose this manipulation that Proposition

6 defines outcomes as "net benefit — not idea quality alone, but including the avoidance of

excessive risk and the reduction of rework, and net of the cost of the decision time required

for verification." That is, the claim of the superiority of multigenerational formation is made

only on the ledger of risk-adjusted, verification-cost-deducted net benefit, not on the single

indicator of idea quality. That Hypothesis H2 requires recording time to decision and

verification effort as cost variables is the measurement-side counterpart of this definition

(Section 9).

The ledger of transaction costs also demands the consideration of pathways that lower

costs. If the infrastructure this section and Section 8 specify — the maintenance of

characteristic measurement and role descriptions, audit assignment and context partition,

the sampling management of dual-track auditing, the compartmentalization of the RAG layer

and the preservation of its diversity, the standardization of contracts and compensation —

were all built in-house by each individual firm, the fixed cost would not be small.

Organizations of a scale at which net benefit exceeds the fixed cost are limited, and the

applicability of this model shrinks to that range — this limit of external validity is confronted

head-on in Section 10.5. On the other hand, much of this infrastructure — measurement

tasks, audit-assignment algorithms, standard contract templates, sampling designs — is low in

firm-specificity and lends itself to standardization. If, therefore, the measurement and

governance infrastructure were provided to multiple organizations as common

infrastructure (standardized SaaS-type services), the fixed costs of development and

maintenance would be shared among the user organizations, and the marginal cost per firm

could fall. If this pathway materializes, the lower bound of the organizational scale at which

Working Paper | Ageless Management in the AI Era 105

this model is applicable comes down. This, however, is the presentation of a possibility, not an

assertion. Whether such a service market actually forms, and how far standardization is

compatible with each organization's contextual specificity, are both empirical questions not

derivable from this paper's theory, and they should be read paired with the statement of

limitations in Section 10.5.

Empirical findings that should be written separately from team outcomes concern

intergenerational knowledge transfer. That knowledge transfer in age-mixed dyads raises the

motivation and organizational commitment not only of the receiver but of the sender

(Burmeister et al. 2020), that reverse mentoring can yield skill development for both younger

mentors and older learners (Kaše et al. 2019), and that intergenerational learning is

bidirectional (Gerpott et al. 2017) all have peer-reviewed evidence. These are pathways — set

apart from performance effects — to retention, skill formation, and the formation of the

brain-capital stock (K), and they constitute adjacent evidence for Proposition 8 (bidirectional

capital formation). Not conflating "the performance effects of diversity" and "the effects on

knowledge transfer and retention" into a single virtue is the descriptive discipline of this

area.

Finally, the scope of the connection to FVT is demarcated. FVT (Kadowaki 2026a) is a

framework that inverts the origin of firm value from past accumulation to the creation of

future value, and the organizational conditions that raise the probability of the occurrence of

discontinuous innovation are its central concern. Two pathways can be considered by which

the design argument of this section connects to FVT. The first pathway is that

multigenerational formations satisfying the conditions raise the production of ideas that

combine novelty and feasibility; this is tested directly as Hypothesis H2. The second pathway

is that social impacts — the lightening of the social-security burden through the continued

participation of super-seniors, the lifetime-income effects of early youth participation — are

reflected in firm value via ESG evaluation and the cost of capital. A proposal exists to expand

investment in late-life brain capital as an object of ESG investment (Dawson et al. 2022), but it

is a proposal paper, and to this paper's knowledge there is no empirical study of this pathway.

The second pathway is stated in this paper only as a pathway hypothesis, and no value claim

premised on its holding is made. The original draft's expression that the multigenerational

ecosystem is a "robust infrastructure" of Future Value is replaced, in this paper's framework,

by "a conditionally testable design hypothesis."

The design argument of this section can be summarized as follows. (i) The role placement

of Table 5 is a default design, not fixed roles; its operation is dynamic role allocation, which

includes the pair of indirect audit feedback and reverse mentoring, the cognitive

apprenticeship loop with micro-ownership (Proposition 8) and its responsibility layering

(HOTL confined to coaching, with attributed responsibility on the junior), multi-role

concurrency (identity buffer — the engagement protection of Proposition 7), and the context

partition for hybrid domains (the operationalization of boundary condition (ii) of Proposition

4). (ii) The supervisor group is formed as a portfolio with the correlation of errors as a design

Working Paper | Ageless Management in the AI Era 106

variable, but that formation is subject to WP8's capacity constraints, carries no guarantee of

detection probability, and requires the three conditions of Proposition 5 — an unmediated

channel (a randomly sampled raw-audit slot), the flattening of the power gradient, and

diversity on the AI side of the audited objects — together with the bulwarks against cognitive

hold-up (independent double-checking; periodic blind evaluation of audit validity). (iii)

Connection beyond the boundary of employment is possible but not unconditional, requiring,

in addition to the conditional design of education, accessibility, intermediary functions, and

compensation, the defense of appropriability (the modularization of interpretive context and

the compartmentalization of RAG-layer access rights). (iv) The outcome effects of

multigenerational formation are conditional; AI-mediated complementarity is this paper's

theoretical addition to that list of conditions, and the claim of its superiority is made on net

benefit after deducting the transaction costs of verification — with the pathway of lowering

that cost through common infrastructure left open as a possibility (Section 10.5). And all of

these designs are sustainable only on the premise that participants' brain capital is protected.

That the institutional conditions of participation are treated in the next section (Section 7),

and the design criteria of protection in the one after (Section 8), follows from this order.

7. Institutional Hurdles: An International Comparison

The theory and design arguments of the preceding sections have been, so to speak, matters

internal to the organization. The implementability of Ageless Management (Definition 1),

however, is strongly conditioned by the legal institutions outside the organization. Even if an

organization seeks to remove calendar age from the allocation of roles, if the pension system

imposes a de facto marginal tax rate on work beyond a certain age, and if the choice of

contractual form determines the presence or absence of protection, dynamic role allocation

(Proposition 12) runs into an institutional wall. For three settings — employment-based work

at older ages (Section 7.1), non-employment participation at older ages (Section 7.2), and

youth participation (Section 7.3) — this section compares the institutions of Japan, the United

States, Germany, the EU, Singapore, South Korea, and the ILO (Section 7.4, Table 6), and

confirms the institutional foundations of Proposition 9 (the protection vacuum of nonemployment

forms) and Proposition 10 (the symmetry of Brain Safety) (Section 7.5).

Two caveats at the outset. First, this section is a description and comparison of institutions,

not legal advice. The application of any particular institution is determined by the facts of the

case and the latest statutes and circulars, and practical decisions require consultation with

professionals. Second, the statute names, article numbers, effective dates, and monetary

amounts in this section are limited to matters checked against primary sources — the e-Gov

statute database, the Ministry of Health, Labour and Welfare, the Japan Pension Service, the

U.S. Social Security Administration (SSA), the German Pension Insurance (DRV), EUR-Lex, and

the ILO — or reliable professional commentary (evidence grade: primary legal and

administrative sources). Details that could not be so verified are not stated.

Working Paper | Ageless Management in the AI Era 107

7.1 Employment-Based Work at Older Ages — Mandatory Retirement,

Employment Security, and Pension Work Disincentives

Japan — the Two-Tier Structure of the Act on Stabilization of Employment of Elderly

Persons and the Emergence of Non-Employment Options

The basic statute governing employment-based work at older ages in Japan is the Act on

Stabilization of Employment of Elderly Persons, etc. (Act No. 68 of 1971; commonly, the Act on

Stabilization of Employment of Elderly Persons). Article 8 of the Act provides that where a

mandatory retirement age (teinen) is set, it may not be below 60 (with an exception for work

in which elderly persons have difficulty engaging). Article 9 obliges employers that set a

mandatory retirement age below 65 to take one of the following employment-security

measures for elderly persons: (i) raising the retirement age, (ii) introducing a continuedemployment

system, or (iii) abolishing the mandatory retirement age. This is a legal

obligation.

By contrast, Article 10-2, newly established by the amendment under Act No. 14 of 2020,

prescribed the securing of work opportunities from age 65 to 70 as an obligation to make

efforts (effective April 1, 2021). The difference in the nature of the obligations — the measures

up to 65 (Article 9) being a legal obligation, the measures up to 70 (Article 10-2) an effort

obligation — is important, and the age-70 measures are not a mandate of retirement at 70.

The Ministry of Health, Labour and Welfare, too, states explicitly that the amendment does

not obligate raising the mandatory retirement age to 70.

From this paper's standpoint, what most deserves attention is the composition of the menu

of measures for securing work up to age 70. There are five options: in addition to (i) raising

the retirement age to 70, (ii) abolishing the mandatory retirement system, and (iii)

introducing a continued-employment system up to 70 (reemployment or extended

employment), they include (iv) a scheme for continuously concluding outsourcing contracts

up to 70, and (v) a scheme under which the person can continuously engage up to 70 in the

employer's social-contribution projects (or the social-contribution projects of entities the

employer commissions or invests in). Options (iv) and (v) are called the entrepreneurshipsupport

measures (Article 10-2), and options for securing work outside employment were

thereby given explicit statutory standing. That is, Japan's older-age employment legislation

has, for the 65–70 phase, officially recognized participation forms outside the employment

boundary as formal policy instruments. This points in the same direction as the participation

structure spanning employment, outsourced engagement, and social-contribution projects

envisaged by Definition 7 (multigenerational ecosystem). Introducing the entrepreneurshipsupport

measures, however, requires a procedure of preparing a plan and obtaining the

consent of a majority union or equivalent, and, as discussed below (Section 7.2), the measures

carry a protection vacuum precisely because they are non-employment forms.

Working Paper | Ageless Management in the AI Era 108

Japan — the In-Work Old-Age Pension Offset and the April 2026 Revision

The other institutional variable of employment-based work at older ages is Japan's in-work

old-age pension offset (zaishoku rลrei nenkin) — the in-work suspension of the old-age

employees' pension under the Employees' Pension Insurance Act. The mechanism is as

follows. When the sum of the basic monthly amount of the old-age employees' pension and

the total-remuneration-equivalent monthly amount exceeds the suspension-adjustment

amount (the threshold), one half of the excess is suspended (suspended amount = (basic

monthly amount + total-remuneration-equivalent monthly amount − threshold) ÷ 2). This

structure, in which increased earnings from work bring a reduction of the pension, has long

been identified as an institutional factor generating work adjustment among older persons —

the perception that "working means losing."

This institution changes substantially in April 2026. Under the Act Partially Amending the

National Pension Act, etc. for Strengthening the Functions of the Pension System in Light of

Social and Economic Changes (Act No. 74 of 2025; enacted June 2025), the threshold of the inwork

old-age pension offset is raised, effective April 1, 2026. The monetary relationships

require precision. The pre-amendment threshold was at the 500,000-yen level (in actual

terms, 510,000 yen for FY2025). The amending Act raised this to 620,000 yen, but this 620,000

yen is the statutory value based on wage levels at the time of the Act's enactment in 2025.

Because the threshold is automatically revised each fiscal year in line with wage movements

(wage indexation), the threshold actually applied in FY2026, the fiscal year of entry into force,

became 650,000 yen. That is, "from 510,000 yen to 620,000 yen" is the statutory increase (at

2025 prices), while "650,000 yen" is the actual amount applied in FY2026 after wage

indexation; the two are not in contradiction. The Japan Pension Service's dedicated page

likewise states the figures as 510,000 yen per month before the amendment (FY2025) and

650,000 yen per month after (FY2026).

One further point requires precision. Persons aged 70 or over who work at establishments

covered by employees' pension insurance have lost insured status under employees' pension

insurance upon reaching 70 and therefore bear no premiums. The suspension under the inwork

old-age pension offset, however, continues to apply to employed persons aged 70 or

over under the same formula (for employed persons aged 70 or over, it is calculated using an

amount corresponding to the standard monthly remuneration). The understanding that "past

70 one falls outside the in-work old-age pension offset" is mistaken, and the end of premium

liability must not be confused with the continued application of the suspension. For the

employment-based participation of the super-seniors (aged 60 to their 90s) envisaged by

Ageless Management, this suspension is an institutional variable that continues beyond age

70.

United States — Abolition of the Earnings Test at and after FRA, and the ADEA

The United States is a representative example of a contrasting institutional choice. The

retirement earnings test on Social Security old-age benefits was a mechanism that withheld

Working Paper | Ageless Management in the AI Era 109

benefits when earnings from work exceeded exempt amounts, but the Senior Citizens'

Freedom to Work Act of 2000 (Public Law 106-182, signed April 7, 2000) abolished the

earnings test at and after full retirement age (FRA). From FRA onward, old-age benefits are

not reduced no matter how much one earns from work.

One must not, however, simplify this to "the United States abolished the earnings test." For

beneficiaries before FRA the earnings test remains in force. In years before the year of FRA

attainment, $1 of benefits is withheld for every $2 of earnings above the lower exempt

amount, and in the year of FRA attainment (through the month before the attainment

month), $1 is withheld for every $3 of earnings above the upper exempt amount. The SSA's

official 2026 exempt amounts are $24,480 per year (lower) and $65,160 per year (upper).

Moreover, because withheld benefits are effectively recovered through benefit recomputation

after FRA attainment, writing "forfeited" is also inaccurate. The institution's implication lies

in the point that it is a design in which the work-disincentive effect disappears at the age

boundary of FRA.

On the employment-discrimination side, the Age Discrimination in Employment Act of 1967

(ADEA) prohibits age discrimination in employment (hiring, discharge, compensation,

promotion, and the like) against workers aged 40 and over. The upper age limit on protected

status that originally existed was removed by the 1986 amendment, whereby mandatory

retirement itself became unlawful for most occupations (with limited exceptions such as

pilots). The United States is the representative country that adopts "no mandatory retirement"

as a matter of legal institution, and it can be called the jurisdiction that has carried the

direction of Definition 1 — removing calendar age as a criterion for exit from roles — furthest

at the level of employment law.

Germany — Complete Abolition of the Earnings Limit on Early Pensions

Germany is the example that has most thoroughly removed work disincentives on the

pension side. The earnings limit (Hinzuverdienstgrenze) applicable while drawing an early

old-age pension was abolished on January 1, 2023 by the Eighth Act Amending Book IV of the

Social Code (8. SGB IV-ÄndG). The official FAQ of the German Pension Insurance (DRV) states

explicitly that this limit has been abolished entirely (ganz entfallen). Since then, even those

drawing an early pension before the statutory pension age can work without limit while

receiving the full pension, regardless of the amount earned. It is a permanent measure and

applies to all recipients irrespective of when the pension began. Until the abolition, through

2022, an annual limit of 46,060 euros per year (a level raised under COVID special rules)

applied. This abolition, however, concerns early old-age pensions: the earnings limits on

disability pensions (Erwerbsminderungsrente) have not been abolished but have shifted to

dynamic limits linked to wage trends. One must not generalize to "Germany has abolished all

pension earnings limits."

At the EU level, the Employment Equality Directive (Council Directive 2000/78/EC, adopted

November 27, 2000) prohibits direct and indirect discrimination in employment and

Working Paper | Ageless Management in the AI Era 110

occupation on grounds of religion or belief, disability, age, and sexual orientation. Article 6 of

the Directive, however, provides that differences of treatment on grounds of age can be

justified where there is a legitimate aim of employment policy, the labor market, or

vocational training and the means are appropriate and necessary. This is a special exception

granted to age alone among the four grounds of discrimination, and it is the basis provision

for the case law of the Court of Justice of the EU on the permissibility of member states'

mandatory retirement systems. Age is given special treatment, even inside antidiscrimination

law, as a "justifiable distinction" — this asymmetry itself indicates how

institutionally deep-seated the variable of calendar age is.

Singapore and South Korea — Extending Working Lives through Mandates

Asia's leading jurisdictions on older-age work take the route of extending working lives by

strengthening mandates. Singapore's Retirement and Re-employment Act adopts a two-tier

structure of a statutory retirement age (below which forced retirement is prohibited) and,

beyond it, an age up to which re-employment must be offered. From July 1, 2022 the

retirement age has been 63 and the re-employment ceiling 68, and from July 1, 2026 they are

raised to a retirement age of 64 and a re-employment age of 69 (the government has stated a

policy of raising them to 65 and 70 by 2030). In imposing an obligation to offer re-employment

rather than an obligation to retain employment, the structure is close to Japan's continuedemployment

system.

South Korea, through the 2013 amendment of the Act on Prohibition of Age Discrimination

in Employment and Elderly Employment Promotion (the retirement-age extension act,

enacted June 2013), made it mandatory to set the retirement age at 60 or above. It was phased

in from 2016 for workplaces with 300 or more employees and public institutions, and from

2017 for those with fewer than 300. The Act also prescribes the prohibition of employment

discrimination on grounds of age. In recent years, a further extension of the statutory

retirement age (to 65) has been under discussion between labor and management and in the

National Assembly.

These employment-based institutions form a three-stage spectrum in the design of in-work

pensions. Japan maintains partial suspension while raising the threshold (2026), the United

States abolished the test at and after FRA (2000), and Germany abolished it entirely, including

for early claimants (2023). The direction is in every case a shrinking of work disincentives,

but the end points differ. This point is taken up again in Section 7.5.

7.2 Non-Employment Participation at Older Ages — the Protection Vacuum and

Its Partial Filling

The participation forms on which Ageless Management depends are not limited to

employment. Definition 7 (multigenerational ecosystem) explicitly includes outsourced

engagement, advisory roles, PBL-based educational partnerships, and NPO partnerships, and

the design arguments of Section 6 conceived many of the audit and advisory roles of super-

Working Paper | Ageless Management in the AI Era 111

seniors in non-employment form. As seen in Section 7.1, Japan's Act on Stabilization of

Employment of Elderly Persons has itself formalized a non-employment option in the

entrepreneurship-support measures. Here, however, lies a structural problem. Most of the

protections of labor law and social security law are designed with the employment

relationship (the labor contract) as their unit. Workers in outsourced, advisory, or gig-type

arrangements, because they are not based on labor contracts, are in principle not reached by

the protections of labor law, beginning with the automatic application of the Labor Standards

Act and workers' accident compensation insurance. The entrepreneurship-support measures

extended work opportunities into non-employment, but protection does not automatically

follow. Proposition 9 (the protection vacuum of non-employment forms) refers to this

structure: current institutions are designed with the employment relationship as the primary

unit of protection and restraint, non-employment participation falls into a vacuum of

institutional protection, and Ageless Management that leaves the vacuum unattended can

degenerate into exploitation.

Overlaid on this vacuum, beyond the negative problem of absent protection, is the positive

problem of an asymmetry of institutional incentives. As confirmed in Section 7.1, the

suspension under the in-work old-age pension offset does not end at 70. Even after a person,

upon reaching 70, loses insured status under employees' pension insurance and ceases to

bear premiums, so long as the person works in employment at an establishment covered by

employees' pension insurance, the suspension continues to apply under the same formula, as

an "employed person aged 70 or over," based on the basic monthly amount and the totalremuneration-

equivalent monthly amount (calculated using an amount corresponding to the

standard monthly remuneration). If, on the other hand, the same person shifts to nonemployment

work under an outsourcing contract — the entrepreneurship-support measures

of Article 10-2 of the Act on Stabilization of Employment of Elderly Persons — the person does

not qualify as an employee under employees' pension insurance and is therefore not subject

to the suspension of the in-work old-age pension offset. That is, under current institutions

there exists a structure of institutional arbitrage in which, even where the same individual

supplies the same kind of services, whether the pension is suspended or not divides

according to whether the contractual form is employment or outsourcing. This asymmetry

financially steers organizations and individuals who would supply Gc-type roles at

remuneration levels above the threshold in employment form toward moving from the

contractual form with thicker protection to the contractual form with thinner protection.

With the formalization of opportunity (the entrepreneurship-support measures) and the

pension arbitrage pointing in the same direction, the shift to non-employment is not an

institutionally neutral choice but an induced one. What makes the protection vacuum of

Proposition 9 grave is that the destination of this inducement is precisely the place where

protection is thinnest.

The reverse side of this arbitrage is the risk of employee misclassification — so-called

disguised subcontracting. Even if the contract is titled an outsourcing agreement, whether a

Working Paper | Ageless Management in the AI Era 112

person qualifies as a worker is judged by substance. The basic framework of administrative

interpretation, the Report of the Labor Standards Act Study Group, "On the Criteria for

Determining 'Worker' Status under the Labor Standards Act" (December 19, 1985), sets out a

framework of holistic judgment centered on the subordination test — the presence or

absence of freedom to accept or refuse work requests and instructions to engage in tasks, the

presence or absence of direction and supervision in the performance of work, the presence

or absence of constraints on place and hours of work, and the substitutability of the labor

supplied — and the remuneration's character as compensation for labor, with the presence or

absence of entrepreneurial character, the degree of exclusivity, and the like as reinforcing

elements; the judgment of the relationship of use and subordination in the internship

circular examined in Section 7.3 (Circular Kihatsu No. 636) belongs to the same lineage.

Accordingly, if a super-senior on an outsourcing contract is obligated to engage in constant

audit work, has working hours and place managed, is not allowed to accept or refuse

individual requests, and has the performance of work placed under the organization's

direction and command, worker status can be found as a matter of substance regardless of

the name of the contractual form. The consequences do not stop at employment

responsibilities — application of the Labor Standards Act, the Minimum Wage Act, and

workers' accident compensation insurance — and the retroactive incurrence of applicable

social insurance premiums: through qualification as an "employed person aged 70 or over,"

the very pension treatment the preceding arbitrage presupposed can be overturned.

Organizations that design super-seniors' audit and advisory roles in non-employment form

therefore need to build into the design a legal safeguard that aligns the substance of the

contract with the requirements of the non-employment form — this paper calls it the protocol

for the exclusion of direction and control — rather than choosing the contractual form by

looking only at the pension advantage. Concretely: (i) define the unit of engagement not by

hours worked but by tasks and deliverables (such as audit reports on a specified set of AI

outputs); (ii) impose no constraints on the time or place of work; (iii) substantively guarantee

the freedom to accept or refuse each individual engagement; and (iv) issue no concrete

direction or command as to the method of performing the work, giving feedback as

evaluation of deliverables. This is not a technique of evasion but its opposite. If the substance

is constant labor under direction and command, the person should be contracted as an

employee and given the protections of employment; if the non-employment form is chosen,

its discretion and freedom of refusal must be guaranteed in substance, not in name. That the

indirection of audit in Section 6.2 — the design that places the addressee of feedback on the

AI side rather than the person — is consistent with this protocol is no coincidence, and the

ceilings on load and the adequacy of compensation demanded by Brain Safety in Section 8,

only together with this protocol, close the path (Proposition 9) by which the arbitrage-driven

inducement turns into exploitation.

The exclusion of direction and control, however, itself generates a trade-off with

responsibility for quality. An ordering organization that has renounced direction and

Working Paper | Ageless Management in the AI Era 113

command cannot secure audit quality through the management of work performance, and

the discipline of quality migrates to the design of contractual liability. Here neither pole holds.

If unlimited liability for damages arising from an independently contracted super-senior's

erroneous audit — an oversight — is imposed, the individual takes on damage risks orders of

magnitude beyond the engagement fee, and a rational contractor will not accept the

engagement. Older contractors, who have little temporal room to recover losses through

work opportunities or asset formation after damage occurs, have rational grounds to behave

risk-aversely, and this chilling operates all the more strongly. Conversely, complete

exculpation for erroneous audits destroys the incentive to maintain the level of care, and

moral hazard arises. This paper's solution is the combination of two principles. The first is a

contractual cap on liability: the contractor's liability for damages is confined within an agreed

ceiling benchmarked to the engagement fee or the like, keeping it at an assumable level of

risk. The second is the principle that ultimate risk rests with the operating entity. As the legal

analysis of WP8 (Kadowaki 2026h) shows, supervisory responsibility in HOTL rests with the

operating entity, and shifting responsibility onto external auditors does not function as an

externalization of the duty of oversight — audit is an input into the operating entity's

decision-making, not a transfer of decision-making responsibility. Accordingly, of the ultimate

damage arising from errors in AI outputs that slipped past audit, the portion exceeding the

cap on liability is borne by the operating entity. If damages are not the instrument of

discipline, what then controls the moral hazard of erroneous audits? The division of roles is

clear. The allocation of catastrophic risk is the province of contract design including the cap

on liability, while the control of everyday levels of care is the province of the operational

governance of Section 8 — monitoring of each auditor's over-flagging rate (false-alarm rate),

temporary suspension and recalibration of audit authority when thresholds are exceeded,

and the reputation mechanism based on records of audit performance (Section 8.2). This

division of labor, which does not use damage claims as an instrument of everyday control,

prevents the chilling of engagement and, at the same time, places the effectiveness of control

on the side of ex-ante measurable operational indicators rather than on ex-post damages that

are difficult to prove.

In Japan, legislation partially filling this vacuum has recently begun to move. First, the Act

on Ensuring Proper Transactions Involving Specified Entrusted Business Operators (Act No.

25 of 2023; commonly, Japan's Freelance Act) came into force on November 1, 2024. The Act

prescribes, as obligations of ordering businesses when outsourcing work to specified

entrusted business operators (freelancers who employ no workers): (i) clear indication, in

writing or equivalent, of transaction terms (the content of the deliverable, the amount of

remuneration, and so on); (ii) setting a remuneration payment date (within 60 days of receipt

of the deliverable) and payment by that date; (iii) prohibition, in continuing outsourcing

arrangements, of refusal of receipt, reduction of remuneration, returns, unreasonably low

pricing, and the like; (iv) accurate display of recruitment information; (v) accommodation for

balancing work with childcare, family care, and the like; (vi) establishment of systems for

Working Paper | Ageless Management in the AI Era 114

harassment countermeasures; and (vii) advance notice at least 30 days before mid-term

termination and the like of continuing outsourcing arrangements. Jurisdiction lies with the

Fair Trade Commission, the Small and Medium Enterprise Agency, and the Ministry of Health,

Labour and Welfare. It is the first statute to extend, cross-cuttingly, subcontracting-act-style

transaction fairness and work-environment improvement (a partial extension of labor-lawtype

protection) to non-employment workers outside employment labor law, and it can be

called the starting point of the institutional infrastructure for non-employment work.

Second, the scope of special enrollment in workers' accident compensation insurance

(Articles 33 et seq. of the Industrial Accident Compensation Insurance Act) was expanded on

November 1, 2024, the same day the Freelance Act came into force. Freelancers engaged, as

specified entrusted business operators, in business performed under outsourcing from

enterprises and the like became newly eligible for special enrollment (work of the same kind

entrusted by consumers is also covered), and, going beyond the existing individually

designated sectors (construction, IT freelancers, and so on), essentially all freelancers became

able to enroll voluntarily in workers' accident compensation insurance. Premiums are

entirely self-funded; the Class II special enrollment premium rate is the basic daily benefit

amount × 365 × 3/1000 (0.3%), and the basic daily benefit amount is chosen from 16 tiers

between 3,500 yen and 25,000 yen.

These two pieces of legislation show that the vacuum of Proposition 9 has been recognized

by legislators as well, and that its filling has begun. But the filling is partial. First, what the

Freelance Act provides is fairness of transactions, not a floor on remuneration corresponding

to the minimum wage. Clear indication of the remuneration amount is mandated, but the

adequacy of its level is left to the market. Second, special enrollment in accident insurance is

voluntary and fully self-funded in premiums, asymmetric with the automatic application and

employer contribution of the employment form. The decision whether to enroll, and its cost,

are shifted onto the side with weaker bargaining power. Third, no correction of collective

bargaining power (a mechanism corresponding to trade-union law) has been put in place for

non-employment forms. The labor–management consent procedure required by the

entrepreneurship-support measures of the Act on Stabilization of Employment of Elderly

Persons is a control at the entrance to conversion to outsourcing; it does not answer the

continuing disparity in bargaining power after conversion. Accordingly, the refutation

condition of Proposition 9 — the existence of jurisdictions granting non-employment

participants protection equivalent to employment (accident compensation, remuneration

adequacy, and correction of bargaining power) — remains unmet, at least in Japan. The

vacuum has shrunk, but it has not disappeared. The anti-exploitation side of Brain Safety

(Definition 8) in Section 8 seeks to fill this residual vacuum with the organization's design

norms.

Working Paper | Ageless Management in the AI Era 115

7.3 Youth Participation — Child-Labor Regulation and the Design Space for

Education-Based Participation

For youth, who stand at the other end of the multigenerational ecosystem, the character of

the institutions is inverted. Whereas the institutional problem on the older side is the

removal of work disincentives, the institutions on the youth side take protection from labor

as their first principle. This protection is not an obstacle to be circumvented but a premise of

design. Can forms be designed that connect the cognitive assets of youth to the ecosystem

without impairing the protective principle of the priority of education and development —

this is the institutional question on the youth side.

At the base of the international standards lie two ILO conventions. The Minimum Age

Convention, 1973 (No. 138) prescribes a three-tier structure for the minimum age for work.

First, the minimum age may not be below the age of completion of compulsory schooling and

in no case below 15 (Article 2(3)). Second, national laws may permit light work (work not

harmful to health, development, or school attendance) by persons aged 13 to 15 (Article 7).

Third, the minimum age for hazardous work likely to jeopardize health, safety, or morals is

18 (Article 3; an exception in Article 3(3) permits it from 16 subject to protection and

training). For member states whose economic and educational institutions are insufficiently

developed, there is a developing-country exception reading these as 14 in principle and 12–14

for light work (Article 2(4)). The one-sentence summary "the ILO minimum age is 15" is

inaccurate because it drops this three-tier structure (15 in principle, 13 for light work, 18 for

hazardous work). Japan ratified Convention No. 138 in 2000, and the minimum age is

consistent, in a form connected to completion of compulsory education, with Article 56 of the

Labor Standards Act discussed below.

The Worst Forms of Child Labour Convention, 1999 (No. 182) obliges immediate measures

for the prohibition and elimination, for children under 18, of the "worst forms of child

labour": slavery, forced labor, and human trafficking; child soldiers; sexual exploitation; illicit

activities such as drug trafficking; and dangerous and harmful work. On August 4, 2020, the

Kingdom of Tonga ratified as the 187th member state, making it the first convention in ILO

history to be ratified by all member states (universal ratification). Adoption (1999) and the

attainment of unanimous ratification (2020) are distinct events. Japan ratified in 2001, and

the United States, which has not ratified No. 138, has ratified No. 182. The fact that the

protection of children is the domain that has exceptionally reached universal consensus

among labor standards whose ratification status is otherwise divided indicates the

international foundation of the design principle Proposition 10 places on the youth side — the

priority of education and health. At the EU level, the Young Workers Directive (Council

Directive 94/33/EC, adopted June 22, 1994) prescribes the prohibition in principle of work by

children (under 15 or in compulsory education), an exception structure of permits for

cultural and artistic activities, work practice and light work from 14, and light work for

Working Paper | Ageless Management in the AI Era 116

limited weekly hours from 13, together with regulation of night work and hazardous work

for those under 18, giving the same three-tier idea form in EU law.

In Japanese domestic law, Chapter 6 "Minors" of the Labor Standards Act (Act No. 49 of

1947) (Articles 56–63) corresponds to this. Article 56 prohibits the use of a child as a worker

until the end of the first March 31 after the child reaches 15 (there is an exception whereby, in

non-industrial undertakings, children aged 13 or over may be employed outside school hours,

with the permission of the administrative agency, in light work not harmful to health and

welfare, and in film production and theatrical undertakings the same permission makes this

possible even under 13). Article 57 mandates keeping age certificates and the like for persons

under 18, Article 58 prohibits the conclusion of labor contracts by parents or guardians on

behalf of a minor, and Article 59 provides that minors may claim wages independently.

Article 60 prohibits, in principle, overtime and holiday work for minors, and Article 61

prohibits, in principle, night work (from 10 p.m. to 5 a.m.) by persons under 18. Article 62

prescribes restrictions on employment in dangerous and harmful work, and Article 63 the

prohibition of work in mine pits. The range of article numbers requires care: dangerous and

harmful work is Article 62, mine work Article 63, and the whole of minor protection is

"Articles 56–63."

How, then, does youth participation outside employment — internships, PBL (project-based

learning) educational partnerships, contest-based participation — intersect with this

regulation? The key is the determination of worker status. Article 9 of the Labor Standards

Act defines a worker as "one who is employed at a business ... and to whom wages are paid,"

and for internships an administrative circular (Circular Kihatsu No. 636 of September 18,

1997) supplies the framework of judgment. That is, where the benefit or effect of the work in

question accrues to the establishment, as when the student engages directly in production

activity, and a relationship of use and subordination is found between the establishment and

the student, the student qualifies as a worker. Conversely, where the training is observational

or experiential and no relationship of use and subordination is found, as where the student is

not considered to receive business-related direction and command from the employer, the

student does not qualify as a worker. The judgment is made by substance, not by name

(internship, PBL, practicum), and considers (i) the presence or absence of direction and

command, (ii) the accrual of outcomes and benefits to the establishment, and (iii) the

presence or absence of attendance management and sanctions, among other factors. If judged

a worker, the Labor Standards Act, the Minimum Wage Act (the obligation to pay wages at or

above the prefectural minimum wage), and workers' accident compensation insurance apply.

The three-ministry agreement of the Ministry of Education, Culture, Sports, Science and

Technology, the Ministry of Health, Labour and Welfare, and the Ministry of Economy, Trade

and Industry (Basic Approach to the Promotion of Internships, revised 2022) is a policy

document organizing the relationship with recruiting activities; the criteria for worker status

themselves are those of the circular above.

Working Paper | Ageless Management in the AI Era 117

This two-sidedness directly yields the design principles for youth participation. The first

side — one cannot say "student interns may go unpaid." If a student is made to work under

direction and command in a manner whose benefits accrue to the business, that is labor, and

it receives the application of the minimum wage and accident insurance. The procurement of

unpaid labor borrowing the name of education is impermissible both legally and under the

normative argument of Proposition 10. The second side — for observational, experiential, or

learning-type programs designed for educational purposes that avoid incorporation into

operations and direction and command, worker status is denied as a rule, and here lies the

legal space to design education-purpose PBL institutionally as "learning" rather than "labor."

The youth PBL participation conceived in Section 6 — youth-perspective feedback on AIgenerated

outputs, prototyping exercises, contest-style problem submissions — can be

designed inside this space so long as educational purpose is placed first, business use of the

outputs second, and no structure of direction and command is adopted.

This design space, however, has three limits. First, the boundary is continuous, and there is

a constant danger that participation begun as education-based slides into de facto labor as

dependence on its outputs grows. Because worker status is a judgment of substance, the name

given at design time is no rampart. Second, where worker status is denied, young participants

are placed outside the protections of labor law (wages, accident compensation). This is a

protection vacuum of the same shape as the non-employment form on the older side

(Proposition 9), and it needs to be filled by the educational institution's management and

insurance and by the host's duty of care for safety. Third, contest-based participation (prizebased

open calls for problem solutions and the like) adopts a structure in which, out of the

unpaid work products of many participants, compensation goes only to a few winners, so that

designs poor in educational benefit approach the de facto procurement of unpaid labor.

Proposition 10's claim — that the protective principle on the youth side (the priority of

education and health) and the prevention of exploitation on the older side are responses to

the same structure, the asymmetry of bargaining power and exit costs — takes concrete form

in these three limits. Section 8's Brain Safety translates this symmetry into design norms.

7.4 Summary Table of the Institutional Comparison

Table 6 organizes the institutions treated in this section across jurisdictions. The details of

each institution (article numbers, amounts, sources) are left to Table A2 in Appendix B; here,

the implications for the implementation of Ageless Management are mapped to each.

Table 6 International comparison of institutions surrounding older-age and youth work, and their

implications for Ageless Management

Jurisdiction

Institution (effective/

amended year) Key points

Implications for

Ageless Management

Japan Act on Stabilization of

Employment of Elderly

Prohibition of mandatory

retirement below 60 (Article 8).

Legal obligation to secure

Participation forms

outside the employment

boundary formalized

Working Paper | Ageless Management in the AI Era 118

Persons (age-70 measures

in 2021)

employment up to 65 (Article 9).

Effort obligation to secure work

up to 70 (Article 10-2). Options

include outsourcing contracts and

social-contribution projects

(entrepreneurship-support

measures = non-employment

forms)

for ages 65–70. But this

remains an effort

obligation, and nonemployment

forms

carry no accompanying

protection (Section 7.2)

Japan Increase of the threshold

of the in-work old-age

pension offset (April

2026)

Threshold raised from 510,000 yen

(FY2025) to 650,000 yen (FY2026;

statutory value 620,000 yen plus

wage indexation). The mechanism

itself of suspending one half of the

excess is maintained. The

suspension also applies to

employed persons aged 70 or over

The range of "working

means losing" shrinks

but does not vanish.

Supplying highremuneration

Gc-type

roles in employment

form must take account

of the residual workadjustment

incentive

Japan Freelance Act (November

2024); expansion of

special enrollment in

accident insurance (same)

Obligations toward nonemployment

workers: clear

indication of transaction terms,

payment within 60 days, 30 days'

advance notice of mid-term

termination, and so on. Voluntary

special enrollment in workers'

accident insurance expanded to

essentially all freelancers

(premiums self-funded)

Partial filling of

Proposition 9's vacuum.

But there is no

remuneration floor or

bargaining-power

correction, and accident

compensation is

voluntary and selffunded

— the vacuum

remains

United

States

Senior Citizens' Freedom

to Work Act (2000)

Earnings test abolished at and

after FRA. Before FRA, $1 per $2

($1 per $3) of earnings above the

exempt amounts (2026: $24,480/

year; $65,160 in the year of FRA

attainment) withheld (effectively

recovered after FRA)

A precedent removing

work disincentives from

FRA onward. The design

bounded by an age

threshold itself remains

United

States

ADEA (1967; amended

1986)

Prohibition of age discrimination

against those 40 and over.

Removal of the upper age limit

made mandatory retirement

unlawful in most occupations

(limited exceptions)

Legal foreclosure of

forced exit by calendar

age — the example that

carries Definition 1's

direction furthest in

employment law

Germany Abolition of the

Hinzuverdienstgrenze

(January 2023)

Earnings limit while drawing an

early old-age pension completely

abolished (permanent measure).

For disability pensions, not

abolition but dynamic limits

An example of full

abolition of pensionside

work disincentives.

Demonstrates that

complete removal of

"working means losing"

is legislatively possible

EU Employment Equality

Directive 2000/78/EC

(2000)

Prohibition of employment

discrimination including age. But

Article 6 gives age alone a

Age receives special

treatment even inside

anti-discrimination law

Working Paper | Ageless Management in the AI Era 119

justification exception (the basis

provision for permitting

mandatory retirement)

— evidence of the

institutional deeprootedness

of the

variable of calendar age

Singapore Retirement and Reemployment

Act (to 64/69

in July 2026)*

Two-tier structure of statutory

retirement age plus mandatory reemployment-

offer age. Policy of

65/70 by 2030

The gradualist type of

extending work through

mandates. The reemployment-

offer

obligation is structurally

close to Japan's

continued employment

South Korea Mandatory retirement

age of 60 (phased in

2016/2017)

Mandated setting the retirement

age at 60 or above. The same Act

also prescribes the prohibition of

age discrimination. Extension to

65 under discussion

An example of

mandating retirement

ages in progress. The

very existence of a

statutory retirement age

is also a preservation of

the calendar-age

criterion

ILO Convention No. 138

(1973); Convention No.

182 (1999)

Three-tier minimum-age structure

(15 in principle, 13 for light work,

18 for hazardous work;

developing-country exception

14/12). No. 182 ratified by all

member states in 2020

Youth participation

must be designed inside

the protective principle

(priority of education

and health) — the

international

foundation of

Proposition 10

Japan Labor Standards Act,

Chapter 6 on minors

(Articles 56–63); Circular

Kihatsu No. 636 (1997)

Minimum age, prohibition of night

work (10 p.m.–5 a.m.), restrictions

on dangerous and harmful work,

and so on. Worker status of

interns and the like judged by the

substance of the relationship of

use and subordination

Both faces: the legal

space to design

education-purpose PBL

as "learning," and the

risk of sliding into

unpaid labor (Section

7.3)

Note: The formal statute names, act numbers, article numbers, amounts, and sources for each institution are given in

Appendix B (Table A2). This table includes only institutions, verified against primary sources or professional

commentary as of August 2026, that are in force or whose effective date is fixed. *Singapore's 2026 amendment rests

on multiple cross-checks of law-firm and HR-media commentary; final confirmation against primary sources of the

responsible ministry (MOM) remains outstanding (see the evidence-grade note at the end of this section). The

descriptions of the institutions are summaries; the details of application are governed by the respective statutes and

circulars.

7.5 Implications of Propositions 9 and 10 — Ageless Management as a Function

of Institutional Trends

Three implications are drawn from this section's comparison. First, the institutions that have

restrained employment-based work at older ages are moving toward contraction across

jurisdictions. On the work disincentives of in-work pensions, the United States moved in 2000

to abolition at and after FRA, Germany in 2023 to complete abolition including early

Working Paper | Ageless Management in the AI Era 120

claimants, and Japan in 2026 to a substantial increase of the threshold. On the side of

employment opportunity as well, the outer bound of working ages is being extended in each

country: Japan's age-70 effort obligation (2021), Singapore's increases (2026), and South

Korea's mandatory retirement age (2016–2017). Institutions that create "working means

losing" and forced exit by calendar age are, at least as a legislative trend, in retreat. For

Ageless Management, this is a tailwind among the givens.

Second, however, this trend is skewed toward the employment form. As Proposition 9

points out, protection for the non-employment participation on which Ageless Management

depends is incomplete. Japan's Freelance Act and the expansion of special enrollment in

accident insurance (both November 2024) show the beginning of the filling of the vacuum,

but lack all of a floor guarantee on remuneration, correction of bargaining power, and

automatic application of accident compensation. While the entrepreneurship-support

measures (2021) formalized non-employment work opportunities, the very time lag —

protection catching up only partially, three and a half years later — shows that the expansion

of opportunity and the development of protection move at different speeds. Organizations

that build multigenerational ecosystems on non-employment contracts must therefore treat

the law's required level as a floor of protection, not a ceiling, and bear the responsibility of

filling the residual vacuum with their own design — Section 8's Brain Safety. Ageless

Management that neglects this can, as Proposition 9 warns, degenerate into exploitation: on

the older side into unpaid advisorship, and on the youth side into unpaid labor borrowing the

name of education.

Third, the institutions on the youth side and the older side, while apparently opposite —

protection from labor on one side, inclusion into labor on the other — can be read as

responses to the same structure. What the three-tier structure of ILO No. 138 and the minor

provisions of the Labor Standards Act protect is the condition that opportunities for

education and development not be eroded by labor under an asymmetry of bargaining

power. What exploitation prevention on the older side protects is the condition that the

compensation and cognitive load of those in positions with high exit costs be kept adequate.

Proposition 10 (the symmetry of Brain Safety) integrates the two as symmetric design

principles (priority of education and health, ceilings on load, adequacy of compensation)

responding to the same structure of bargaining-power asymmetry and exit costs. This

section's institutional comparison has shown that one side of this symmetry (youth) has

reached nearly universal consensus in international law, while the other side (older, nonemployment)

is still in formation in each country.

In sum, the implementability of Ageless Management is a function of institutional trends.

The more institutions shrink "working means losing" and expand protection for nonemployment

forms, the lower the implementation costs of Definition 1's dynamic role

allocation and Definition 7's multigenerational ecosystem. Conversely, the longer the

protection vacuum is left unattended, the more implementation deepens its dependence on

the organization's autonomous norms, and implementation without norms raises the risk of

Working Paper | Ageless Management in the AI Era 121

exploitation. That this paper positions Brain Safety in Section 8 not as a merely desirable

consideration but as a condition for the validity of Ageless Management (the conditional

clauses of Propositions 8 and 11) is a consequence of this institutional reality.

Evidence-grade note: The institutional descriptions in this section are limited to matters checked against

primary statutes and administrative sources (e-Gov, the Ministry of Health, Labour and Welfare, the

Japan Pension Service, the SSA, the U.S. Congress, the EEOC, the DRV, EUR-Lex, and the ILO) and against

public research institutes (JILPT) and professional commentary. Singapore's 2026 amendment rests on

multiple cross-checks of law-firm and HR-media commentary; final confirmation against primary

sources of the responsible ministry (MOM) remains outstanding. This section is a description of

institutions, not legal advice.

8. Brain Safety: Occupational Health for the Brain

Definition 8 defined Brain Safety as occupational health and safety standards that protect the

brain capital of ecosystem participants from depreciation. It is bidirectional: (i) protection

from cognitive load and brain fatigue, and (ii) protection from conversion into unpaid labor

and from exploitation arising from asymmetries of bargaining power. As the preceding

section showed, the non-employment forms of participation on which Ageless Management

depends fall into the protection gap of current institutions (Proposition 9). Until institutions

catch up — and after institutions are in place as well — the designers of an ecosystem must

have health and safety standards of their own. This section presents the skeleton of those

standards. Section 8.1 addresses the cognitive-load side and Section 8.2 the anti-exploitation

side; Section 8.3 shows the theoretical connection to the framework of this series' BCM

(Kadowaki 2026e), and Section 8.4 consolidates the results into Table 9 (the skeleton of the

guidelines).

8.1 The Cognitive-Load Side — The Individual Version of the Solvency Condition

and Dual-Track Auditing

Conventional occupational health and safety has been designed around physical load and

hazardous work. The audit-type roles that Ageless Management assigns as the default (Section

6) reduce physical load while shifting the center of gravity of the load to the cognitive side.

Verifying AI output continuously consumes working memory and attentional resources in the

form of sustained attention, document reading, and responses to alarms. In light of the

structure of cognitive aging confirmed in Section 2, this load structure constitutes, for superseniors

whose age-related changes lie on the side of processing speed and attention, a health

and safety problem of a different kind from physical load.

This paper formalizes this problem as the individual version of the solvency condition of

WP8 (Kadowaki 2026h). WP8's solvency condition states that there is an upper bound on the

cognitive resources, time, and cost an organization can devote to oversight, and that oversight

can be sustained only within that bound. The same logic operates inside the individual. An

Working Paper | Ageless Management in the AI Era 122

individual's attentional resources and cognitive effort are also subject to a budget constraint,

and auditing can be sustained only within that budget. What happens when auditing is piled

up beyond the budget has also already been described by WP8: alarm fatigue. The more

outputs there are to check, the more sluggish the auditor's responses become, and the

detection probability falls. That is, neglected cognitive load is at once a health and safety

problem that depreciates the auditor's own brain capital and a governance problem that

undermines the detection probability that is the source of oversight value. An overloaded

audit regime damages the person and fails to protect what it is meant to protect. It fails twice.

From this formalization, the three design principles listed in the original draft can be rederived.

The first is control of working time. Dividing an audit session into short units

(roughly 20–30 minutes as a guide) with rest in between is a budget-execution discipline for

keeping the rate of consumption of the attention budget within its ceiling. The commissioning

unit for audit tasks is likewise designed per session, not as continuous engagement. The

second is a multimodal work environment. Auditing that depends on screen gaze and reading

concentrates its consumption on visual attention. The combined use of voice interfaces (Voice

UI) and read-aloud is a candidate design for dispersing this concentration of consumption.

The third is interactive auditing. A format that progressively narrows the object of

verification through AI summaries and question-and-answer, rather than the bulk reading of

long documents, limits the amount of information that must be held at once. What must be

stated explicitly, however, is that all three principles are design hypotheses derived from the

attention-budget formalization, and their effects on audit quality and brain health have not

been tested. That testing belongs to the measurement framework of Section 9.

Managing the budget constraint should not be left to the individual auditor's self-restraint.

The parties who commission audits usually do not know how much of an auditor's attention

budget each request consumes, and if multiple commissioners pile requests onto the same

auditor, the total exceeds the budget even when each individual request is reasonable. The

cognitive-load management of Brain Safety must therefore be designed as a pair: session

design (execution discipline on the individual side) and total-volume management of audit

demand (commissioning discipline on the organizational side). Concretely, the demand-side

management tools include visualizing each auditor's acceptance ceiling, prioritizing requests

(not having every output audited equally), and AI-side design that reduces the volume of

output requiring audit in the first place (such as flagging low-confidence outputs). To

acknowledge the scarcity of oversight is the same thing as to use oversight with care. The

effects of this prioritization and pre-screening on the independence of the audit channel,

however, are addressed explicitly at the end of this section.

The rarity effect that WP8 identified can also be reread at the individual level. The rarer

errors are in an environment, the harder it is to sustain a supervisor's vigilance. As AI

performance rises, errors become rarer, auditing becomes the task of confirming "almost

always fine," and attention adapts to that environment and declines. The experimental

finding seen in Section 4.2 — that human vigilance slackens most under imperfect but high-

Working Paper | Ageless Management in the AI Era 123

performing AI — is consistent with this mechanism. The design implication is not to entrust

the maintenance of vigilance to the auditor's will. Candidate devices include exercises that

deliberately mix known errors into the audit stream to confirm detection (an operational

version of the "embedded errors" adopted by the experimental task of Hypothesis H3),

rotation of audit targets to prevent habituation, and the sharing of detection cases to update

the fact that "errors are real." These, too, are design hypotheses, and testing their

effectiveness belongs to Section 9.

At the same time, load ceilings must not be made uniform by age. Age-based protection

such as "halving audit time across the board because someone is old," however well

intentioned, is an operation that reintroduces, in the name of protection, the chronologicalage

variable that Definition 1 removed — the same slide as in dynamic role allocation

(Section 6.2): protective stereotyping. As confirmed in Section 2, age-related changes in

attention and processing speed show large individual differences and overlapping

distributions, and load tolerance is no exception. The load management of Brain Safety

should be designed not as uniform age-specific standards but as individualized ceilings based

on measured states (fatigue, performance degradation, and the person's own reports). The

individualization of protection is the consistent application of the principle of Ageless

Management to occupational health.

How, then, is the state of load to be measured? There are three candidate layers. First, selfreport

(periodic short scales of fatigue and perceived load). Second, behavioral indicators

(degradation signals derived from work logs, such as changes in audit response time, declines

in detection rate, and increased misses in the latter half of a session). Third, participation

patterns (the frequency of session interruptions and declinations). Behavioral indicators may

be more sensitive than self-report, but continuous collection is one step away from turning

auditors into objects of surveillance, and the design requirements are the separation of the

purpose of measurement (protection) from its use (evaluation and selection), feedback to the

person, and consent to the scope of collection. The danger that measurement is diverted from

protection to selection is taken up again as self-critique in Section 10.

Operational restriction of unannounced verification — anonymized system

calibration. The danger of surveillance appears most sharply in the unannounced, real-time,

blind in-situ verification (unannounced in-situ verification) that Proposition 12 specifies as an

operational requirement of dynamic role allocation. This section fixes, from the Brain Safety

side, the operational principle stated in Section 5 as a restrictive clause. Unannounced in-situ

verification is operated, in the first instance, in anonymized and aggregated form for systemlevel

calibration — testing the divergence between backtest results and real-time observation.

Its use for individual treatment (such as changes in role allocation) is permitted only through

aggregation over long windows and due process including disclosure to the person and an

opportunity to respond; demotion or revocation of authority on the basis of a single

unannounced result is not performed. This restriction is not merely the general fairness

requirement of restraining arbitrariness; it is an intrinsic requirement of Brain Safety. Under

Working Paper | Ageless Management in the AI Era 124

a regime in which one cannot know when one is being observed and a single observation can

directly determine one's treatment, participants are forced to behave under the assumption

of constant surveillance — the panopticon effect. What the sense of constant surveillance

supplies is not discipline but a permanent state of vigilance, which continues to occupy

cognitive bandwidth as an unerasable claim on attentional resources. That is, individually

surveillant unannounced verification erodes, in the name of governance, the very thing this

section seeks to protect as the extension of Base — cognitive bandwidth — and contradicts the

purpose of Brain Safety head-on. The deterrence of gaming is achieved by publicizing the fact

that the system is calibrated, and does not require diverting individual observations to

individual surveillance (Section 5, commentary on Proposition 12). This two-layer separation

— anonymized system calibration, and use for treatment only over long windows with due

process — is included in Table 9 as the only operational mode that reconciles Proposition 12's

anti-gaming safeguard with Brain Safety.

Consent-based load adjustment. How, then, should protection be activated when a

measured state crosses the protective threshold? This paper does not adopt unilateral forced

cutoff by the firm — immediate exclusion from audit work or unilateral termination of the

contract on the grounds of fatigue or degradation indicators — as the mode in which Brain

Safety is activated. There is a conflict here that must be made explicit. Unilateral cutoff is an

operation that overrides, in the name of protection, the person's will to work and right of selfdetermination,

and it can slide into a variant of ageism that merely replaces the

chronological-age variable removed by Definition 1 with fatigue indicators — even when

"because you are old" becomes "because the data say so," it remains exclusion that bypasses

the person's will. On the other hand, leaving degradation signals unattended damages both

the person's brain capital and the detection probability of the audit (earlier in this section).

This paper's design for this conflict is risk alerts and consent-based dynamic contract

renewal. That is, objective data — degradation signals derived from work logs and measured

values of fatigue — are used first as the presentation of a risk alert to the person, and load

adjustments — reduction of session counts, changes of format, pause and resumption — are

implemented as renewals of contract terms under the person's consent (informed consent),

accompanied by an explanation of the meaning of the measurements and of the options, and,

where necessary, with the involvement of an intermediary organization (Sections 8.2 and

6.4). The involvement of the intermediary is a bulwark against "consent" turning into de facto

coercion under asymmetric bargaining power. Even this mode, however, leaves a residual

tension. If the person does not consent to adjustment while degradation progresses, the

collision between responsibility for audit quality and the person's self-determination does not

disappear. And this situation is not exceptional but the normal case the design must

anticipate. Underestimation of one's own functional decline — anosognosia (reduced illness

awareness) and inflated self-assessment — can itself be a degradation of metacognitive

calibration, so consent-based adjustment carries a structural blind spot: consent is hardest to

obtain precisely in the situations that most require adjustment. Choosing neglect here is a

Working Paper | Ageless Management in the AI Era 125

collapse of governance (damage to both audit quality and the person's brain capital);

choosing forced cutoff is a relapse into the violation of self-determination. Against this

deadlock, this paper specifies a third path as an escalation procedure: an objective second

opinion by a third-party body independent of the interests involved — an occupational

physician, an external intermediary organization (Sections 6.4 and 8.2), or the like. The third

party evaluates the validity of the measurement data and the reasonableness of the proposed

adjustment independently of both the person and the commissioning organization, and its

findings are used both for re-presentation to the person and for revision of the adjustment

proposal. Only if agreement still fails to form after this procedure are organizational

measures justified, in an order that puts reassignment according to the risk level of the

audited domain (transfer to low-risk domains) first and termination of participation as the

last resort. That no general solution to the collision exists remains true. But specifying the

escalation path and the order of measures in advance is a bulwark against both arbitrary

cutoff and irresponsible neglect.

The cognitive-load side has one further aspect. As seen in Section 4.3, dependence on

assistance can depreciate skills that are not exercised (Budzyล„ et al. 2025). If the reduction of

audit load is pursued to the point of full reliance on AI pre-screening and pre-summarization,

the auditor's own detection skill depreciates and the source of oversight value is lost.

Restraining cognitive load and maintaining skill therefore involve a trade-off, and load design

must incorporate deliberate opportunities for exercise with AI assistance removed — what

BCM calls protected unassisted practice (Section 8.3). Designing the ceiling on load and the

floor on exercise simultaneously is the requirement of the cognitive-load side of Brain Safety.

There is another, independent channel of damage to be placed alongside skill depreciation:

damage to the independence of the audit channel. The load-reduction designs listed in this

section — AI summaries, interactive auditing, confidence-based prioritization — lower

cognitive load while at the same time making the information channel of the act of auditing

dependent on AI. As Proposition 5 makes explicit, generational decorrelation widens the

group's detection set under the condition that auditors access the original output in

unmediated form; if all auditors share summaries and pre-screening by the same AI,

generational heterogeneity is homogenized at the entrance of the audit, and misses recorrelate

at the channel level (Section 6.3). That is, pushing load reduction to the limit not

only depreciates the individual auditor's detection skill (previous paragraph) but can also

damage the group's independence — the other factor of oversight value. Load reduction and

independence stand in a trade-off, and this trade-off must be managed explicitly within the

design.

Dual-track auditing. Neither pole of this trade-off is viable as a design. Full-volume

unmediated auditing exceeds the ceiling on the attentional resources that can be devoted to

oversight — the solvency condition of the organization and of the individual — and collides

head-on with the load-management requirements of the first half of this section. Full-volume

AI-summarized auditing lowers the load but breaks the channel condition of Proposition 5

Working Paper | Ageless Management in the AI Era 126

and nullifies generational decorrelation at the entrance of the audit. This paper's design

therefore takes a dual structure. As the first track, for a portion sampled at random from all

outputs, a raw-audit slot is preserved in which the raw output is audited without summaries

or pre-screening (unmediated, unsummarized, unscreened). As the second track, the

remaining outputs receive screening by AI summaries, interactive auditing, and

prioritization, economizing on the attention budget. That the sampling is random is essential:

raw auditing of the "seemingly important parts" selected by an AI reintroduces channel

sharing through the selection itself, and equally excludes from the first track precisely the

places where the AI errs with confidence (Sections 6.3 and 5).

Sampling design — the breakdown of fixed rates and risk-weighted dynamic sampling.

How the sampling rate of the first track is designed determines the success or failure of the

dual track. First, fixed-rate random sampling necessarily breaks down as processing volume

grows. The capacity arithmetic is simple. The attention required for raw auditing grows in

proportion to (sampling rate × total output volume), whereas the attention that can be

supplied is capped at (number of auditors × per-person attention budget). As long as AI

throughput keeps growing, any fixed sampling rate will eventually exceed this supply ceiling

— a fixed rate collides head-on with the solvency condition. Conversely, if the sampling rate is

lowered continually to stay within the supply ceiling, coverage declines in inverse proportion

to total volume. The limit on oversight coverage that WP8 identified is not dissolved by the

dual track; it is merely reallocated in the form of sampling.

For that reason, the allocation of the scarce raw-audit slots should not be indiscriminate.

This paper's design is risk-weighted dynamic sampling. Outputs are stratified by impact,

based on the severity of the consequences of error — irreversibility (the possibility of ex-post

correction), scope of propagation (the number and kinds of parties affected), and monetary

scale — with the sampling rate raised for high-impact strata and relatively lowered for lowimpact

strata. Sampling within each stratum remains random. Two disciplines attach to this

weighting. First, impact weights must be defined from the attributes on the consequence side

of the output, independently of the AI's self-assessment. In particular, the AI's confidence

must not be used as grounds for lowering the sampling rate as "evidence of low risk." Errors

accompanied by high confidence are precisely the typical manifestation of hallucination, and

confidence-linked pre-screening systematically removes from the audit the places where the

AI errs with confidence — exactly the configuration against which the findings on automation

bias and alarm fatigue in Section 4.2 and WP8 warn. Using low confidence as a trigger for

additional checking (earlier in this section) is permissible; using high confidence as grounds

for exemption from audit is not — the discipline is asymmetric. Second, a purely random

floor that obeys no impact weighting whatsoever (a minimum sampling rate applied across

the entire output domain) is maintained. If all outputs were selected according to the impact

function, the selection function itself would become a new channel shared by all auditors,

and errors in the domains left outside the impact assessment would reach no one's eyes. The

role of the floor is twofold. Its first role is the insurance against fallibility just mentioned — a

Working Paper | Ageless Management in the AI Era 127

safeguard against errors and blind spots in the stratification design itself. Its second role is an

exploration slot for unknown unknowns. The impact weighting by irreversibility, scope of

propagation, and monetary scale is constructed in reliance on knowledge of what brings

about grave consequences — that is, on the lessons of past accidents and failures. Novel

modes of breakdown outside that weighting — errors classified as low-impact precisely

because they resemble none of the past accident types, or errors that do not appear in the

variables of the impact function at all — therefore, by definition, never ride the risk-weighted

track. Only a purely random floor that relies on no existing risk knowledge can, in principle,

catch errors of this kind — exploratory sampling beyond the distribution (exploration) as

against sampling optimized to the known distribution (exploitation). In this twofold role, the

floor is the last safeguard of the channel condition of Proposition 5 — the preservation of

decorrelation.

Forced queue. Sampling design must also include control of the later stage of sampling —

the phase in which the sampled outputs are actually read. Even if sampling is perfectly

random, if auditors can themselves choose what to read from among their assigned outputs, a

second selection by preference arises inside the track. An auditor under an attention-budget

constraint digests short, plain, readable outputs first and defers long and complex ones — and

the deferred outputs are then often left unaudited when the session ends. This auditor-side

selection bias (cherry-picking) is a natural budget-saving behavior requiring no ill will, yet its

consequences are systematic: long and complex outputs are, in many cases, high-risk outputs

with intricate judgments, and selection by preference erodes the coverage that random

sampling was supposed to guarantee precisely from the high-risk side. The countermeasure is

algorithmic delivery control. Assignments within the raw-audit track are delivered as a

forced queue in random order that the auditor cannot choose, and the auditor processes

them in the order received. In addition, the volume of each delivery is calibrated to what can

be digested within the session time ceiling (the 20–30 minutes earlier in this section) — for if

the volume exceeds the budget, truncation of the tail revives selection even when the order is

forced. Audit difficulty arising from low readability is handled not by deferral but by splitting

the session and assigning an additional session — since this is the raw-audit slot, lightening

by summarization would violate the principle of non-mediation. The forced queue is the

enforcement mechanism that insulates the design decision of random sampling from the

individual auditor's preferences.

Even under this refinement, the residue of Type II error — errors slipping through — does

not disappear. It can be disclosed quantitatively. If the raw-audit sampling rate of a given

stratum is p and the miss rate of AI-summarized screening is β, the probability that an error

in that stratum is captured by neither track is approximately (1−p)×β, which is positive as

long as p is below 1 and β is positive. Risk weighting lowers the slip-through probability of

high-impact errors, but does not bring it to zero, and slip-through in the domains classified as

low-impact is, indeed, tolerated by design. Accordingly, even after the introduction of risk

weighting, it remains the case that an operation claiming "we are safe because a fixed rate of

Working Paper | Ageless Management in the AI Era 128

raw auditing is secured" supplies only reassurance detached from the reality of coverage —

no fixed guarantee of safety exists. Ex-ante disclosure of the sampling rates (including

stratum weights and the floor) with periodic revision in response to changes in processing

volume and risk, disclosure of estimated slip-through probabilities, and measurement of

detection performance by channel (raw audit vs. AI summary) (Section 9) are the minimum

requirements for keeping this design variable open to verification.

Dual-track auditing is the independence version of the protected unassisted practice of BCM

(Kadowaki 2026e). Just as unassisted practice is the floor on exercise that prevents the

depreciation of an individual's skill, the raw-audit slot is, in the same form, the floor on

independence that prevents the erosion of the group's decorrelation. Load reduction, skill

maintenance, and independence maintenance cannot all be maximized at once, and the

cognitive-load design of Brain Safety is operated as a constrained design problem that

simultaneously specifies the ceiling on load, the floor on exercise, and the floor on nonmediation

(the raw-audit slot including the random-sampling floor).

Layer separation of audit feedback. The errors and corrections detected by audits are fed

back indirectly into the improvement of AI prompts, evaluation criteria, and verification

procedures (Section 6.2). To the governance of this feedback path, this paper adds one further

requirement: layer separation — human corrections and their interpretive context are not

fed in full volume into the retraining of the AI's base model. The ground for concern lies in

the demonstration of model collapse: when a generative model is trained continually on

recursively generated data, the tails of the original data distribution are lost and the model

loses diversity (Shumailov et al. 2024). Since the object of audit feedback is AI-generated

output and its corrections, an operation that recycles this into retraining without selection

has the structure of recursively biasing the training data toward self-generated artifacts and

their vicinity. What the demonstration directly addressed was recursive training on modelgenerated

data, and extrapolation to feedback containing human corrections carries a margin

of uncertainty; but if degradation of generative diversity does occur, it works in the direction

of deepening from within the systematic blind spots of a single model against which model

condition (iii) of Proposition 5 warns. The design requirement is therefore a separation of

layers: the interpretive context of the audit — context metadata such as the reasons for

corrections, contextual information, and conditions of application — is managed at the layer

of prompts, retrieval references, and evaluation criteria, and input into the training of the

base model is limited to selected and recorded portions. Layer separation alone, however, is

not enough. The retrieval-reference (RAG) layer that manages interpretive context is itself an

accumulating database; if it saturates with the past judgments of a small number of auditors,

then even with the base model protected, the context the retrieval returns converges on

particular auditors and judgment types, and every subsequent audit and generation refers to

it — a loss of diversity of the same form as model collapse is reproduced at the RAG layer. This

paper therefore places, in addition to layer separation, two operational requirements on the

RAG layer: (i) temporal decay of logs — attenuating the reference weight of old judgments

Working Paper | Ageless Management in the AI Era 129

over time to prevent the entrenchment of past judgments — and (ii) source-diversity

weighting — imposing diversity constraints on the composition of references so that retrieval

results are not biased toward particular auditors, generations, or judgment types. These

requirements are included in Table 9.

8.2 The Anti-Exploitation Side — The Symmetry of Proposition 10

The second aspect of Brain Safety is economic protection. The participation forms of Ageless

Management have their center of gravity in non-employment types — outsourcing contracts,

advisory roles, PBL, and NPO partnerships. As Proposition 9 pointed out, these lie outside the

employment relationship that is the principal unit of labor law's protection and restraint, and

an ecosystem left unattended can degenerate into exploitation. There are three typical modes

of degeneration. The first is conversion into unpaid advisors. The practice of procuring the

experiential audits of super-seniors without pay or for nominal consideration, under

honorific titles such as "advisor" or "counselor," is nonpayment for the scarce resource of

experiential audit capacity (Sections 4–5). So-called exploitation through purpose (yarigai

exploitation) — using the logic that "social participation is itself the reward" or "it gives you

purpose in life" as a substitute for compensation — makes the intrinsic motivation of

participation a cover for asymmetric bargaining power. The second is the conversion of youth

PBL participation into unpaid labor in the name of education. The moment the commercial

acquisition of deliverables becomes the primary purpose and education becomes

subordinate, it is no longer PBL but disguised labor. The third is the consumption of the

"insider perspective" of socially marginalized groups as opinion-gathering unaccompanied by

compensation or attribution.

Proposition 10 claims that these apparently different problems share a single structure.

Labor protection for youth and the anti-exploitation and cognitive-load protection of superseniors

are both responses to the same structure — asymmetric bargaining power and exit

costs — and demand symmetric design principles. Youth are prone to accepting unfavorable

terms because of asymmetries of experience, information, and legal status; super-seniors

because of the scarcity of re-employment opportunities and the desire for social recognition;

marginalized groups because of the scarcity of participation opportunities themselves. What

the symmetry claim means in practice is that protection can be designed not as a bundle of

stratum-specific exceptions but as the application of a single design principle — priority of

education and health, ceilings on load, and the adequacy of compensation — to all

participating strata. Concretely: (i) require written contracts and explicit compensation for all

forms of participation, and set compensation for the provision of experiential audit and

insider perception with reference to market levels; (ii) require the primacy of educational

purpose for youth participation, institutionalizing the priority of schooling, restrictions on the

commercial use of deliverables, and the involvement of educational institutions (Section 7);

(iii) apply load ceilings (Section 8.1) to all strata; (iv) contractually provide non-employment

participants with accident compensation and consultation channels comparable to those of

Working Paper | Ageless Management in the AI Era 130

employees. Note that, as the refutation condition of Proposition 10 specifies, if the protection

requirements of the two strata are shown to differ structurally, this symmetric design turns

out to be an oversimplification. Symmetry is a hypothesis that simplifies design, not a

doctrine.

The primacy of educational purpose in youth participation needs operational criteria for

judgment. This paper proposes three. First, documentation of learning objectives. For each

PBL project, the learning objectives of the participating youth are defined in advance, and

their attainment is assessed by the educational institution. Second, subordination of

deliverable use. The firm's acquisition of deliverables is positioned as a byproduct of the

learning process, and the commercial value of deliverables does not govern project design.

The moment the initiative of design shifts to "what should we have them make so that it

sells," the primacy of educational purpose is lost. Third, educational design of AI use. As the

evidence in Section 4.3 shows, AI use without guardrails can raise assisted performance while

impairing learning. "Youth participation that delivers results fast with AI," convenient for the

firm, can be the most dangerous design from the standpoint of education. A connection that

does not explicitly manage this tension does not satisfy the requirement of the primacy of

educational purpose.

For the participation of socially marginalized groups as well, the application of the

symmetric principle should be made concrete. Insider problem perception (Table 5) is a

cognitive asset for the ecosystem and should not be procured as unpaid "cooperation with

interviews." There are three design requirements. First, compensation. If the provision of

insider perception is used to improve products or businesses, it is professional service and

the object of commensurate compensation. Second, attribution. Contributions to ideas and

improvements originating in insiders' perception should be recorded, and attribution to the

person made explicit. Third, consideration for load. Repeatedly narrating the experience of

one's own hardship itself carries cognitive and emotional load. Placing the choice of

frequency and format of participation on the person's side is the insider version of the load

management of Section 8.1. "Insider participation" lacking these is mere extraction wearing

the appearance of inclusion.

The adequacy of compensation faces one practical difficulty: for the task of experiential

audit, no established labor market or reference price yet exists. The absence of a market price

lends itself easily to the pretext for markdown that "there is no going rate, so an honorarium

will do." For the time being, the design can only hold, as self-imposed standards, the

application by analogy of the compensation levels of adjacent professional services (audit,

advisory, expert committee membership, and the like), ex-ante disclosure of per-task

compensation, and periodic review of compensation levels. Note that if Proposition 1 (the

marginal value shift) is correct, referenceable market prices should form as demand for

evaluation- and audit-type tasks increases, and whether that formation occurs is itself one of

the observable implications of Proposition 1.

Working Paper | Ageless Management in the AI Era 131

Governance of the over-flagging rate — suspension and recalibration of audit

authority. The design of the anti-exploitation side needs, as its pair, a bulwark against the

failure mode running in the opposite direction. As the compensation and status of audit-type

roles become established, an incentive can arise to pile up the volume of flags as proof of

contribution. If the response criterion drifts toward "suspect everything," false alarms inflate

verification costs and alarm fatigue, and organizational decision-making heads toward

gridlock through over-auditing. Moreover, if collusion arises in which multiple senior

auditors endorse one another's flags — logrolling — over-auditing becomes fixed as a group

equilibrium beyond individual tendency. The division of labor that does not assign this

control to ex-post damages claims was set out in Section 7.2: liability doctrine handles the

allocation of catastrophic risk, and everyday discipline is handled by operational governance.

That is, the standard operation incorporates a procedure of continuously monitoring each

auditor's false-alarm rate — the estimated bias of the response criterion c in signal detection

theory (Section 9; Appendix A) — temporarily suspending the audit authority of an auditor

who exceeds a pre-set and published threshold, and reinstating that auditor upon passing a

recalibration task in which known errors and sound outputs are mixed in under blinding.

Basing the judgment on performance in blinded tasks rather than on mutual evaluation

among auditors is a bulwark against collusive mutual endorsement contaminating the

calibration judgment. Likewise, ex-ante disclosure of thresholds and procedures, and basing

the judgment solely on calibration indicators rather than on the content of the flags, are

bulwarks against suspension being diverted into the suppression of inconvenient flags — the

erosion of the voice channel that organizational condition (ii) of Proposition 5 protects. In

addition, a reputation mechanism based on the recorded history of detections, false alarms,

and recalibrations supplies everyday discipline against faulty auditing without recourse to

damages claims that are difficult to prove (Section 7.2).

A word is also in order on who implements protection. Individual participants, above all

non-employment individuals, cannot be expected to negotiate load ceilings or the adequacy of

compensation with firms — precisely because of asymmetric bargaining power. Here

intermediary organizations such as NPOs, educational institutions, and employment agencies

(Section 6.4) can perform the bargaining-power-correcting functions of collective negotiation,

the development of standard contracts, and the provision of grievance channels. Correction

by private ordering, however, is no substitute for institutional protection. The protection gap

that Proposition 9 identifies should ultimately be filled by institutional design (such as the

extension of freelance-protection legislation seen in Section 7), and the guidelines of this

section are positioned as the organization-side self-imposed standards to hold until

institutions catch up, and to lay on top of them.

Working Paper | Ageless Management in the AI Era 132

8.3 Connection to BCM's Three Bs — Brain Safety Is the Lifelong Extension of

Base

The BCM of this series (Kadowaki 2026e) formalized a firm's brain capital as stock (K) ×

utilization rate (u), and organized the sequential conditions for individual cognitive ability to

become firm capability into three constraints — Belonging, Base, and Build: that the ability

exists and is maintained (failure mode: decay), that it is available at the moment of work

(failure mode: depletion), and that it does not remain inside the head but is transmitted to the

organization (failure mode: silence). The core discipline of BCM lay in the distinction between

states and measures — the point that what links to value is not expenditure on measures but

the measured state of the workforce.

Brain Safety has a precise location within this framework: Brain Safety is the lifelong

extension of Base. BCM's Base formalized, for the workforce within employment, the claim

that resolving outstanding economic claims on attention (economic hardship and liquidity

constraints) releases cognitive bandwidth. This paper extends this Base in two directions. The

first is the extension of scope. BCM's unit was the employed workforce, but the participants of

a multigenerational ecosystem extend beyond the boundary of employment. The economic

foundation of non-employment participants (adequate compensation and the provision of

accident coverage, Section 8.2) is nothing other than Base extended beyond employment. The

second is the extension of the sources of claims. Claims on attention do not come from

economic hardship alone. The cognitive load, alarms, and brain fatigue emitted by the audittype

role itself are a standing claim on attentional resources, and the load design of Section

8.1 is an intervention that keeps this claim within budget. In Ageless Management, which

presupposes lifelong participation, the protection of Base extended in these two directions is

a precondition for restraining the depreciation of participants' brain-capital stock (K) and

maintaining the utilization rate (u) (Proposition 8).

The phrase "lifelong extension" also carries a temporal implication. BCM's Base addressed

the current-period cognitive bandwidth of the workforce during employment. Participation

in Ageless Management is not a single period of employment but a decades-long process of

repeated increases and decreases in participation volume, interruptions, and re-entries. On

this time axis, the protection of Base is not limited to load management during participation

but includes the maintenance of brain capital in transition periods — the shrinking of roles,

temporary exit, and preparation for re-entry. The evidence since Mental Retirement, seen in

Section 3, suggests that the sharp drop in cognitive engagement accompanying exit is the

principal risk of this transition period, and the gradual adjustment of participation volume

that an ecosystem can offer (gradual decrease and increase of session counts) is a design

alternative to the "all-or-nothing" exit of the employment system. Whether this gradual

adjustment actually restrains the depreciation of K is, however, precisely the object of testing

of Hypothesis H1.

Working Paper | Ageless Management in the AI Era 133

The extension does not stop at Base. In the dimension of Belonging, experiential audit

capacity becomes oversight value only when a detected error reaches others as voice and is

heard. In WP8's terms, detection leads to the correction of error only through reporting and

acceptance. In an organization where an auditor's dissent is discounted on grounds of status,

age, or contract form, assembling the decorrelation portfolio (Section 6.3) realizes no value.

The failure mode of silence is most likely to arise precisely among participants outside the

employment boundary who are in weak positions. In the dimension of Build, as stated at the

end of Section 8.1, the design of unassisted exercise opportunities prevents the depreciation

of experiential audit capacity. Applying the protected unassisted practice that BCM proposed

not only to youth in the course of learning but to the Gc-type roles of super-seniors is an

extension newly claimed by this paper, and its effectiveness is untested (Section 9). BCM's

distinction between states and measures likewise applies to Brain Safety as it stands: Brain

Safety is not to be deemed achieved by the introduction of a list of measures, but is to be

evaluated by the measured states of participants' load, fatigue, and engagement.

Finally, the relation between Brain Safety and the brain-health hypothesis (H1) should be

made clear. As seen in Section 3, the evidence on the brain-health effects of work is mixed,

and Proposition 7 qualified the claim: if an effect exists, it is mediated not by the length of

working hours but by the maintenance of cognitive engagement. Brain Safety is the designside

counterpart of this qualification. That is, this paper does not claim that "participation is

good for the brain." Overloaded auditing, chronic alarm fatigue, and participation under

exploitative terms supply not cognitive engagement but chronic stress and attrition, and can

instead depreciate brain capital. What Hypothesis H1 takes as its object of testing is

engagement in roles that are Gc-exercising and AI-complemented, and its formalization

implicitly includes the conditions of this section — load managed within budget and the

absence of exploitation. The expansion of participation lacking Brain Safety is a path that

worsens participants' brain capital under the banner of Ageless Management, and

Proposition 8 promises nothing about the consequences of such a design. What the refutation

condition of Proposition 8 asks is whether, under a design judged to satisfy the conditions

prior to the observation of outcomes, a deterioration of brain-capital indicators is observed

ex post (Section 5).

8.4 Table 9: The Skeleton of the Brain Safety Guidelines

The foregoing discussion is consolidated into the guideline skeleton of Table 9. Each row

shows the domain, the design principle, examples of concrete measures, and the

corresponding propositions of this paper.

Table 9 The skeleton of the Brain Safety guidelines

Domain Principle Concrete measures (examples)

Corresponding

propositions

Do not exceed the

budget constraint on the

Time ceiling per audit session

(roughly 20–30 minutes as a guide)

Propositions 7 and

8 (WP8's solvency

Working Paper | Ageless Management in the AI Era 134

Cognitive-load

management

individual's attentional

resources (the

individual version of the

solvency condition)

with rest; per-session

commissioning of tasks;

management of the volume of items

to check and of alarms

condition and

alarm fatigue)

Consent-based

load adjustment

(informed

consent)

Protection is activated

not by unilateral forced

cutoff but by consentbased

dynamic contract

renewal

Presentation of risk alerts based on

objective data; renewal of contract

terms involving the person and an

intermediary organization; an

objective second-opinion procedure

by a third-party body (occupational

physician, external intermediary

organization) when agreement fails

to form; an order of measures that

puts reassignment to low-risk

domains first and termination of

participation as the last resort

Propositions 7, 9,

and 10 (managing

the conflict with

self-determination)

Workenvironment

design

Distribute the modes of

load and make work

cognitively accessible

Combined use of Voice UI and readaloud;

interactive auditing through

AI summaries and question-andanswer;

reduction of the cognitive

load of displays and documents

Propositions 2 and

11

Skill

maintenance

Design the ceiling on

load and the floor on

exercise simultaneously

Regular incorporation of protected

unassisted practice; avoidance of

full reliance on AI pre-screening;

assessment of ability under both

assisted and unassisted conditions

Propositions 3 and

8 (extension of

BCM's Build)

Sampling design

for the raw-audit

slot (dual-track

auditing)

Risk-weighted dynamic

sampling — stratify by

impact while

maintaining a purely

random floor (minimum

sampling rate)

(independence of the

audit channel; the floor

is insurance against

fallibility and an

exploration slot for

unknown unknowns)

Design of stratum-specific sampling

rates by impact (irreversibility,

scope of propagation, monetary

scale) with random sampling within

strata; the discipline of not using

the AI's confidence as grounds for

lowering sampling rates;

application of the random-sampling

floor to the entire output domain

(exploratory sampling for novel

modes of breakdown outside the

impact weighting); ex-ante

disclosure of sampling rates and the

floor and disclosure of the slipthrough

probability ((1−p)×β);

measurement of detection

performance by channel (Section 9)

Proposition 5

(WP8's

decorrelation

condition and

solvency condition;

the independence

version of protected

unassisted practice)

Forced queue Insulate assignments

within the raw-audit

track from auditors'

preferences (prevention

of auditor-side selection

bias)

Delivery in a random forced order

that cannot be chosen; calibration

of volume to what can be digested

within the session time ceiling (20–

30 minutes); monitoring of deferral

and tail truncation; handling of

Proposition 5

(enforcement

mechanism of

random sampling)

Working Paper | Ageless Management in the AI Era 135

hard-to-read outputs by session

splitting

Independent

double-checking

(exclusion of

solo auditing)

Do not make the audit

dependent on the

judgment of a single

auditor (a bulwark

against cognitive holdup)

Independent audits by multiple

heterogeneous seniors with crosschecking

of flags; exclusion of solo

auditing by a single senior; plural

evaluation of the validity of flags

(Section 6.3)

Proposition 5

(application of

Section 6.3)

Periodic blinded

evaluation of

audit validity

Base the evaluation of

auditors on the validity

of flags, not their

volume

Blinded mixing-in of known errors

and unproblematic outputs;

periodic measurement of the

detection rate and the over-flagging

rate (operational version of H3

dependent variable (iv)); feedback

of results to the person

Propositions 3, 4,

and 12 (backtesting

on real work logs)

Over-flagging

governance

(suspension and

recalibration of

audit authority)

Control the moral

hazard of over-flagging

by operational

governance, not by

damages claims

(division of labor with

the cap on liability of

Section 7.2)

Continuous monitoring of each

auditor's false-alarm rate

(estimated bias of the response

criterion c); temporary suspension

of audit authority upon exceeding

pre-set and published thresholds;

recalibration by blinded tasks with

reinstatement judgment; a

reputation mechanism based on the

history of detections, false alarms,

and recalibrations; non-reliance on

mutual evaluation as a bulwark

against collusion (logrolling)

Propositions 5 and

12 (monitoring of

the SDT response

criterion c —

Section 9; Appendix

A)

Operational

restriction of

unannounced

verification

(anonymized

system

calibration)

Use unannounced insitu

verification, in the

first instance, for

anonymized and

aggregated system-level

calibration, and do not

divert it to individual

surveillance (exclusion

of the panopticon effect

— consistency with the

protection of cognitive

bandwidth)

System calibration through

anonymized and aggregated

results; use for individual

treatment limited to long-window

aggregation with due process

(disclosure to the person and an

opportunity to respond);

prohibition of demotion or

revocation of authority based on a

single result; gaming deterrence

through publicizing the fact that

calibration is performed (Section 5,

commentary on Proposition 12)

Propositions 7 and

12 (separation of

gaming deterrence

from surveillance

pressure)

Layer separation

of feedback and

diversity

preservation at

the RAG layer

Separate by layer the

management of the

audit's interpretive

context from the

training of the base

model, and preserve

judgment diversity at

the RAG layer

Management of interpretive

context (context metadata) at the

prompt and evaluation-criteria

layers; avoidance of full-volume

input into base-model retraining;

selection and recording of what is

fed into training; temporal decay of

logs at the RAG layer and sourcediversity

weighting

Proposition 5

(model condition

(iii); Shumailov et

al. 2024)

Working Paper | Ageless Management in the AI Era 136

(preservation of

generative diversity)

Compensation

and contracts

Adequacy of

compensation; do not

substitute honorific

titles or a sense of

purpose for

compensation

Written contracts with explicit

compensation; market-referenced

remuneration for experiential

audit; contractual provision of

accident compensation and

consultation channels for nonemployment

participants

Propositions 9 and

10

Youth

participation

Primacy of educational

purpose

Priority of schooling; restrictions on

the commercial use of deliverables;

involvement of educational

institutions; guardrail design for AI

use

Proposition 10

Voice channels

(flattening the

power gradient)

Audit dissent reaches

decision-making

regardless of status, age,

or contract form

(organizational

condition (ii) of

Proposition 5)

Recording of dissent and detected

items with escalation paths; a duty

to respond to audit opinions

(recording the reasons for

rejection); operations that exclude

the speaker's attributes and

contract form from the evaluation

of flags (Section 6.3)

Proposition 5

(extension of BCM's

Belonging)

Measurement Measure states, not the

introduction of

measures

Continuous measurement of load,

fatigue, and cognitive engagement

with threshold-based operation;

brain-capital indicators by

participation stratum (Section 9)

Propositions 7 and

8 (BCM's distinction

between states and

measures)

Note: This table is a skeleton, not a regulatory standard or a finished code. The effectiveness of each concrete measure

(its effects on audit quality, brain health, and the prevention of exploitation) is untested; the cognitive-load side

connects to Hypotheses H1 and H3, and the economic-protection side to the institutional comparison of Section 7 and

the testing frameworks of Propositions 9 and 10. Numerical values such as the time ceiling are initial design values

inherited from the practical guides of the original draft, not empirically grounded thresholds. Further, that each item

of this table is operated in a form whose satisfaction can be judged prior to the observation of outcomes is a

requirement for the testability of Propositions 8 and 11 (the refutation conditions of Section 5; the measurement

framework of Section 9).

The order of implementation also inherits BCM's discipline. BCM derived the conclusions

that the three constraints operate multiplicatively, that the weakest link governs the whole,

and that the foundation comes first and development comes after. The ecosystem version has

the same form. Introducing sophisticated load measurement while the adequacy of

compensation and contracts (the extended Base) is lacking only means that participants are

measured under exploitative terms; assembling a decorrelation portfolio while voice

channels (the extended Belonging) are lacking means that detection ends in silence. The

domains of Table 9 are not a parallel checklist: they have an order in which the foundational

domains of compensation-and-contracts and cognitive-load management come first, and the

refinement of skill maintenance and measurement rests on top of them.

Working Paper | Ageless Management in the AI Era 137

Brain Safety is not an ancillary welfare benefit of Ageless Management. Following the logic

of Sections 4 and 6, it is the very supply condition of oversight value. Without budget

management of cognitive load, detection probability collapses (Section 8.1); without the

adequacy of compensation, the ecosystem degenerates into exploitation and the sustainability

of participation collapses (Section 8.2); without the design of exercise opportunities,

experiential audit capacity itself depreciates (Section 8.3). It is in this sense that Proposition 8

asserts bidirectional capital formation only for multigenerational ecosystems that satisfy exante

verifiable conditions — the connection modes of Definition 7 and the satisfaction of

Definition 8. The framework for measuring and verifying the satisfaction of these conditions

is the subject of the next section.

Regarding the operation of this conditional claim, a normative statement and an epistemic

statement should be clearly separated. At the normative level, this paper claims the following:

an organization that cannot meet the standards of Brain Safety has no standing to claim the

benefits of Ageless Management. This is a statement about the legitimacy of an organization

attaching this paper's name to its own practice. The epistemic level — the judgment of

condition satisfaction in testing Propositions 8 and 11 — is a different matter. Condition

satisfaction is judged by an ex-ante checklist preceding the observation of outcomes, that is,

by the satisfaction of the items of Table 9 operated in a form that a third party can judge

before seeing the results. If, under a design judged ex ante to satisfy the conditions, a

systematic ex-post deterioration of the brain-capital indicators of any participating stratum is

observed, Propositions 8 and 11 are rejected. Retroactive denial of the conditions from expost

outcomes — claiming that "the conditions were not satisfied after all" — is not admitted

as a defense of the propositions (see the refutation conditions of Propositions 8 and 11). The

normative statement disciplines organizations' claims; the epistemic statement disciplines

this paper's theory. An operation that conflates the two and reclassifies every failure case as a

deficiency of Brain Safety is a maneuver that immunizes the theory against refutation, and is

not a correct application of this paper's framework. That this paper asserts its propositions

only conditionally is not rhetoric: the conditions — the design standards of this section — are

the substance of implementation, and at the same time the entrance to verification.

9. Measurement and Verification

The definitions and propositions of Section 5 were developed into design theory, institutional

analysis, and safety and health in Sections 6 through 8. Most of this paper's propositions,

however, remain untested, and a theory stands as a scientific proposal only when it is

equipped with a map of verification. This section draws that map. Section 9.1 organizes three

hypotheses (H1–H3) together with their verification designs, identification threats, and costs

(Table 8), and Section 9.2 presents the design of H3, the highest-priority and lowest-cost

experiment (Figure 5). Section 9.3 extends the K × u framework of Brain Capital Management

(Kadowaki 2026e) to organization-level indicators for the multigenerational ecosystem, and

Working Paper | Ageless Management in the AI Era 138

Section 9.4 sets out a roadmap for implementation. Two principles run through this section.

First, begin with the tests that are lowest in cost and that strike directly at the theory's core

mechanisms. Second, measurement is an instrument for role allocation and protection, not

an instrument of selection and exclusion. The risk of the latter misuse is flagged at various

points in this section and confronted head-on in Section 10.

9.1 The Map of Verification — Three Hypotheses

This paper's twelve propositions were each presented with a refutation condition attached

(Section 5). But enumerating refutation conditions and having an executable verification plan

are two different things. This section bundles the core of the propositions into three testable

hypotheses and specifies a verification design for each. H1 targets the individual-health side

of Proposition 7 (the cognitive-engagement pathway) and Proposition 8 (bidirectional capital

formation); H2 targets the organizational-outcome side of Proposition 5 (generational

decorrelation) and Proposition 6 (the AI-mediated diversity effect); and H3 targets the

theory's core mechanism, Proposition 3 (the non-compressibility of experience) and

Proposition 4 (the formation of experiential audit capacity).

Hypothesis H1 (Brain-Health Hypothesis)

Super-seniors engaged in Gc-exercising, AI-complemented roles show a lower rate of

cognitive decline (MoCA, etc.) than same-age peers who have exited work, mediated by

the maintenance of cognitive engagement. Identification strategy: use the exogenous

variation of pension reforms and mandatory-retirement rules as instrumental variables,

or a quasi-experiment exploiting the exogeneity of reasons for working. Explicitly

address health selection and reverse causation. Rigorously control for years elapsed

since leaving work (time since retirement) as a covariate — in order to distinguish the

effect of working from the cumulative natural depreciation over the period out of work.

H1 is the most expensive to test and the hardest to identify of this paper's hypotheses. As

confirmed in Section 3, the association between work and cognitive function is doubly

contaminated by health selection (the healthy worker effect) and reverse causation (the

prodromal phase of cognitive decline hastens retirement), and, moreover, conclusions can

reverse depending on the choice of instrumental variable. Estimates using international

differences in pension and tax institutions as instruments suggested a large negative effect of

retirement (Rohwedder & Willis 2010), whereas estimates using employer-offered earlyretirement

incentive windows as instruments did not accept the association as causal and, for

blue-collar workers, instead reported a positive relationship between time in retirement and

cognition (Coe et al. 2012). A systematic review summarizes the evidence as mixed (Meng et

al. 2017), and the average reading of recent causal-inference research likewise goes no

further than "most estimates find that cognitive skills decline after retirement, but the effects

are highly heterogeneous by occupation and by the voluntariness of retirement" (van Ours

Working Paper | Ageless Management in the AI Era 139

2022). Testing H1 must therefore not depend on a single instrumental variable; it must

combine multiple identification strategies — the exogenous variation of institutional reforms

and the exogeneity of reasons for working — in a design that reports the discrepancies

among estimates themselves. H1 furthermore contains a mediation hypothesis. To test the

structure of Proposition 7 — that the mediator is not working hours themselves but the

maintenance of cognitive engagement through occupation in Gc-type roles — measurement

not only of whether one works but of the quality of work (the share of Gc-type components in

the role, the presence or absence of AI complementation) is indispensable. On this point, the

data of existing retirement studies cannot substitute, and the longitudinal study and medical

collaboration of Section 9.4 must be awaited.

Hypothesis H2 (Multigenerational Team Outcome Hypothesis)

Under AI use, mixed teams of "experiential audit (Gc) × youthful prototyping (Gf) × midcareer

orchestration" show significantly higher scores than homogeneous teams on both

the novelty and the feasibility of the ideas produced. Blinded external evaluation. In

team assignment, other demographic attributes such as gender, ethnicity, and cultural

background are handled by stratified randomization or covariate control, identifying

the effect of age-and-experience heterogeneity apart from other diversity dimensions.

Analysis plan: the significance of the interaction alone is not counted as support for

Proposition 6; a decomposition of simple main effects is used to confirm an absolute

improvement in the mixed teams' scores (excluding the false positive in which the

interaction arises solely from the deterioration of homogeneous teams through AI

overtrust). Time to decision, verification effort, and the human-to-human

communication time required to reach agreement (time-to-consensus) are recorded as

cost variables, and outcomes are reported as net benefit after deducting them. A positive

net benefit is a requirement for support; superiority in idea quality alone is not counted

as support.

Testing H2 presupposes an honest recognition of the empirical baseline. As confirmed in

Section 6, the average relation between age diversity and team outcomes is near zero (Joshi &

Roh 2009; Schneid et al. 2016; Wallrich et al. 2024), and positive effects appear only under the

conditions of task complexity and creativity and an inclusive climate (Backes-Gellner & Veen

2013; Wegge et al. 2012). H2 is therefore not a claim of average effect — that

multigenerational composition raises outcomes — but a test of Proposition 6's moderator

structure: the effect turns positive only in the presence of AI-mediated complementarity. The

ideal design is a 2 × 2 team-level randomization crossing team composition (mixed/

homogeneous) with AI use (with/without), testing Proposition 6 up to and including the

disappearance or reversal of the effect in mixed teams without AI use. The composition of the

homogeneous control teams is stipulated here. The primary control is a homogeneous team

composed solely of the mid-career generation; where feasible, a multi-arm comparison adds

Working Paper | Ageless Management in the AI Era 140

homogeneous conditions of juniors only and seniors only — because which generation the

homogeneous teams are drawn from changes both the meaning of the hypothesis and the

number of teams required. In assignment to teams, other demographic attributes such as

gender, ethnicity, and cultural background are handled by stratified randomization (or

covariate control where cell sizes are insufficient) — because if the mixed/homogeneous

contrast is confounded with diversity dimensions other than age-and-experience

heterogeneity, the observed effect cannot be attributed to the axis Proposition 6 asserts; the

outline of the assignment procedure is supplemented in Appendix A (A.6). The analysis plan

additionally specifies a decomposition of simple main effects to identify the source of the

interaction's sign. An apparent interaction produced by AI mediation lowering the

performance of homogeneous teams and an interaction produced by raising the performance

of mixed teams are not equivalent as support for Proposition 6. The task is a creative

divergent task such as drafting new-business proposals, and evaluation is by blinded external

evaluators — with team composition and participant attributes concealed — who score

novelty and feasibility independently. The procedure of blinded third-party scoring follows

the standard method of generative-AI experiments (Noy & Zhang 2023; Dell'Acqua et al. 2023).

The identification threats are the endogeneity of team formation (addressed by

randomization), insufficient statistical power from small team numbers (the team is the unit

of analysis, so the required sample is large), behavioral change from participants sensing the

experimental intent, and breach of blinding when evaluators infer team composition from

style or the like. In light of the meta-analysis finding that human-AI combinations on average

fall below the best single agent (Vaccaro, Almaatouq & Malone 2024), there is no guarantee

that the AI-use condition automatically raises outcomes. Moreover, cost-side measurement is

built into the hypothesis's very criterion of support. Generationally heterogeneous teams may

require longer than homogeneous teams to reach agreement, in order to process and

coordinate the points raised, and unless this cost is booked, the outcome that Proposition 6

defined as net benefit (Section 5) is not measured. H2 therefore records, in addition to time to

decision and verification effort, the human-to-human communication time required to reach

agreement (time-to-consensus) as cost variables, and makes a positive net benefit after their

deduction a requirement for support — if only superiority in idea quality is observed and net

benefit after cost deduction is not positive, H2 is not counted as supported. H2 is a design that

can produce results unfavorable to this paper's theory — no mixed-team advantage, no AImediation

effect, a quality advantage eaten up by costs — and it is for that reason that it

deserves the name of a test.

Working Paper | Ageless Management in the AI Era 141

Hypothesis H3 (Direct Test of the Non-Compressibility of Experience)

A 2 × 2 factorial design [experience level: long-term domain experts vs. juniors] × [AI

assistance: with vs. without]. The task is the verification of AI-generated business

proposals and documents with misinformation and contextual risks embedded. So that

the embedded errors do not depend on the experimenters' preconceptions, tasks blindly

generated from an empirical failure-case dataset of accidents, scandals, and failures that

actually occurred are included. Some of the errors are constructed in both a form

conforming to auditors' industry received views (confirmation-conforming) and a form

departing from them (deviating), simultaneously testing boundary condition (i) of

Proposition 4. The fluency of the task documents (high fluency = stylistically polished

output vs. low fluency) is controlled as a factor or covariate, simultaneously testing

boundary condition (iii). Dependent variables: (i) detection rate of surface errors; (ii)

detection rate of contextual and practical risks; (iii) quality of proposed corrections

(blinded evaluation); (iv) over-rejection rate — the rate of erroneously rejecting the

embedded "correct but convention-defying innovative proposals" (measuring the

boundary at which experience turns into over-auditing). Years elapsed since leaving

work are recorded and controlled. Predictions: on (i), the experience gap narrows with

AI assistance (consistent with prior research), but on (ii) and (iii) a main effect of

experience persists and does not narrow even under AI assistance (a direct test of

Proposition 3). No prediction is placed on (iv); it is an exploratory indicator. Appendix A

gives the detailed experimental protocol.

H3 has the highest priority of the three hypotheses. There are three reasons. First, its cost is

the lowest. It is a task experiment with the individual as the unit of analysis, executable

online, requiring neither longitudinal tracking nor team-level randomization. Second, it

strikes directly at the theory's core. The claim that experiential audit capacity (Definition 4) is

real and is not compressed by AI assistance is the keystone of this paper's entire construction

— the supply-side theory of oversight, the design of the multigenerational ecosystem, the

conversion of social problems into resources — and if H3 is rejected, the theory requires

revision from its foundations. Third, it is prior to the interpretation of H1 and H2. If the

existence of experiential audit capacity cannot be confirmed, H2's mixed-team design loses its

basis, and the definition of H1's treatment — a Gc-exercising role — becomes ambiguous.

Verification should proceed from the cheap and decisive, and H3 is the only hypothesis that

meets that condition. The details of the design are given in the next section.

Table 8 The map of verification — verification designs and identification threats for the three

hypotheses

Hypothesis Verification design

Dependent

variables

Identification

threats

Cost and order

of execution

Working Paper | Ageless Management in the AI Era 142

H1 (brain-health

hypothesis)

Corresponds to

Propositions 7 and

8

Quasi-experiment.

Panel analysis using

the exogenous

variation of pension

reforms and

mandatoryretirement

rules as

instrumental

variables, or

comparison

exploiting the

exogeneity of

reasons for working.

Mediation analysis

used in combination

Rate of cognitive

decline (MoCA, etc.).

Cognitive

engagement as

mediating variable.

The share of Gc-type

components in the

role and the

presence or absence

of AI

complementation

included in the

definition of

treatment

Health selection

(healthy worker

effect), reverse

causation,

voluntariness of

retirement. Prior

examples of

conclusions

reversing with

the choice of

instrumental

variable

(Rohwedder &

Willis 2010 vs.

Coe et al. 2012).

Measurement

error in the

mediating

variable

High (requires

years of

longitudinal

tracking and

medical

collaboration).

Third in order

— conducted in

the third and

fourth stages of

Section 9.4

H2

(multigenerational

team outcome

hypothesis)

Corresponds to

Propositions 5 and

6

Team-level

randomized

experiment. 2 × 2 of

team composition

(mixed/

homogeneous) × AI

use (with/without).

In team assignment,

other demographic

attributes such as

gender, ethnicity,

and cultural

background handled

by stratified

randomization or

covariate control.

Creative divergent

task. Blinded

external evaluation.

Decomposition of

simple main effects

specified in the

analysis plan

(excluding spurious

interactions arising

solely from the

deterioration of

homogeneous

teams)

Blinded scores for

the novelty and

feasibility of the

ideas produced. Time

to decision,

verification effort,

and communication

time to reach

agreement (time-toconsensus)

recorded

as cost variables,

with outcomes

reported as net

benefit after their

deduction. A positive

net benefit is a

requirement for

support. Detectionoverlap

rate as

auxiliary measure

(Section 9.3)

Insufficient

power from the

number of teams.

Behavioral

change from

experiment

participation.

Breach of

blinding

(inferring

composition from

style, etc.).

Artificiality of the

task. The

empirical

baseline of a

near-zero average

effect.

Misattribution of

the source of the

interaction's sign

(addressed by

decomposition of

simple main

effects)

Medium

(requires

simultaneous

execution at the

scale of dozens

of teams).

Second in order

— designed in

light of H3's

results

H3 (noncompressibility

of

experience)

Corresponds to

Individual-level 2 ×

2 factorial

experiment.

Experience level ×

(i) Detection rate of

surface errors; (ii)

detection rate of

contextual and

Confounding of

experience with

chronological age

(addressed by

Low (individual

task, executable

in weeks). First

in order —

Working Paper | Ageless Management in the AI Era 143

Propositions 3 and

4

AI assistance.

Verification task on

AI-generated

documents with

embedded errors.

Embedded errors

include tasks blindly

generated from an

empirical failurecase

dataset,

constructed in both

confirmationconforming

and

deviating forms

(simultaneous test of

boundary condition

(i) of Proposition 4).

Fluency of the task

documents (high/

low fluency)

controlled as a

factor or covariate

(simultaneous test of

boundary condition

(iii)). Years since

leaving work

recorded and

controlled.

Executable online

practical risks; (iii)

quality of proposed

corrections (blinded

evaluation); (iv)

over-rejection rate

(exploratory

indicator)

recruiting highage

lowexperience

cells,

etc.). Construct

validity of the

embedded errors

(experimenter

bias mitigated by

blind generation

from the failurecase

dataset).

Artificiality of the

task.

Measurement

error in domain

knowledge

highest priority,

lowest cost

Note: Cost levels are relative assessments. Each hypothesis is designed so that an unsupportive result connects directly

to the refutation condition of the corresponding proposition (Section 5). For details of H1's identification threats see

Section 3.2; for H2's empirical baseline see Section 6.

The three rows of Table 8 are not independent items of verification; they carry a cascade

structure of rejection. If H3 is rejected, H2's mixed-team design and H1's definition of

treatment, both of which depend on the reality of experiential audit capacity, require

reconstruction, and this paper's theory is revised from its core. If H3 is supported and H2 is

rejected, experiential audit capacity is real at the individual level, but the organizational

exploitation of it through team composition (the design theory of Section 6) is in error. If H3

and H2 are supported and H1 is rejected, Ageless Management survives as a theory of

organizational outcomes but must abandon the claim of spillover to brain health (the

individual side of Propositions 7 and 8). Specifying in advance which combination of results

kills which part of the theory — that is what the phrase "the map of verification" means.

9.2 The Experimental Design of H3

The H3 experiment is a 2 × 2 factorial design (Figure 5). The first factor is experience level,

comparing practitioners with long-term experience in the target domain against juniors with

shallow experience in the same domain. The second factor is AI assistance, with one

Working Paper | Ageless Management in the AI Era 144

condition permitting and one condition not permitting the use of generative AI while

performing the verification task. The task is the verification of business proposals and

practical documents for the domain, produced with generative AI, into which the

experimenters have systematically embedded two types of error. The first type is surface

errors — numerical inconsistencies, missing steps, formal breakdowns — detectable by

careful cross-checking even without domain knowledge. The second type is contextual and

practical risks — reproductions of past failure patterns, collisions with regulation and

commercial practice, missing consideration for stakeholders, ethical risks — corresponding to

what Definition 4 specified as the detection targets of experiential audit capacity. The design

of measuring AI users' performance on tasks with deliberately embedded errors adapts the

outside-the-frontier task method of Dell'Acqua et al. (2023) to the measurement of

experiential audit capacity. Two disciplines are imposed on the construction of the embedded

errors. First, so that the errors do not become reflections of the experimenters'

preconceptions, tasks blindly generated — by producers ignorant of the hypotheses — from

an empirical failure-case dataset of accidents, scandals, and failures that actually occurred

are included. Second, some of the errors are constructed in both a form conforming to

auditors' industry received views (confirmation-conforming) and a form departing from

them (deviating) — in order to test simultaneously, in the same experiment, boundary

condition (i) of Proposition 4: that experience-derived priors aid detection for conventiondeviating

errors, whereas for convention-conforming errors detection can instead be

impeded by the synergy of confirmation bias and automation bias (Section 5). In addition,

fluency is controlled as a property of the task documents themselves. Two versions of

identical content, manipulating only stylistic polish — a high-fluency and a low-fluency

version — are prepared, and fluency is treated as a factor or covariate — in order to test

simultaneously, in the same experiment as the main test, boundary condition (iii) of

Proposition 4: that high-fluency output can, through the effect of processing fluency, raise the

threshold of auditors' cognitive sense of unease and impede detection without any deficit in

the auditor's capacity (the manipulation procedure and manipulation checks are in Appendix

A).

There are four dependent variables. (i) The detection rate of surface errors; (ii) the

detection rate of contextual and practical risks; (iii) the quality of the proposed corrections

for the problems detected, scored by blinded evaluators from whom participant attributes

and conditions are concealed. (iv) is the over-rejection rate — the rate at which the "correct

but convention-defying innovative proposals" embedded in the task were erroneously

rejected as problems — measuring the boundary at which experience turns into overauditing

(excessive risk aversion). The predictions are asymmetric. On (i), AI assistance

narrows the experience gap — this is the direction consistent with the compression evidence

organized in Section 4.1, and this paper's theory in fact predicts compression here. On (ii) and

(iii), a main effect of experience persists and does not narrow even under AI assistance. This

is the direct test of Proposition 3. If, under the AI-assisted condition, the gap in contextual

Working Paper | Ageless Management in the AI Era 145

detection rates between long-term experts and juniors disappears or reverses, Proposition 3

is rejected exactly as its refutation condition states, and this paper's theory, resting on

experiential audit capacity, loses its core. Conversely, if the main effect of experience persists

on (ii) and (iii), that is the first direct evidence supporting the existence of experiential audit

capacity. No directional prediction is placed on (iv); it is an exploratory indicator — whether

over-rejection increases with experience, and whether it is distributed as the flip side of

confirmation-conforming misses, is information that demarcates the location of Proposition

4's boundary, and it is reported whatever the results.

Three devices are built into this design. The first is the separation of experience from

chronological age. As argued in Section 4.4, this paper's reattribution thesis asserts that "what

predicts detection capacity is experience, not age" (Proposition 4). Recruitment of participants

therefore deliberately breaks the correlation between experience and age — including older

career-changers and returners with shallow experience in the domain, and young early

specializers with long experience — making it possible to estimate the independent effect of

age controlling for years of experience. Proposition 4's refutation condition (that

chronological age still independently predicts detection capacity after controlling for years of

experience) becomes testable only through this design. The second is the explicit

manipulation of AI assistance as a factor. This transcribes into the measurement design the

divergence of Section 4.3 — that assisted performance does not guarantee unassisted

competence (Bastani et al. 2025) — and is also a requirement for bringing onto the test bench

the relation between BCM's (Kadowaki 2026e) "protected unassisted practice" and H3, which

this paper newly asserts and which is untested. The third is the recording and control of years

elapsed since leaving work. The low performance of participants long away from practice

may reflect not the absence of experiential audit capacity but depreciation over the idle

period, and unless these two are distinguished, the experience-level factor is confounded

with the effect of timing of exit. The same reason that H1 demands rigorous control of years

since retirement applies to H3 at the level of individual differences (eligibility criteria and

controls are detailed in Appendix A).

Working Paper | Ageless Management in the AI Era 146

Figure 5 The 2 × 2 factorial design of Hypothesis H3. Experience level (long-term domain experts /

juniors) is crossed with AI assistance (with / without), and a verification task on AI-generated documents

with embedded errors is imposed. The predictions are asymmetric across dependent variables: on the

detection rate of surface errors (i) the experience gap narrows with AI assistance, but on the detection

rate of contextual and practical risks (ii) and the quality of proposed corrections (iii) the main effect of

experience persists even under AI assistance. This asymmetry constitutes the direct test of Proposition 3

(the non-compressibility of experience). The over-rejection rate (iv) is recorded as an exploratory

indicator with no prediction placed on it. Note: The detailed experimental protocol — participant

requirements, task construction, the typology of embedded errors, evaluator blinding, the testing plan,

and an approximate required sample size — is given in Appendix A.

Only the essentials of execution are noted in the main text. The task domain must match

the participants' domain of experience, and execution in a single domain (for example,

business planning in a specific industry) is the first step. Generalization of the results requires

replication in multiple domains. Evaluator blinding includes a robustness check against the

possibility that experience level is inferred from the style of participants' responses. Because

chance hits contaminate performance on detection tasks, the rate of flagging non-error

locations as errors (the false-alarm rate) is recorded in parallel, and sensitivity d′ and

response criterion c are estimated separately within the framework of signal detection theory

(SDT) — because detection rates alone cannot distinguish true sensitivity from a "suspect

everything" response bias (the details of the analysis plan are in Appendix A). These, together

with participant requirements, task construction, the typology of embedded errors, evaluator

blinding, the testing plan, and the approximate required sample size, are detailed in the

experimental protocol of Appendix A. Furthermore, execution requires preregistration of the

hypotheses and testing plan and publication of the task materials, the list of embedded errors,

and the evaluation criteria. The compression evidence this paper relies on includes

preregistered experiments (Noy & Zhang 2023; Dell'Acqua et al. 2023), and in the present case,

where the theory's proposer may be involved in its verification (see the conflict of interest in

H3: 2×2 factorial design (experience level × AI assistance)

Junior × no AI assistance

Baseline

Junior × AI assistance

Surface detection improves (predicted)

Experienced × no AI assistance

Contextual-detection reference

Experienced × AI assistance

Contextual gap persists (predicted)

Experience

level

DVs: (i) surface-error detection (ii) contextual/practical-risk detection (iii) quality of corrections (blind-rated)

(iv) over-rejection rate — errors span confirmatory/anomalous and high/low fluency — prediction: gap shrinks on (i), persists on (ii)(iii); (iv) Working Paper | Ageless Management in the AI Era 147

Section 10.6), making after-the-fact substitution of hypotheses structurally impossible is

demanded even more strongly than in an ordinary experiment.

9.3 Organization-Level Indicators — Extending K × u to the Ecosystem

In parallel with hypothesis testing, an organization implementing Ageless Management needs

indicators for measuring its own state. This paper extends the framework of brain capital =

stock (K) × utilization rate (u), formalized by BCM (Kadowaki 2026e), from the employees of a

single firm to the participants in a multigenerational ecosystem not limited to the

employment boundary (Definition 7). What Proposition 8 asserts is that an ecosystem

satisfying its conditions increases K and u bidirectionally — restrained depreciation of K

among super-seniors, early formation of K among youth and socially marginalized groups,

and a rise in u for the organization. Making this claim measurable requires participant-level

K indicators, organization-level u indicators, and, as the third dimension this paper adds,

indicators of decorrelation. Furthermore, since these indicators become inputs to role

allocation, the ramparts that keep measurement itself robust (the operational requirements

of Proposition 12) must be built into the operation of the indicators.

First, participants' K-maintenance indicators. As a premise, the structure of K should be

restated. As made explicit in the commentary on Proposition 8 (Section 5), K is not a single

number, nor is it an additive composite score of three dimensions — (1) clinical cognitivefunction

scores (standardized tests such as MoCA), (2) structural and functional brain

measures, and (3) standardized domain-knowledge measures. K is conceptualized as a

hierarchical function K = 1[Kbase ≥ θ] × f(Kbase, Kdomain), gated by the foundational

cognitive function Kbase measured by (1) and (2) exceeding a threshold θ, on top of which the

domain knowledge and operational schemata Kdomain of (3) are multiplied (an inheritance

and refinement of BCM's measurement framework). This structure has two implications for

measurement procedure. The dimensions must be measured and reported independently,

and no additive composite may be constructed in which high domain-knowledge scores offset

declines in clinical scores — that would resurrect, on the measurement side, the substitution

the hierarchical function forbids. Also, the gate judgment on Kbase logically precedes the

measurement of the other dimensions — this threshold is the same gate as the cognitivescreening

lower bound of Proposition 2(b) (the theory's internal consistency), and its

judgment is also the judgment of the boundary for allocation to experiential-audit roles. The

"systematic deterioration of brain-capital indicators" in Proposition 8's refutation condition is

adjudicated by dimension-wise measurement in accordance with this hierarchical structure.

Of these, direct measurement of (1) and (2) should be conducted only under medical

collaboration (Section 9.4); what an organization can handle day to day is limited to proxy

indicators for them. There are three candidates. (1) Measurement of cognitive engagement —

the share of Gc-type components in the role occupied, and the subjective absorption in and

challenge level of role occupation. This is the mediating variable of Proposition 7 and also

connects to the testing of H1. (2) Periodic performance on detection tasks under unassisted

Working Paper | Ageless Management in the AI Era 148

conditions — small H3-type tasks administered periodically without assistance, tracking the

level and trajectory of experiential audit capacity. This can function as an early warning

against the skill-depreciation risk discussed in Section 4.3, but it must be stated repeatedly

that whether unassisted practice is itself effective for maintaining capacity is untested. (3)

Continuity of participation and changes of role — tracking whether exits due to the cognitive

bottleneck (Definition 3) are actually decreasing. All of these are apprehensions of state, not

measurements of the effect of interventions, and the distinction between state and

intervention that BCM made a discipline is maintained here as well.

Second, the organization's u indicators. The utilization rate (u) is operationalized as the

share of the cognitive assets the ecosystem could connect that are actually connected to roles.

Concretely: (1) the fit rate between participants' held assets (measured cognitive

characteristics and domain experience) and their current role requirements; (2) the fill rate

of Gc-type and audit-type roles — how far the supply of experiential audit capacity is

connected to oversight demand (the verification function shown in Section 4 to be growing

scarce); and (3) the distribution of participation and remuneration by contractual form —

data to be read alongside the Brain Safety indicators (Section 8, Table 9) to check that nonemployment

participation is not falling into the protection vacuum (Proposition 9). A rise in

u, if it becomes an end in itself, slides into exploitation (Definition 8(ii)). u indicators must

always be reported paired with indicators of load and remuneration.

Third, the measurement of decorrelation. Proposition 5 (generational decorrelation) can be

measured directly within an organization. The core of the proposed measurement is the

detection-overlap rate. The same set of AI outputs is audited independently by multiple

supervisors, and the sets of missed errors are compared across supervisors. The overlap rate

of misses for generationally homogeneous supervisor pairs is contrasted with that for

heterogeneous pairs, and if the latter is systematically lower, Proposition 5's claim — that the

correlation of blind spots is lower between generations than within them — is supported.

Conversely, if there is no difference in overlap rates, Proposition 5 is rejected exactly as its

refutation condition states. This measurement adds the audit channel as a factor — a

condition in which auditors access the raw output unmediated, and a condition in which they

access it through summarization and pre-screening by the same AI. Proposition 5 is

formulated conditional on unmediated access (Section 5), and the claim of its latter half —

that sharing an AI-mediated channel destroys, at the channel level, the decorrelation supplied

by generational heterogeneity — can be directly tested through this design as the difference

in overlap rates between the unmediated and mediated conditions. For the same reason, a

foundation-model condition is also added as a factor — a condition in which the audited AI

outputs derive from a single foundation model and a condition in which they derive from

multiple models of different architectures and developers. As condition (iii) of Proposition 5

asserts, in a situation where a single model's systematic blind spots dominate all outputs,

generational heterogeneity on the human side should not be able to override them, and this

claim becomes testable as an interaction in which the reduction of overlap rates from

Working Paper | Ageless Management in the AI Era 149

generational heterogeneity shrinks or disappears under the single-model condition. This

measurement is an operationalization of the independence side of WP8's (Kadowaki 2026h)

oversight value = independence × detection probability; the detection-probability side is

measured separately with H3-type tasks — to conflate the two is to commit the error of rating

highly a supervisor pool whose blind spots merely fail to overlap while detecting nothing. It

should be noted that experimental evidence that information exchange broadens and factual

errors decrease in diversely composed groups exists in the context of racial diversity

(Sommers 2006), but replication on the age-and-generation axis is unconfirmed, and the

measurement of detection-overlap rates is precisely an attempt to fill that gap.

Fourth, the robustness of measurement itself — the ramparts against Goodhart's law. Since

this section's indicators, above all detection-task performance and role-fit rates, become

inputs to dynamic role allocation (Definition 1), measurement connects directly to allocation

and is therefore structurally exposed to gaming (Proposition 12). The operational

requirements of Proposition 12 are here given concrete form as measurement procedures.

First, the non-periodic replacement of detection and measurement tasks. Repetition of the

same tasks permits overfitting to the tasks — improvement in test-taking rather than in

experiential audit capacity — so the task pool is refreshed without notice, and discontinuities

in performance across refreshes are monitored as an indicator of overfitting. Second,

backtesting against real-work logs. Whether performance on the verification tasks used as

grounds for role allocation diverges from audit performance in real work — the record of

detections and misses measured by ex-post verification through unmediated audit of realwork

logs — is checked periodically, and any indicator for which divergence is systematically

observed is removed from the grounds of allocation. Only operation accompanied by ex-post

verification, not one-off test scores, can sustain measurement as a ground of allocation. Third,

unannounced, real-time blind in-situ verification (unannounced in-situ verification). What

backtesting verifies is the logs that remain, and as long as how logs are kept is itself under the

control of those subject to allocation, a deeper level of gaming remains — adaptation toward

"ways of keeping logs that backtesting does not detect" (Proposition 12). Accordingly, a

procedure of blindly observing, in real time and without advance notice, randomly selected

audit scenes in real work, and collating the judgments made on the spot with the entries in

the logs, is operated as a pair with backtesting — a double rampart against the two levels of

gaming: adaptation to the test and adaptation of log formation. In addition, monitoring of

each auditor's false-alarm rate is built into the operation of the indicators. Monitoring only

the detection rate (hit rate) cannot distinguish a shift of the response criterion toward

suspecting everything — a criterion shift that raises apparent detection while increasing false

alarms and verification costs — from a true improvement in detection capacity. Sensitivity d′

and response criterion c are separated within the framework of signal detection theory

(Appendix A), and each auditor's false-alarm rate is continuously monitored as an estimate of

the response criterion c. This measurement is the monitoring of the same quantity as the

rampart against gridlock through excessive flagging (Section 8), and it is also a condition for

Working Paper | Ageless Management in the AI Era 150

preserving the interpretability of detection performance as an indicator. These are, however,

ramparts, not guarantees. No means exists to completely prevent divergence between

measurement and true capacity, and the implications of that residual risk are discussed in

Section 10.3.

Finally, the risk of measurement misuse is stated explicitly. This section's indicators are

designed for improving role allocation and monitoring Brain Safety; the moment they are

diverted into instruments for rating individuals, adjudicating exit, or cutting remuneration,

Ageless Management degenerates into a device that merely replaces discrimination by

chronological age with discrimination by measurement. A decline in K-maintenance

indicators is an occasion for support and role redesign, not a ground for exclusion. This

danger is inherent in this paper's theory, and it is confronted head-on in Section 10.3.

9.4 Implementation Roadmap

Verification and implementation cannot be separated. The propositions of Ageless

Management include some that can be tested only inside an implemented multigenerational

ecosystem (Propositions 8 and 11), and conversely, implementation without verification does

not rise above the level of an ideal. This section presents a four-stage roadmap that raises cost

and the strength of causal inference step by step.

The first stage is a pilot. In the small-scale ecosystems of the author's own organization and

its partners, the H3-type verification tasks and the measurement of detection-overlap rates

are trialed, and the collection procedures for the indicators of Section 9.3 — proxy indicators

of K maintenance, u indicators, Brain Safety indicators — are established. The purpose of this

stage is not the estimation of effects but the verification of the feasibility of measurement. For

super-seniors, for whom participation in the tasks is itself a load, applying Brain Safety

(Section 8) — load ceilings and the adequacy of compensation — to the design of the

measurement itself is also among the procedures to be established at this stage.

The second stage is case studies. WP3 of this series (Kadowaki 2026c) collated theory with

observation for the theoretical framework of enterprise redefinition through structured

observation of company cases based on public information. The same observational method

is re-applied to this paper's framework: cases of organizations in the process of implementing

multigenerational role composition, non-employment participation, and AI-mediated

complementarity are described in structured form along the design variables of Section 6 (the

default design of Table 5 and deviations from it, the operation of dynamic role allocation, the

organizational design of decorrelation). Case studies do not settle causation, but they reveal

whether the propositions take observable form in real organizations and whether there are

unanticipated failure modes. The selection bias by which observed cases skew toward

successes is unavoidable here, and its implications must be interpreted together with the

survivorship-bias discussion of Section 10.4.

Working Paper | Ageless Management in the AI Era 151

The third stage is a longitudinal study. Ecosystem participants across multiple organizations

are formed into a panel, and the K and u indicators together with data on roles, contracts, and

load are tracked continuously. Because a single organization lacks both sample size and

exogenous variation, the formation of a multi-organization consortium is a precondition. At

this stage, H1's identification strategy comes into view. Exogenous institutional variation —

reform of Japan's in-work old-age pension offset (Section 7) or changes to mandatoryretirement

rules — is a candidate instrumental variable linking changes in work and role

occupation to cognitive outcomes, and if the panel is maintained across an institutional

change, it can be used as a quasi-experiment. As seen in Section 3, however, this field has a

history of conclusions reversing with the choice of instrument. From the outset, the design

incorporates the discipline of not betting on a single identification strategy and of reporting

including the discrepancies among estimates. The verification tasks at this stage include

estimating the depreciation function of detection capacity with years since leaving work as a

continuous variable. For the stratum long past exit, to which H3's eligibility criterion (within

five years of leaving practice) does not permit extrapolation, the speed at which experiential

audit capacity depreciates with idle time is an unresolved empirical question that directly

demarcates the scope of the resource-conversion claim (Propositions 2 and 11).

The fourth stage is medical collaboration. Measuring the rate of cognitive decline (MoCA,

etc.), H1's dependent variable, requires the ethics review, clinical expertise, and long-term

follow-up infrastructure of medical research, and cannot be carried out by management

studies alone. There is precedent. The effects of structured productive social engagement by

older adults on cognitive function and brain structure have been tested jointly by medicine

and social science in the randomized controlled trials of Experience Corps (Carlson et al.

2008; Carlson et al. 2015). These, however, are effects of fifteen hours per week of structured

volunteering in a limited population (chiefly low-income urban populations in the United

States), and extrapolation to work requires bridging through the superordinate concept of

productive social engagement (Section 3). What corresponds to this in the context of Ageless

Management is the randomized or quasi-experimental evaluation of the intervention of

occupation in Gc-exercising, AI-complemented roles. Including the design of ethically

permissible forms of assignment — such as randomizing the order in which roles are offered

to those wishing to participate — this stage presupposes joint design with medical

researchers.

This roadmap is at the same time a declaration of the provisional character of this paper's

claims. The first and second stages can be undertaken by practitioners including the author's

own organization, but the execution of the third and fourth stages — and above all the

interpretation of results — must be open to execution and replication by independent

researchers. A configuration in which the proposer of a theory monopolizes its verification is

to be avoided, particularly under the conflict of interest disclosed in Section 10.6. The results

of each stage — above all a rejection of H3 — connect directly to the revision or abandonment

of the propositions. Until verified, this paper's propositions remain proposals.

Working Paper | Ageless Management in the AI Era 152

10. Limitations and Self-Critique

This section is not a ritual enumeration of the paper's limitations. The theory of Ageless

Management carries an internal structure by which it can itself turn into a new apparatus of

discrimination, selection, and exclusion, and the proposer of the theory knows the pathways

of that transformation most concretely. This section discusses six limitations — the limits of

the evidence, new stereotyping, the risks of measurement, survivorship bias, economic

presuppositions, and conflict of interest — naming names wherever possible. The paper's

core thesis bears restating. Whereas conventional senior-employment and D&I arguments

have rested on a paradigm of "accommodation and compensation (cost/CSR)," this paper

presents a management model that, through bidirectional complementarity between

heterogeneous cognitive abilities mediated by AI, converts multigenerational and diverse

talent into co-creating agents of the managerial resource of brain capital. The sections that

follow specify the conditions under which this thesis fails, and the conditions under which,

even succeeding, it does harm.

10.1 The Limits of the Evidence — Most of the Propositions Are Untested

Begin with the most basic limitation. Of this paper's twelve propositions, not one has been

directly tested. The empirical work this paper relies on is in every case not a test of the

propositions themselves but peripheral evidence consistent with them. And that peripheral

evidence itself has the following three weaknesses.

The first is misalignment of axes. What the empirical work on compression, oversight, and

depreciation organized in Section 4 measured was tenure, skill, and expertise, not

chronological age, and no study that operationalized age or years of tenure and measured

performance in AI oversight and verification exists within the range of this paper's search

(Section 4.4, search record). Neither the reality of experiential audit capacity (Definition 4)

nor the claim that it is not compressed by AI assistance (Proposition 3) has direct evidence

until Hypothesis H3 is carried out. That the keystone of this paper's theory is placed at the

point where the evidence is currently thinnest should be stated explicitly.

The second is the mixed character of the evidence. The relation between work and brain

health (the background of Propositions 7 and 8) is a field in which the sign of the estimates

has reversed with the choice of instrumental variable (Rohwedder & Willis 2010 vs. Coe et al.

2012), and the summation of the systematic reviews goes no further than "the evidence is

mixed, with large research gaps" (Meng et al. 2017; van Ours 2022). On age diversity and team

outcomes (the background of Propositions 5 and 6), the meta-analytic average effect is near

zero (Joshi & Roh 2009; Schneid et al. 2016; Wallrich et al. 2024), and this paper's Proposition 6

is a hypothesis that stacks an untested moderator — AI mediation — on top of this headwind

baseline. Citing only the favorable side of the estimates would make this paper's claims look

strong, but that is not the state of the evidence.

Working Paper | Ageless Management in the AI Era 153

The third is the mixture of evidence grades. This paper has referred, alongside peerreviewed

research, to preprints, technical reports, surveys by NGOs and membership

organizations, and practitioners' essays, with the grades made explicit. In particular, the

headwind data of Section 4.5 and parts of the description of institutional operation in Section

7 rest on grey literature, and these are not measurements of ability or outcomes. Making

grades explicit is a requirement of honesty, but it does not cure the weakness of the evidence

itself. In addition, there are limits of external validity. Much of the empirical work this paper

cites is data from the United States and Europe, often from a single firm, a single occupation,

or a limited population (noted individually at various points in Sections 3, 4, and 6), and its

transferability to the context of Japanese employment practice, pension institutions, and

elderly employment is itself an assumption requiring verification. The institutional

comparison of Section 7 has done no more than roughly demarcate the conditions of that

transfer. Taken together, the accurate positioning of this paper is not a report of established

facts but a proposal of theory equipped with refutation conditions and a map of verification

(Section 9). This qualification, stated in the abstract, becomes all the more important when

read together with the COI disclosure at the end of this section (Section 10.6).

Alongside the limits of the evidence, the limits of the concepts should also be stated. This

paper's theory is built on the two-axis classification of intelligence into Gf and Gc, but as

Definition 2 made explicit, this classification is a relative weighting on a continuum, not a

binary, and many real tasks are inseparable bundles of both components. If, as the terms Gftype

and Gc-type recur through this paper, they harden in the reader's mind into binary

categories, that is a misreading of Definition 2 — and at the same time a misreading invited

by this paper's own mode of exposition. Moreover, the construct validity of experiential audit

capacity (Definition 4) is not established. Years of experience, used to operationalize longterm

domain experience, is a coarse proxy variable, and there is no guarantee that the same

years of experience form the same detection capacity — the quality of experience (exposure

to failure cases, the density of feedback) is very likely the true explanatory variable, and this

refinement can only be carried out once H3's results are in.

10.2 The Risk of New Stereotyping — A Self-Critique of Table 5

This paper has repeatedly used the role image of "seniors as auditors." Although this image

was introduced in order to dismantle the presumption of ability from age, it carries the risk

of itself hardening into a new presumption. The statement that older people are suited to

auditing carries on its reverse side the implication that older people are not suited to

execution, and it revives the very operation Definition 1 was meant to remove: inferring role

aptitude from chronological age. Even when the direction is celebratory, the structure of

inferring an individual's characteristics from age is isomorphic with age discrimination, and

the fixation of seniors-as-auditors can become age discrimination's inverted twin.

This danger lies not outside this paper but inside it. To name names, it is Table 5 of Section

6. Table 5 presented as a "default design" a role composition assigning experiential audit and

Working Paper | Ageless Management in the AI Era 154

contextual evaluation to super-seniors, rapid prototyping to youth, orchestration to the midcareer

generation, and lived-experience problem perception to socially marginalized groups.

This paper positioned it as a convenience of departure and made deviation through dynamic

role allocation (Proposition 12) the rule, but the danger remains that practice will operate the

table as a norm. Tables circulate faster than theories. If Table 5 alone is excerpted and

transcribed into training materials and HR systems, then older people who want and can

perform Gf-type roles, early specializers with high experiential audit capacity at a young age

(the thought experiment of Section 4.4), and socially marginalized people who want to

participate in execution rather than audit are rendered invisible once again inside the new

template. That low age stereotyping is a condition for the success of age diversity is also

empirically suggested (Wegge et al. 2012), and if this paper's schema supplies new

stereotypes, that is a self-destructive consequence that undermines the paper's own

theoretical premises.

The danger of fixation is not limited to super-seniors. Table 5 assigns youth to rapid

prototyping, but if this role image hardens, youth will keep specializing in AI-assisted

execution, and by the crutch-effect logic seen in Section 4.3 they risk being deprived of the

opportunity to form contextual verification capacity — that is, the acquisition pathway of

future experiential audit capacity. This stands in direct tension with Proposition 8, which

asserts the early formation of K among youth. Unless the role design of the multigenerational

ecosystem includes intergenerational role transition (the staged entry of youth into audit

experience), this paper's schema fixes the present division of labor and dries up the future

supply. Likewise, the description that positions socially marginalized groups as suppliers of

lived-experience problem perception carries a pathway of sliding into the exploitation of

lived experience — tokenism in which opinions alone are solicited while decisions and

rewards are withheld. For the provision of a lived-experience perspective to be respected as a

role means that it carries compensation and decision rights (Section 8), and it must be

distinguished from the staging of symbolic participation.

Three counter-devices are made explicit. First, Definition 1 removes chronological age from

the criteria of role allocation, leaving only measured characteristics, experience, health

status, and the person's own intent. Second, Proposition 12 asserts the inferiority of age-fixed

allocation, on the grounds of intra-individual variation in cognitive characteristics and the

overlap of distributions across age groups, and attaches a refutation condition. Third, the

operating model of Section 6, as the name dynamic role allocation indicates, requires periodic

re-measurement and re-allocation of roles. But the existence of counter-devices does not

mean the disappearance of the danger. What this paper can do extends no further than to

state here in plain words: any reading that cites Table 5 as a norm is a misreading of this

paper.

Deeper still beneath the danger of stereotyping lies the critique of commodification. This

paper's vocabulary — brain capital, stock and utilization rate, the decomposition of job

bundles into Gf-type and Gc-type components, the connection of cognitive assets — can be

Working Paper | Ageless Management in the AI Era 155

read as a schema that disassembles human beings into parts of brain function, selects the

parts with market value, and reuses them as managerial resources, and this reading cannot

be dismissed as a mere misreading. This paper has in fact consistently discussed the

operations of decomposing jobs into components, measuring individuals' cognitive

characteristics, and optimizing their connection, and what the critique points at is not this

paper's periphery but its method itself. This paper's answer to it is placed, as stated in Section

1.4, in the normative anchor that connects its concept of capital to Sen's (1999) capability

approach. The measurement and connection of brain capital can be legitimate only insofar as

it serves not the maximization of organizational output for its own sake, but the return to

individuals of the freedom of participation that the coarse proxy variable of chronological

age has taken from them — the substantive freedom to live a life one has reason to value.

That Definition 1 includes the person's own intent among the allocation criteria, that

Definition 8 makes protection a precondition of participation, and that Propositions 8 and 11

demand the bidirectionality of benefits across all participating strata are the expressions of

this limitation inside the theory. But the existence of the anchor does not extinguish the

danger of the vocabulary. The words capital, asset, and utilization rate can circulate stripped

of the limiting clauses of intent and dignity, and at that point this paper's framework

functions as a lexicon for the instrumentalization of human beings. This danger too, like

Table 5, is one this paper has itself supplied.

10.3 The Risks of Measurement — Transformation into an Apparatus of

Selection

Definition 1 placed measured cognitive characteristics in the vacancy left by removing

chronological age. This substitution is the core of this paper's theory, and at the same time its

most dangerous point. If measurement is diverted into selection for hiring, treatment, and

exit, Ageless Management will not have abolished age discrimination but merely replaced it

with trait discrimination. Discrimination by chronological age is at least visible, and is

established as an object of legal regulation (Section 7). Selection based on the measurement of

cognitive characteristics, clothed in the appearance of objectivity, is that much harder to

make visible and to regulate. This paper cannot deny the possibility that Definition 1 will be

read as the blueprint of a new apparatus of selection.

There are four concrete dangers. The first is mismeasurement. The measurement of

cognitive characteristics is accompanied by measurement error, situational dependence, and

practice effects, and a low score at a single point in time is misread as a permanent deficit of

ability. The second is gaming. If measurement connects directly to treatment, the

optimization of measured performance becomes an end in itself, and the validity of the

measurement itself collapses. The third is cultural and attribute bias. When the language,

format, and context of a test work against particular groups, measurement automates bias

under the appearance of neutrality. The observation that algorithmic hiring tools can

perpetuate discrimination against people with disabilities (El Morr et al. 2024) is a real

Working Paper | Ageless Management in the AI Era 156

instance, in an adjacent field, of the pathway by which a measurement apparatus that

proclaims inclusion reproduces exclusion. The fourth is use beyond purpose. The pathway by

which cognitive-characteristic data collected for role allocation is diverted to the person's

detriment — exit inducement, insurance, credit — is, under the protection vacuum of nonemployment

participation (Proposition 9), especially unclosed.

The safety devices internal to this paper are the refutation condition of Proposition 12 — if

the costs of trait measurement, mismeasurement, and gaming exceed the gains of dynamic

allocation, the superiority of dynamic allocation is rejected — together with the explicit

prohibition of misuse in Section 9.3 and the Brain Safety of Section 8 (a design that restricts

measurement to an instrument of protection). As to gaming, Proposition 12 wrote into the

proposition itself, as operational requirements, the ramparts against Goodhart's law — the

non-periodic replacement of measurement tasks and ex-post verification through

unmediated audit of real-work logs (backtesting) — and Section 9.3 gave these concrete form

as procedures of indicator operation. But ramparts do not extinguish residual risk. As long as

measurement is connected to allocation and treatment, no means exists to completely

prevent divergence between measured performance and true capacity — overfitting to the

tasks, manipulation of the very real-work logs that backtesting targets, the skewing of effort

toward measurable components. The formula that when a measure becomes a target, it

ceases to be a good measure (Strathern 1997) can be mitigated by replacement and

backtesting but not repealed. The superiority of dynamic role allocation (Proposition 12) must

therefore be read as a net-benefit claim that prices in this residual cost of gaming, and that is

why its refutation condition explicitly names the costs of measurement, mismeasurement,

and gaming. On the institutional side, the minimum line is purpose limitation of

measurement (for role allocation, not for exit adjudication), the person's right of access to

results and right to re-measurement, and the guarantee of alternative pathways upon refusal

of measurement. But explicit statement is not prevention. The observation of an operational

reality in which the misuse of measurement exceeds its gains is a legitimate ground for the

normative criticism that the organization in question lacks the standing to claim the benefits

of Ageless Management. This normative statement, however, must not be diverted into a

defense of the propositions. In testing Propositions 8 and 11, the adjudication of whether the

conditions were satisfied is by ex-ante determination prior to the observation of outcomes

(Section 5). Reclassifying failure cases after the fact as "not having been implemented" in

order to protect the propositions is, as the refutation conditions of both propositions state

explicitly, not admitted as a defense, and the moment this epistemic discipline is broken, the

refutability that is this paper's signboard loses its substance.

Before the risk of misuse lies a design-level problem: the legality of the measurement

apparatus itself. Section 7 compared institutions organized around age, but the measurement

apparatus Definition 1 puts in age's place itself intersects with a different lineage of legal

regulation. First, it is a general demand of anti-discrimination law that cognitive testing in the

employment context be job-related, and comprehensive cognitive measurement unconnected

Working Paper | Ageless Management in the AI Era 157

to role requirements may not be justifiable as a selection procedure. Second, using health

status as an allocation criterion can stand in tension with the legal regimes prohibiting

discrimination on the basis of disability — the ADA (Americans with Disabilities Act) in the

United States; in Japan, the Act on Employment Promotion of Persons with Disabilities and

the Act for Eliminating Discrimination against Persons with Disabilities. The pathway by

which restriction of roles based on health status collides with duties to provide reasonable

accommodation and with limits on medical examinations cannot be excluded. Third, data on

cognitive characteristics and health status are of a class that may constitute special carerequired

personal information under Japanese law, subject to heightened discipline in

acquisition and use. That is, the apparatus introduced to remove discrimination by

chronological age can be assessed as unlawful or improper as trait discrimination or healthinformation

discrimination — a problem of the institutional viability of the design itself, prior

to the misuse discussed at the opening of Section 10.3. Furthermore, because years of

experience correlate strongly with chronological age, experience-based role allocation, even

where at the individual level it is reattribution away from age (Section 4.4), can appear at the

group level as a difference in treatment correlated with age — indirect discrimination. The

implementation of Ageless Management presupposes passing these legality reviews at the

design stage. The description in this section is an identification of issues, not legal advice, and

individual assessments depend on the statutes of each jurisdiction and the facts of each case

(the same reservation as at the opening of Section 7).

The handling of data carries its own dangers. In the non-employment participation on

which Ageless Management depends (outsourced engagement, advisory roles, PBL, NPO

partnership), the disciplines protecting workers' personal information and governing health

information that presuppose an employment relationship may not apply, or their application

may be unclear. The data whose collection this paper's framework demands — cognitive

characteristics, health status, load — belong by their nature to the class whose misuse carries

the heaviest consequences. Minimization of collection, limitation of retention periods, giving

substance to the person's consent (under asymmetric bargaining power, consent easily

becomes a formality), and prohibition of data linkage across organizations are the minimum

line (Section 8), but these remain voluntary standards without legal backing. The protection

vacuum Proposition 9 identified extends not only to workers' accident compensation and

remuneration but to data.

10.4 Survivorship Bias — The Re-Invisibilization of Those Who Cannot Work

This paper's theory is, by its very materials, skewed toward survivors. The evidence on

capacity maintenance in Section 2 is data from those who could keep participating in testing;

the association between work and health in Section 3, from those who could keep working;

the evidence on older workers' productivity in Section 6, from those who remained in the

factory (Börsch-Supan & Weiss 2016 corrects statistically for survivorship bias, but complete

removal is impossible). A theory standing on this skew carries the danger that, by speaking of

Working Paper | Ageless Management in the AI Era 158

the possibilities of older people who can work, it renders invisible once again those who

cannot.

One should have a sense of scale. Healthy working life expectancy at age 50 in England —

the years a person can expect to spend both healthy and in work — is estimated at about 9

years, below the years remaining to state pension age (Parker et al. 2020). And this figure has

steep gradients: 10.9 years for men against 8.3 for women, 6.8 years in the most deprived

areas, with regional differences reaching about 4.5 years. In the sample of a large Japanese

longitudinal survey as well, 66.0% of those aged 65 and over are retirees, and 10.2% have no

work experience at all (Takeuchi et al. 2024). The image this paper draws — "super-seniors

exercising experiential audit capacity into their 90s" — is legitimate as a description of

possibility, but if it circulates as a standard image, it turns into a pressure to reinterpret the

condition of the many who cannot participate because of health, caregiving, poverty, or

disability as if it were a matter of will or effort. The pathway by which the vocabulary of

Ageless Management is appropriated to justify policies such as raising the pension eligibility

age or tightening work requirements lies outside this paper's control, but being foreseeable, it

is warned against explicitly here.

Furthermore, this paper's theory has, even regarding change in cognitive function itself,

shone its light on the side that can be maintained — Gc-type abilities and metacognition. But

the changes of cognition that accompany aging do actually include decline to a level at which,

if it progresses, experiential audit capacity itself ceases to hold. For people in states of

dementia or long-term care need, this paper's prescription of participation in audit-type roles

is no answer. What Ageless Management can speak to extends only to role design within the

range where cognitive participation-capability is preserved, and to the provision of the

environment that preserves that capability (Brain Safety, cognitive engagement); dignity and

care after the capability is lost are a separate task outside this paper's theory — and one that

does not rank below it. Talk of "lifelong active engagement" that blurs this boundary runs

counter to this paper's intent.

The second boundary condition of Proposition 2 (Section 5) writes this fact into the theory

in plain terms. AI's complementation of Gf-type components holds only within the range in

which standard cognitive screening does not fall below the threshold of mild cognitive

impairment, and this model cannot complement a state in which attentional resources

themselves are exhausted — this lower bound is the acknowledgment that the

complementation model has a neurological limit point. This paper makes this boundary

explicit not to justify exclusion but to draw the theory's honest boundary. A theory that does

not write down its lower bound would, by the overclaim that AI complementation is possible

for people in every cognitive state, instead blur the distinctiveness of the real support — care,

medicine, income security — needed by those outside the boundary. At the same time, as

Proposition 2 itself states explicitly, this lower bound is a boundary on allocation to

experiential-audit roles, not a constraint on participation in the multigenerational ecosystem

in general. The possibility that a person below the threshold participates in forms other than

Working Paper | Ageless Management in the AI Era 159

audit — the provision of lived-experience problem perception, the handing down of

narrative, loose social connection — remains open outside the boundary, and its concrete

design remains a task this paper has not discharged. The rampart against survivorship bias is

not to hide the existence of the limit point. It is to make the location of the boundary explicit

in measurable form, and to condition the dignity of people on neither side of it on

participation — the distinction drawn in the preceding passage is the minimum line for that

purpose.

The breakwaters within this paper's framework are three. First, Definition 1 includes

health status and the person's own intent among the allocation criteria, and participation is a

right, not a duty. Second, Definition 8 (Brain Safety) requires a design that caps the intensity

and load of participation, and Proposition 10 requires that this protection be designed

symmetrically as a response to asymmetric bargaining power. Third, the socioeconomic

gradient of healthy life expectancy requires that the ecosystem's design object include not

only "the utilization of those who can participate" but "the distribution of participationcapability

itself" — this is also the reason the institutional analysis of Section 7 is a

component of the management theory, not an external condition to it. Even so, this paper

presents no theory of its own about the welfare of people in states where they cannot work.

Ageless Management is a theory of inclusion through work and participation, and it is no

substitute for inclusion that does not pass through participation — income security, care,

dignity without participation. Making this boundary explicit is the limit of the honesty this

paper can offer.

10.5 The Limits of the Economic Presuppositions

This paper's theory places several economic presuppositions implicitly. In domains where

they break down, implementation fails even if the propositions are theoretically correct. The

first is the cost of AI use. This paper took the fall in the marginal cost of Gf-type components

(Proposition 1) as its point of departure, but what an organization bears is not only the unit

price of inference. Systems that present output and grounds in auditable form, work

environments compliant with Brain Safety, and the cost of keeping up with model updates

remain as fixed costs. The second is the cost of task decomposition. The work of decomposing

job bundles into Gf-type and Gc-type components and re-allocating them across AI and

multiple participants itself generates costs of analysis, coordination, and contracting.

Moreover, the boundary of AI's zone of competence is hard to see in advance (the jagged

frontier of Section 4), and errors of decomposition surface only after the fact, as oversight

failures. The third is the transaction cost of running the ecosystem. Managing the multiple

contractual forms of employment, outsourced engagement, advisory roles, PBL, and NPO

partnership; measuring and re-allocating participants; and maintaining a composition that

preserves decorrelation all require administrative cost.

A large part of these costs diminishes with organizational scale, and this paper's design

theory may be tilted in favor of large organizations or consortia. This point must be accepted

Working Paper | Ageless Management in the AI Era 160

explicitly as a critique of external validity. The cost of building and maintaining, firm by firm,

the measurement and governance apparatus this paper demands — measurement and remeasurement

of cognitive characteristics, dual-track auditing, risk-weighted sampling,

unannounced in-situ verification, a layer-separated feedback infrastructure — is enormous,

and this paper cannot exclude the possibility that net benefit turns positive only in large

organizations. The response to this critique rests on two pathways. The first is the platform

pathway presented as a possibility in Section 6.5. If the administrative and measurement

infrastructure is standardized as common infrastructure, the marginal cost per firm can fall,

and if this holds, the scale threshold of net benefit comes down — though this is the

presentation of a possibility, not an assertion that it holds. The second is the net-benefit

requirement of Hypothesis H2. Because H2 records time to decision and verification effort as

cost variables and makes a positive net benefit after their deduction a requirement for

support, if administrative overhead eats up the benefits, Proposition 6 is by design not

supported — that is, the testing of this external-validity critique is built in not as an objection

from outside the theory but as an adjudication criterion inside this paper's verification plan.

Implementability in small, medium-sized, and regional organizations could not be examined

in this paper, and it is an important observation item for the pilot and the case studies in the

roadmap of Section 9. Also, regarding the time lag until the marginal value shift of

Proposition 1 is actually reflected in the labor market's reward structure, and the possibility

that the reward premium of Gc-type tasks is concentrated in particular occupations and

strata, this paper has nothing beyond theoretical prediction. In domains where the economic

presuppositions are not met, the multigenerational ecosystem will remain a cost center, and

pressure to revert to the "accommodation and compensation" paradigm will operate. That is

not a refutation of this paper's thesis but a demarcation of its scope — yet narrowness of

scope directly diminishes the value of a theory.

Finally, the temporal reach (Temporal Robustness) of this paper's theory should be made

clear. As declared in Section 1.4, what this paper's propositions are anchored to is not the

performance profile of the current generation of generative AI — the transient performance

gap that "today's LLMs are poor at contextual verification." The anchor point is the structure

of verification independence formalized by WP8 (Kadowaki 2026h, Propositions 10 and 11).

Namely: for any intelligent agent whatsoever, a verifier whose errors correlate with its own

training distribution cannot supply itself with statistical independence in detecting its own

systematic blind spots. The independence of verification is not supplied by improvements in

capability — however capable the verifier, so long as it shares an error distribution with the

generator, its misses remain correlated with the generator's misses. Therefore, even in a

world where future AI has made self-correction and self-verification highly sophisticated, (a)

the detection of systematic blind spots deriving from the training distribution will still

require decorrelated verifiers, and (b) the sources of their supply will be both humans with

heterogeneous experience and models of different lineages, so that this paper's framework

Working Paper | Ageless Management in the AI Era 161

persists as the allocation problem between the two. The audit-certification pathway shown in

the commentary on Proposition 1 (Section 5) is likewise anchored to this structure.

This structural anchoring, however, does not guarantee the size of the human contribution

in that allocation. Stated honestly, this paper cannot exclude the possibility that the advance

of machine-side decorrelation — cross-verification by models of different architectures and

developers (Proposition 5's model condition) — and the acceleration of verification-skill

acquisition in AI-use environments will progressively shrink the residual niche of humansupplied

decorrelation. Because this shrinkage would appear not as a one-stroke refutation of

the theory but as a gradual narrowing of its scope, it is all the harder to detect. What this

paper designates as its detector is the second sentence of Proposition 3's refutation condition

— longitudinal evidence that the acquisition of verification capacity accelerates in AI-use

environments and substantially substitutes for the effect of accumulated experience

(longitudinal compression). This paper's propositions should therefore be read as dynamic

claims to be re-tested at each point of the technology's development; the results of H3 must

always be annotated with the model generation at the time of data collection; and the

longitudinal measurement in the roadmap of Section 9 must include, among its objects of

observation, the very speed at which this residual niche shrinks.

10.6 Disclosure of Conflict of Interest and Classification of Self-Citations

Finally, this paper discloses the conditions of its own production. The author of this paper is

affiliated with VURA Capital Innovation Holdings, Inc. The company is a business operating in

the domains of longevity and brain capital, and it stands to gain business benefit from the

spread of the ideas of Ageless Management and brain capital management. That is, this paper

is written not by an observer with no stake in the theory's consequences, but by an interested

party who may benefit from the theory's diffusion. Readers should read this paper with this

structure discounted, and this paper's disciplines — the attachment of refutation conditions

to every proposition, the explicit grading of evidence, the explicit presentation of headwind

data (Section 4.5), and this section — are designed on the premise of that discount. It is for the

same reason that the roadmap of Section 9 makes it a requirement that verification from the

third stage onward, and the interpretation of results, be open to execution and replication by

independent researchers. A configuration in which the proposer of a theory monopolizes its

verification is not permissible under this conflict of interest.

The classification of self-citations is also made explicit. This paper has cited eight papers of

the VURA Working Paper Series, and their positioning divides into three types. The first is

inheritance. The inversion of valuation in Future Value Theory (Kadowaki 2026a); the K × u

framework and the state-versus-intervention distinction of Brain Capital Management

(Kadowaki 2026e); and the oversight value = independence × detection probability, error

decorrelation, the solvency condition, and the non-delegable residual of the Human on the

Loop argument (Kadowaki 2026h) were used by this paper as presupposed frameworks. The

second is reference. Enterprise redefinition and its observation (Kadowaki 2026b; 2026c),

Working Paper | Ageless Management in the AI Era 162

purpose-description-based role design (Kadowaki 2026d), the self-defined society (Kadowaki

2026f), and redefinition capitalism (Kadowaki 2026g) were cited solely to indicate position

within the series. The third is what newly becomes an object of verification in this paper. The

relation between BCM's "protected unassisted practice" and Hypothesis H3, and the extension

of HOTL's decorrelation condition to the generational axis (Proposition 5), are new claims of

this paper, and they are untested.

The point of this classification lies in the following discipline. The inherited frameworks

themselves are all unrefereed working papers, and mutual citation within the series is no

substitute for external evidence. This paper does not make claims for which no independent

supporting evidence outside the series exists appear established through a chain of selfcitations.

The strength of this paper's claims is supported not by the internal coherence of the

series but solely by the independent empirical literature cited in Sections 2 through 4 and by

the executability of the verification plan of Section 9. And since the core of that verification

plan (H3) has not yet been executed, what this paper has reached is a proposal put into

refutable form — this sentence is the conclusion of this section, and the premise of the

conclusion of the next.

11. Conclusion

This paper has presented the theory of a management regime — Ageless Management — that

removes the variable of chronological age from decisions on role allocation, evaluation,

participation, and exit, and that dynamically allocates roles, within an ecosystem not limited

to the boundary of employment, on the basis of measured cognitive characteristics,

accumulated domain experience, health status, and the person's own intent. The skeleton of

the theory consists of three operations. First, the operation of reattributing the oversight

value of older workers from age to "the interaction of long-term domain experience × Gc ×

metacognition" — that is, to experiential audit capacity. Second, the operation of formalizing

generational heterogeneity as an organizational source of error decorrelation in AI oversight.

Third, the operation of converting the three social problems of population aging, constrained

youth participation, and the exclusion of socially marginalized groups into untapped sources

of brain capital (K × u), conditional on AI-mediated complementarity and on Brain Safety

operated in an ex-ante verifiable form. These were cast into eight definitions, twelve

propositions each carrying a refutation condition, and three testable hypotheses (Section 9).

This paper's position within the series should also be confirmed. This paper extends Brain

Capital Management (Kadowaki 2026e), which treated brain capital as an asset of a single

firm's employees, vertically into a multigenerational ecosystem not limited to the boundary of

employment, and at the same time bridges the theory of future value (Kadowaki 2026a) and

the theory of oversight structure (Kadowaki 2026h) along the axis of age. WP8, which

demonstrated the failure of oversight, and this paper, which asks about the supply side of

oversight, divide a single question between demand side and supply side — who bears the

Working Paper | Ageless Management in the AI Era 163

residual that cannot be delegated to AI, under what protections, and at what cost. This paper's

answer — experiential audit capacity and generational decorrelation — is not the only

answer to this question, but the first candidate arranged in a testable form.

The meta-structure in the lineage of theory declared in the introduction (Section 1.4) should

also be reconfirmed here. The main axis of this paper is a theory of firms' competitive

advantage. Experiential audit capacity, as a path-dependent product that can be accumulated

only over time through long-term domain experience, is a candidate for a scarce and hard-toimitate

oversight resource (the resource-based view — Barney 1991), and the configurations

that organize it — the oversight portfolio that designs error decorrelation, and dynamic role

allocation — sit at the level of dynamic capabilities that integrate and reconfigure resources

under change (Teece, Pisano & Shuen 1997). By contrast, the institutional argument (Section

7) and the health argument (Section 3) are not parallel claims of this paper but the

institutional context and complementary assets that make the acquisition of competitive

advantage possible. The protective institutions for non-employment participation are the

institutional context that makes connection to oversight resources possible, and Brain Safety

and brain health are the complementary assets that prevent the depreciation of the resource

and sustain its utilization. That this paper — a theory of management — discussed

institutions and health at length is due to this complementary structure; conversely, when the

institutional context and complementary assets are absent, experiential audit capacity

remains unexpressed as a source of competitive advantage.

What ran through this paper's argument is a discipline of conditionality. This paper did not

claim that "humans are good at oversight"; it claimed only that, given that a residual of

oversight not delegable to AI exists, the supply side of oversight — a scarce and consumable

resource — must be interrogated. It did not claim that "working makes you healthy"; it

presented a mediation hypothesis (H1) conditional on the quality of work. It did not claim

that "multigenerational means more value"; on top of the empirical baseline of a near-zero

average effect, it placed a conditional proposition with AI-mediated complementarity as the

moderator. Nor did it unconditionally claim that "different generations mean fewer blind

spots"; it conditioned the supply of decorrelation on three conditions (Proposition 5) —

unmediated audit channels, flattened power gradients, and cross-validation across multiple

foundation models. On the audit value of experience itself, it drew the boundaries of

impairment under confirmation-congruent errors and the half-life of knowledge (Proposition

4), and the neurological boundary of a cognitive-screening floor (Proposition 2). And it offered

the keystone of the theory itself — the existence of experiential audit capacity — as the object

of the highest-priority, lowest-cost experiment (H3). Most of the propositions are untested,

and this paper is not a report of established facts. In addition, this paper's theory contains

within it pathways of misuse — a new stereotyping of seniors as auditors, the conversion of

measurement into an apparatus of selection, and the re-invisibilization of those who cannot

work — and Section 10 identified them by name.

Working Paper | Ageless Management in the AI Era 164

Even so, the reason this paper presents this theory lies in one point inherited from the

original draft. Ageless Management is not an improvement of measures for making older

people work longer. What it aims at is to dismantle the institutions and notions by which

human beings are uniformly severed from the front line of social participation by the single

variable of chronological age, and to prepare a structure in which human beings carrying

cognitive change can, together with the complementary apparatus of AI, participate in the

creation of value across the lifespan — a state in which human dignity and organizational

value creation are compatible (Human Flourishing). The linear three-stage model of life —

education, work, and retirement — is an institutional product of an era when average

lifespans were shorter than they are now, and the extension of lifespans is already

invalidating this model's premises. The dynamic role allocation this paper has presented is an

attempt to concretize, at the level of management, one candidate for what should come after

this model — a configuration in which roles are continually reallocated across the lifespan

not as a function of age but as a function of measured characteristics, experience, and intent.

Longevity is experienced as a burden when it lacks a structure to support it, and as a

resource when it gains one. Which it becomes is a matter not of biology but of design, and

among its design variables, as this paper has argued, not a few lie within the reach of

management and institutions. And that design can be legitimate only inside the boundary

line of Section 10 — that the dignity of people who are in a state of being unable to work is

not conditioned on participation.

The implications of this paper are stated in minimal form for each type of reader. What this

paper offers to management practice is not a collection of measures whose success is

guaranteed but an inventory of design variables — the decomposition of job bundles into Gftype

and Gc-type components, the measurement of experiential audit capacity and its

connection to roles, generationally heterogeneous supervisor configurations (including the

three conditions of Proposition 5 — raw-audit slots, flattened power gradients, and crossvalidation

across multiple foundation models), state monitoring of K × u, and Brain Safety as

bidirectional protection against both overload and exploitation. A partial adoption lacking

any one of these — above all, utilization without protection — is not recognized as an

implementation within this paper's framework. The implication for policy, as argued in

Section 7, is that institutional development filling the protection gap for non-employment

participation (Proposition 9) is a precondition of Ageless Management. Flexibilization without

protection is a passage to exploitation, and this paper refuses to be cited as an argument for

flexibilization. For research, the map of verification in Section 9 is, as it stands, an inventory

of research problems. H3 in particular can directly test the core of the theory at low cost, and

whichever way the result falls — as the first direct evidence of the non-compressibility of

experience, or as a refutation of the core of this paper's theory — it constitutes a contribution.

Within the scope of this paper's search, no study was found that operationalized age or

tenure and measured performance in AI oversight (Section 4), and this gap is worth filling

beyond the question of whether this paper's theory is right.

Working Paper | Ageless Management in the AI Era 165

Yet the sentences that speak of this prospect must be chosen carefully. What this paper has

shown is not a realized achievement but a structure of possibility. That complementation of

Gf-type components can remove the cognitive bottleneck, that experiential audit capacity can

become a scarce input to oversight, that generational decorrelation can widen an

organization's detection set, that the three social problems can be converted into sources of

brain capital — all of these are written in the modality of "can," and their realization is not

automatic. What is required is the execution of the verification program of Section 9, the

revision or abandonment of the propositions according to its results, and the institutional

design that fills the protection gaps identified by Propositions 9 and 10. Absent verification,

this paper remains a system of aspirations; absent institutional design, implementation can

degenerate into exploitation. What this paper has shown is a structure of possibility, and its

realization is conditional on the verification of the propositions and on institutional design.

Appendix A Experimental Protocol for Hypothesis H3

This appendix details the experimental protocol for Hypothesis H3 (the direct test of the noncompressibility

of experience) presented in Section 9. The design skeleton of H3 is as follows.

A 2×2 factorial design [experience level: long-term domain-experienced practitioners vs.

junior participants] × [AI assistance: present vs. absent]. The task is the verification of AIgenerated

business proposals and documents in which misinformation and contextual risks

have been embedded. So that the embedded errors do not depend on the experimenters'

preconceptions, the tasks include items blind-generated from an empirical failure-case

dataset of accidents, scandals, and business failures that actually occurred (A.2), and part of

the errors are constructed in both a form congruent with auditors' industry conventional

wisdom (confirmation-congruent) and a form contrary to that wisdom (deviating), so as to

simultaneously test boundary condition (i) of Proposition 4. The task documents are prepared

in two versions — a high-fluency version and a low-fluency version, identical in content and

differing only in stylistic polish — and fluency is controlled as a factor or covariate, so as to

simultaneously test boundary condition (iii) of Proposition 4 (A.2). The dependent variables

are four: (i) the detection rate of surface errors, (ii) the detection rate of contextual and

practical risks, (iii) the quality of correction proposals (blind evaluation), and (iv) the overrejection

rate — the rate at which embedded "correct but counter-conventional innovative

proposals" were erroneously rejected. The prediction is that for (i) the experience gap shrinks

under AI assistance (consistent with prior research), whereas for (ii) and (iii) the main effect

of experience persists and does not shrink even under AI assistance (the direct test of

Proposition 3). No prediction is placed on (iv); it serves as an exploratory indicator of the

boundary at which experience turns into over-auditing. The following records, in order,

participant requirements (A.1), task construction (A.2), procedure (A.3), evaluator blinding (A.

4), and the testing plan (A.5), and appends at the end the team-assignment procedure

(stratified randomization) for Hypothesis H2 (A.6). Implementation presupposes

Working Paper | Ageless Management in the AI Era 166

preregistration (registration of hypotheses, exclusion criteria, and the analysis plan) and

research ethics review.

A.1 Participant Requirements and Operational Definitions

Experience level cannot be randomly assigned. The experience factor of this experiment is

therefore a measured quasi-experimental factor, and only the AI-assistance factor is

randomized. With this asymmetry made explicit, experience level is operationally defined as

follows.

Long-term domain-experienced practitioners are those with 15 or more cumulative years

of practical experience in the task domain (the industry to which the business proposals

described below belong), and who have been away from practice in that domain for no more

than 5 years. Junior participants are those with fewer than 5 cumulative years of practical

experience in that domain. Years of experience are measured by self-report combined with

employment-history verification (documentary confirmation of periods of tenure or

certification by the affiliated organization). Those with 5 or more but fewer than 15

cumulative years are excluded from the main analysis (including the boundary-ambiguous

intermediate stratum would blunt the contrast of the factor; the exclusion criterion is stated

explicitly in the preregistration).

The eligibility criterion "away from practice for no more than 5 years" imposes an

important limit on generalizability. This criterion excludes from the experiment the superseniors

long after exit — the stratum well beyond 5 years away from practice — whom this

paper's management model envisions as the principal source of supply. Whatever the results

of H3, therefore, they cannot be extrapolated to that stratum. Even if H3 is supported, what it

shows is non-compressibility among practitioners who are active or recently exited; for the

long-exited stratum, estimating the depreciation function of detection capacity with years

since leaving practice as a continuous variable remains an independent verification task

(Section 9.4).

At the same time, even within this eligibility range (0–5 years after leaving practice), years

elapsed since exit vary and can confound the experience factor through the depreciation of

unexercised abilities (Section 4.3). Accordingly, years elapsed since leaving practice (set to 0

for those currently working) are recorded in months for all participants and controlled as a

covariate (A.5). This is the H3-side response to the same confound — the conflation of the

effects of work and experience with the accumulated natural depreciation over the period

since exit — for which Hypothesis H1 demands strict control of "Time Since Retirement."

Chronological age is not used as an assignment condition. This is a design choice

corresponding to this paper's core distinction (Section 4.4). Age is recorded and used as a

covariate and in auxiliary analyses (A.5 below). Furthermore, to mitigate the confounding of

experience and age to the extent possible, recruitment deliberately strives to include older

persons with little experience (career changers and returners) and younger persons with long

Working Paper | Ageless Management in the AI Era 167

experience (early specializers). If filling both cells proves difficult, this is reported and stated

explicitly as a limit on the separability of experience and age. In addition, for all participants,

a proxy measure of crystallized intelligence (Gc) (vocabulary and knowledge tests), a selfreport

scale of metacognition, and experience using generative AI (frequency and purposes)

are measured in advance and used as covariates. AI-use experience can correlate with

experience level (the usage gap of Section 4.5) and is therefore especially important as a

control variable.

A.2 Task Construction and Error Embedding

The task materials are business proposals and business documents created with generative AI

(market analyses, business plans, customer proposals, and the like), formatted to match the

practice of the task domain. In each document, errors and risks known to the experimenters

are embedded by type. Four types are embedded.

First, factual errors. These are errors verifiable by checking against information external to

the document, such as misstated figures, references to nonexistent sources, and confusions of

proper nouns. Second, contextual risks. These are contents that are coherent on the face of

the document but that cause problems in light of the practical context of the domain —

transaction terms contrary to industry practice, premises that do not hold for the customer

segment in question, measures whose same-pattern failures are known from the past, and the

like. Third, ethical risks. These are problems requiring value judgment, such as potential

conflicts with laws and norms, unjust burdens on stakeholders, and overlooked conflicts of

interest. Fourth, anachronisms. These are premises that were once valid but no longer hold —

reliance on abolished institutions, designs premised on the specifications of a previous

generation of technology, outdated images of changed consumer behavior. The detection of

anachronisms is a window for observing in which direction the experience of different

historical environments (the foundation of the generational decorrelation of Definition 5)

operates; because old experience may aid detection and, conversely, old knowledge may

generate misplaced confidence about the present, the direction is measured without being

fixed in advance.

Blind Generation of Embedded Errors — Eliminating Experimenter Bias

The creation of embedded errors is subjected to procedures that eliminate the experimenters'

preconception bias. As organized in Section 4.2, contextual information can systematically

distort the judgments of experts themselves (Dror, Charlton & Péron 2006). This finding

applies also to the experimenters who create the tasks: errors devised at the experimenters'

desks would be biased toward types congruent with the experimenters' own views of the

industry and of failure, and that bias could work either for or against the experienced

participants. The response consists of two steps.

First, part of the errors are incorporated as tasks blind-generated from an empirical failurecase

dataset of accidents, scandals, and business failures that actually occurred — accident

Working Paper | Ageless Management in the AI Era 168

investigation reports, administrative sanctions and recall cases, public records of

bankruptcies and withdrawals, and the like. That is, a generation team independent of the

experimenters converts the structure of the failure cases (what was believed to be sound, and

what in fact broke down) into errors and risks within the task documents. The generation

process and the conduct and analysis of the experiment are separated in personnel, and the

experimenters and evaluators are not informed which error derives from which case until

analysis is complete. Second, the full list of errors, the generation procedure, and the

correspondence table to the source cases are preregistered with timestamps before the

experiment begins (a sealed registration kept nonpublic until analysis is complete). This

makes ex-post substitution of errors or convenient reinterpretation procedurally impossible.

Confirmation-Congruent and Deviating Types — Simultaneous Test of Boundary

Condition (i) of Proposition 4

Furthermore, part of the embedded errors — above all the contextual risks and

anachronisms — are constructed in two classes according to their congruence with auditors'

industry conventional wisdom. The confirmation-congruent type consists of plausible errors

in the direction of industry conventional wisdom or the success experiences of the auditors'

generation (contents that follow the conventional wisdom but break down in the given

context); the deviating type consists of errors in the direction contrary to that wisdom. The

validity of the classification is confirmed by having a panel of practitioners in the domain

who do not participate in the experiment independently rate, for each statement, "whether it

is natural in light of industry conventional wisdom," and by its meeting a preregistered

agreement criterion.

This classification is the operation for the simultaneous test of boundary condition (i) of

Proposition 4 — that when AI output is congruent with the auditor's own past success

experiences and industry conventional wisdom, experience can, through the synergy of

confirmation bias and automation bias, actually impair detection. If boundary condition (i) is

real, the detection-rate advantage of the experienced will be observed for deviating errors

and will shrink, vanish, or reverse for confirmation-congruent errors. Conversely, if even for

confirmation-congruent errors the detection rate of the experienced does not fall below that

of the inexperienced, boundary condition (i) is empty and the body of Proposition 4 is

supported in a stronger form (the asymmetric structure of Proposition 4's refutation

condition). This experience × congruence interaction test is preregistered as a secondary

analysis (A.5).

Embedding Innovative Proposals — Measuring the Over-Rejection Rate

In addition to errors, a small number of "correct but counter-conventional innovative

proposals" are embedded in the task documents. These are non-error items for measuring the

boundary at which experience goes beyond the detection of errors and turns into the

rejection of sound deviation (over-auditing), and they are the measurement target of

Working Paper | Ageless Management in the AI Era 169

dependent variable (iv), the over-rejection rate — the rate at which these proposals were

erroneously rejected as problems.

The certification of "innovative but correct" is made, by prior agreement before the

experiment begins, by an expert panel independent of both the experimenters and the

outcome evaluators. The panel certifies, under a preregistered agreement criterion

(unanimity or a prespecified supermajority), that each candidate proposal both (a) deviates

from industry conventional wisdom and (b) is nonetheless sound from the standpoints of

fact, logic, and practice, and candidates that fail the criterion are not used. To the extent

possible, the proposals are constructed on the basis of real instances that were initially

dismissed as contrary to conventional wisdom but whose validity was later established, and

the corresponding provenance table is also included in the sealed registration. No theoretical

prediction is placed on the over-rejection rate; it is treated as an exploratory indicator

(Section 9, Hypothesis H3).

The Fluency Manipulation — Simultaneous Test of Boundary Condition (iii) of

Proposition 4

The task documents are prepared in two versions, a high-fluency version and a low-fluency

version, identical in content and differing only in stylistic polish. The two versions keep the

propositional content — including the embedded errors and the innovative proposals (facts,

figures, logical structure, and the location of risks) — completely identical, and the difference

is confined to the level of style. The high-fluency version is a low-processing-resistance style:

grammatically well-formed, with plain vocabulary and syntax and smooth paragraph

transitions. The low-fluency version is a style that, while preserving the propositional

content, raises processing resistance through more complex syntax, uneven paragraph

construction, and redundant or stilted expression. The versions are produced by combining

generative-AI style transformation with human copyediting, and the identity of the two

versions' propositional content is confirmed by an independent checker through collation

against a correspondence table.

As a manipulation check, a panel of raters who do not participate in the experiment rates

the subjective fluency (readability and ease of processing) of each document, confirming that

the high-fluency version exceeds the low-fluency version by at least a preregistered criterion.

Objective readability indicators such as sentence length and syntactic complexity are also

reported, and document pairs that fail the manipulation check are replaced. In assignment,

fluency is treated as a factor (between-participants or within-participants) or as a covariate,

and combinations of documents and versions are balanced across conditions. The theoretical

grounding of this manipulation lies in the cognitive-psychology finding that processing

fluency inflates judgments of truth (Reber & Schwarz 1999; Alter & Oppenheimer 2009). If

boundary condition (iii) of Proposition 4 — that high-fluency output can, through the effect of

processing fluency, raise the threshold of the auditor's cognitive sense of unease and lower

detection performance — is real, detection rates will fall in the high-fluency version, with the

Working Paper | Ageless Management in the AI Era 170

fall predicted above all in the detection of contextual risks. Conversely, a result in which

detection performance does not fall even when fluency is manipulated empties boundary

condition (iii) and leaves the body of Proposition 4 standing in a stronger form (Proposition

4's refutation condition). The corresponding tests are preregistered as secondary analyses (A.

5).

In embedding, the number of errors of each type and the ratio of confirmation-congruent to

deviating errors are balanced across documents, and placement is randomized so that errors

cannot be detected from cues of position or formatting. Sufficient sound passages containing

no errors are also secured, so that false alarms — participants flagging correct passages as

errors — can be measured. This separates the detection rate (hit rate) from the false-alarm

rate and enables analysis within the framework of signal detection theory (SDT) — the

separate estimation of sensitivity d′ and response criterion c (A.5). The validity of the

documents and the embedded errors is confirmed by prior review by a panel of practitioners

in the domain who do not participate in the experiment, and embedded errors not detected

in that review are replaced.

Table A1 Correspondence between embedded error types and the abilities required for

detection (theoretical predictions)

Error type Definition

Ability required for detection

(theoretical correspondence)

Predicted compression

under AI assistance

Factual error Errors verifiable by

checking against

information external to

the document (figures,

sources, proper nouns)

Procedural skills of collation and

search. Relatively large Gf-type

component

Compressed (AI

assistance takes over

search and collation, and

the experience gap

shrinks)

Contextual risk Contents coherent

within the document

but problematic in light

of the practical context

Interaction of long-term domain

experience × Gc-type abilities

(the formation mechanism that

Proposition 4 asserts for the

detection capacity of Definition

4)

Not compressed (the

main effect of

experience persists even

under assistance) — the

direct test point of

Proposition 3

Ethical risk Potential conflicts with

laws and norms, unjust

allocation of interests,

overlooked conflicts of

interest

Gc-type abilities, metacognition,

and value judgment. Experience

contributes as case knowledge

Not compressed (though

predicted to be less

experience-specific than

contextual risk)

Anachronism Reliance on premises

once valid but no longer

holding

Experience of changing

historical environments (the

foundation of Definition 5).

However, overconfidence in old

knowledge may operate in

reverse

Direction not specified in

advance (exploratory

measurement) — the

observation window for

generational

decorrelation

Proposals deviating

from industry

conventional wisdom

The ability to judge the validity

of content beyond conformity to

conventional wisdom. The

No prediction placed

(exploratory

measurement) — the

Working Paper | Ageless Management in the AI Era 171

Innovative

proposal (nonerror

item)

but certified as sound

by an independent

expert panel by prior

agreement

observation window for the

boundary at which experience

turns into over-auditing

measurement target of

dependent variable (iv),

the over-rejection rate

Note: The correspondence in "ability required for detection" is this paper's theoretical prediction and is itself the object

of the experimental test. Types for which the prediction fails are reported as information delimiting the scope of

application of Definition 4. Contextual-risk and anachronism errors are constructed in two classes: confirmationcongruent

(errors in line with the auditors' industry conventional wisdom) and deviating (errors contrary to that

wisdom) (see the main text). In addition, all task documents are prepared in two versions, a high-fluency version and a

low-fluency version, identical in content and differing only in stylistic polish (see "The Fluency Manipulation" in the

main text).

A.3 Procedure

After determination of experience level, participants are randomly assigned to either the AIassisted

condition or the unassisted condition (stratified randomization: stratified by

experience level, age band, and AI-use experience). Participants in the assisted condition may

freely use generative-AI tools during the verification work, and their usage logs (prompts and

responses) are recorded. Participants in the unassisted condition work without AI tools, with

ordinary reference materials (whether search is included depends on the design of factualerror

detection and is fixed in the preregistration).

The task consists of multiple documents, and the order of document presentation is

randomized. The instruction to participants is: "This document may contain errors or

problems. Point out all problems you find, and attach a correction proposal to each." The

types and number of embedded errors are not disclosed. Working time is recorded with an

upper limit, and time itself is treated as a dependent variable (a proxy for the effort invested

in verification). After all tasks are completed, confidence ratings (confidence for each flag) are

collected, and in the AI-assisted condition a self-assessment of dependence on the assistance.

The correspondence between confidence and correctness serves as an indicator of

metacognitive calibration and is used in examining the interaction term (metacognition) of

Proposition 4.

A.4 Evaluator Blinding

The detection rates of dependent variables (i) and (ii) can be scored mechanically because the

embedded errors are known (only the matching of flags to embedded locations is judged,

independently, by two blinded judges, with disagreements resolved by discussion). The overrejection

rate of dependent variable (iv) can likewise be scored mechanically by whether a

"problem" flag was placed on the location of an embedded innovative proposal (the matching

procedure is shared with (i) and (ii)). The quality of correction proposals, dependent variable

(iii), is assessed by blinded external evaluation. The evaluators are external experts with

practical experience in the domain, and they are presented only with the proposal texts, with

the participants' experience level, age, AI-assistance condition, and names concealed. To

prevent conditions from being inferred from style and the like, the proposal texts are

Working Paper | Ageless Management in the AI Era 172

presented after standardization of presentation (correction of typographical errors and

unification of formatting). Each proposal is rated independently by two or more evaluators,

and inter-rater reliability (intraclass correlation coefficient) is reported. If reliability falls

below the preregistered criterion, the rating procedure is revised and re-rating is conducted.

The evaluation axes are three — accuracy of problem apprehension, feasibility of the

correction, and attention to side effects — and the definition and rating scale of each axis are

stated explicitly in the preregistration.

A.5 Testing Plan

The main analysis conducts, for each of dependent variables (i), (ii), and (iii), an analysis of

variance (ANOVA) of experience level (2) × AI assistance (2), testing main effects and the

interaction. The core of the hypothesis is the contrast of interactions. The predictions are as

follows. For the detection rate of surface factual errors (i), the experience × AI-assistance

interaction is significant and the experience gap shrinks under AI assistance (consistent with

the compression evidence of Section 4.1). For the detection rate of contextual and practical

risks (ii) and the quality of correction proposals (iii), the main effect of experience persists

and no interaction-driven shrinkage is observed — that is, the "absence or smallness of the

interaction" in (ii) and (iii) constitutes the supporting evidence for Proposition 3. Because a

claim of an absent interaction depends on statistical power, mere non-significance is not

relied on; an equivalence test against a preregistered smallest effect size of interest (or a

report of estimation precision) is used in combination. In addition, the analysis plan states

explicitly the application of signal detection theory (SDT). An analysis relying on the detection

rate (hit rate) alone cannot, in principle, distinguish differences in true sensitivity from

differences in response bias — a shift of the judgment criterion toward suspecting everything.

The apparently higher detection rate of the experienced may reflect higher discriminability,

or merely a lower threshold for flagging. Therefore, an SDT analysis incorporating the falsealarm

rate estimates sensitivity d′ and response criterion c separately (Macmillan & Creelman

2005; Hautus, Macmillan & Creelman 2022). Whether the advantage of the experienced is due

to detection sensitivity (d′) or to flagging assertiveness (a low c) is discriminated by this

separate estimation.

Dependent variable (iv), the over-rejection rate, is not subjected to hypothesis testing; it is

an exploratory analysis descriptively reporting the rate by condition (experience level × AI

assistance) with confidence intervals. The over-rejection rate is conceptually a kind of false

alarm, but unlike false alarms on sound passages in general, it is confined to rejections of

non-error items defined on the content dimension of deviation from conventional wisdom.

The two are reported separately, so that the general strictness of the experienced participants'

response criterion can be distinguished from selective rejection of deviation from

conventional wisdom. Within the SDT framework, dependent variable (iv), the over-rejection

rate, is not an independent phenomenon but a manifestation of the response criterion c

shifting toward the conservative (suspecting) side, and it is analyzed in the same framework

Working Paper | Ageless Management in the AI Era 173

as the separate estimation of sensitivity and criterion for (i) and (ii). That is, the general bias

of the criterion estimated from false alarms on sound passages in general and the increment

of selective rejection on deviation items are contrasted within the same model. This

increment is the measured quantity of the boundary at which experience turns into overauditing.

As the secondary analysis corresponding to boundary condition (i) of Proposition 4, the

congruence classification of errors (confirmation-congruent / deviating) is added as a withinparticipants

factor, and the interaction of experience level × congruence (and the three-way

interaction adding AI assistance) is tested. The contrast of interest is the difference in

detection rates between the experienced and the inexperienced for confirmation-congruent

errors. That this difference is significantly smaller than that for deviating errors (including

vanishing or reversal) is consistent with boundary condition (i); that the detection rate of the

experienced does not fall below that of the inexperienced even for confirmation-congruent

errors empties boundary condition (i) (Proposition 4's refutation condition). This contrast and

its adjudication criteria are stated explicitly in the preregistration.

As the secondary analysis corresponding to boundary condition (iii) of Proposition 4, the

fluency of the task documents (high-fluency / low-fluency) is added as a factor (or, in designs

treating it as a covariate, its coefficient), and the main effect of fluency and the interaction of

experience level × fluency (and the three-way interaction adding AI assistance) are tested.

Within the SDT framework, whether the effect of fluency appears as a decline in sensitivity d′

or as a conservative shift of the response criterion c — a rise in the threshold of suspicion —

is reported separately. The mechanism assumed by boundary condition (iii) (a rise in the

threshold of cognitive unease) is predicted to appear primarily as a movement of c, but this

prediction is itself an object of the test. A result in which detection performance does not

decline even when fluency is manipulated empties boundary condition (iii) (Proposition 4's

refutation condition). The adjudication criteria are stated explicitly in the preregistration.

As covariate analysis, an analysis of covariance is conducted entering the Gc proxy

measure, the metacognition scale, AI-use experience, and years elapsed since leaving practice

(A.1). In the auxiliary analysis corresponding to Proposition 4, years of experience and

chronological age are entered into the same model, and the partial effect of chronological age

after controlling for years of experience is estimated. The prediction of Proposition 4 is that

this partial effect is indistinguishable from zero (and that the effect of years of experience

persists). Conversely, if chronological age independently predicts detection performance even

after controlling for experience, this paper's reattribution thesis (reattribution from age to

experience) is refuted. Because this analysis depends on the success of the experience–age

confound mitigation described in A.1, the correlation between the two variables and the

variance inflation factor are always reported.

The required sample size is fixed by a power analysis based on the smallest effect size of

interest (set in advance from the range of effect sizes in prior compression studies and from

Working Paper | Ageless Management in the AI Era 174

practically meaningful detection-rate differences) and on the power requirements of the

interaction test. Detecting an interaction typically requires a larger sample than detecting a

main effect, and this point is treated explicitly in the power analysis. This appendix does not

provisionally fix a specific sample size. The number is fixed at preregistration as the result of

the power analysis and included in the registered content.

Finally, the limitations of this protocol are recorded. First, the experience factor is a quasiexperimental

factor, and differences between the experienced and junior participants may be

contaminated by selection and cohort factors other than experience. Covariate control

removes this only partially. Second, the embedded errors are errors known to the

experimenter side. Blind generation from the empirical failure-case dataset and sealed

registration (A.2) mitigate the bias of error selection toward the experimenters'

preconceptions, but the detection of the errors most dangerous in practice — "errors no one

had anticipated," never once manifested in the past — still cannot be measured. Third, results

from implementation in a single domain do not immediately generalize to other domains,

and replication with changed domains is necessary. Fourth, because this design is a crosssectional

2×2 factorial design, it can detect only compression by contemporaneous assistance,

and it cannot, in principle, detect longitudinal acquisition acceleration — the pathway in

which the acquisition of verification capacity accelerates under an AI-use environment and

effectively substitutes for the effect of accumulated experience (the second clause of

Proposition 3's refutation condition). Testing this pathway requires separate longitudinal

measurement (Section 9.4). These limitations are stated alongside the results report as the

frame for interpreting the experimental results.

A.6 Supplementary Note — Team Assignment for Hypothesis H2 (Stratified

Randomization)

The main object of this appendix is H3, but the skeleton of the assignment procedure for the

team-level randomization of Hypothesis H2 (the multigenerational team-outcome hypothesis)

is also appended here (Section 9.1). What H2 seeks to identify is the effect of heterogeneity in

age and experience; if the composition of other demographic attributes — gender, ethnicity,

cultural background, and the like — is skewed across teams, the mixed/homogeneous contrast

is confounded with these diversity dimensions, and the observed effect can no longer be

attributed to the axis of age and experience. Participants are therefore stratified by attributes

such as gender, ethnicity, and cultural background, and within each stratum randomly

assigned to the team-composition condition (mixed / homogeneous) and the AI-use condition

(with / without), thereby balancing the composition of these attributes across teams and

conditions (stratified randomization). For attributes for which sample-size constraints leave

strata underfilled and full stratification infeasible, the attribute in question is controlled at

the analysis stage as a covariate. In either case, post-assignment attribute balance

(standardized differences across conditions) is reported. The list and definitions of the

attributes used for stratification, the criteria for switching from stratification to covariate

Working Paper | Ageless Management in the AI Era 175

control, and the reporting format for balance indicators are stated explicitly in the

preregistration.

Appendix B Detailed Table of the Institutional Comparison

Table A2 lists the details of the institutions compared in Section 7 (Table 6) — official statute

names, statute numbers, effective dates, key points of content, and sources. Entries are

limited to what was checked against primary statutes and administrative materials or

materials of public research institutions. This table is a description of institutions, not legal

advice, and details of application are governed by the latest content of each statute and

circular.

Table A2 Details of institutions related to older-age and youth work (by jurisdiction, as of August

2026)

Jurisdiction

and

institution

Official statute

name and basis

Effective

date, etc. Key points of content Sources

Japan — Act

on

Stabilization

of

Employment

of Elderly

Persons

Act on Stabilization

of Employment of

Elderly Persons,

etc. (Kลnenreishatล

no Koyล no

Antei-tล ni kansuru

Hลritsu; Act No. 68

of 1971). The

age-70 measures

were newly

established by the

amendment under

Act No. 14 of 2020

The measures

for securing

work

opportunities

up to age 70

(Article 10-2)

took effect on

April 1, 2021

Article 8: the mandatory

retirement age (teinen) may

not be set below 60 (with

exceptions). Article 9: legal

obligation to take

employment-securing

measures up to age 65

(raising the retirement age,

continued employment, or

abolition of the retirement

age). Article 10-2: duty to

make efforts to secure work

up to age 70. Of the five

options, outsourcing

contracts (option 4) and

social-contribution projects

(option 5) are the

entrepreneurship-support

measures (Article 10-2) (nonemployment

type).

Introduction requires

consent procedures with a

majority labor union or

equivalent

Ministry of

Health, Labour

and Welfare

(amendment

overview;

statutes

database)

Japan —

revision of the

in-work oldage

pension

offset

Act Partially

Amending the

National Pension

Act, etc. for

Strengthening the

Functions of the

Pension System in

Effective April

1, 2026

Suspended amount = (basic

monthly amount + totalremuneration-

equivalent

monthly amount − threshold)

÷ 2. The threshold is raised

from 510,000 yen (the actual

FY2025 figure) to 620,000 yen

Japan Pension

Service (special

page;

calculation

method),

Ministry of

Health, Labour

Working Paper | Ageless Management in the AI Era 176

Light of Social and

Economic Changes

(Act No. 74 of 2025;

enacted June 2025)

(the statutory value at 2025

wage levels). Through wage

indexation, the actual figure

applied in FY2026 is 650,000

yen. Employed persons aged

70 and over bear no

premiums (loss of insured

status), but the benefit

suspension (offset) continues

to apply

and Welfare (act

overview)

Japan —

minors

provisions of

the Labor

Standards Act

Labor Standards

Act (Rลdล Kijun Hล;

Act No. 49 of 1947),

Chapter 6

"Minors" (Articles

56–63)

Enacted in

1947 (the

current

provisions

reflect

subsequent

amendments)

Article 56: prohibition of

employing children until the

end of the first March 31 after

they reach age 15 (exceptions

for non-industrial businesses:

age 13 or over, light labor,

permission of the

administrative agency, etc.;

for film and theater, even

under 13 on a permit basis).

Article 57: certificates of age,

etc. Article 58: prohibition of

labor contracts concluded on

behalf of minors by persons

with parental authority.

Article 59: minors'

independent claim to wages.

Article 60: prohibition in

principle of overtime and

holiday work. Article 61:

prohibition in principle of

night work (10 p.m. to 5 a.m.).

Article 62: restrictions on

employment in dangerous

and harmful work. Article 63:

prohibition of underground

labor

Hyogo Labour

Bureau

(commentary on

the minors

provisions)

Japan —

employee

status in

internships

Article 9 of the

Labor Standards

Act (definition of

"worker");

administrative

circular Kihatsu

No. 636 of

September 18, 1997

The circular

was issued in

1997

Where the profit or effect of

the work accrues to the

establishment and a

relationship of use and

subordination is recognized,

students also qualify as

workers (the Labor Standards

Act, the Minimum Wage Act,

and workers' accident

compensation insurance

apply). Where participation is

observational or experiential

and not subject to direction

and control, the student does

not qualify as a worker. The

judgment turns on substance

Nagano Labour

Bureau

materials

Working Paper | Ageless Management in the AI Era 177

(direction and control,

attribution of output,

attendance management,

etc.), not on the label

Japan —

Freelance Act

Act on Ensuring

Proper

Transactions

Involving Specified

Entrusted Business

Operators (Tokutei

Jutaku Jigyลsha ni

kakaru Torihiki no

Tekiseika-tล ni

kansuru Hลritsu;

Act No. 25 of 2023)

Effective

November 1,

2024

Obligations of commissioning

businesses: clear indication

of transaction terms in

writing or equivalent; setting

a remuneration due date

(within 60 days of receipt of

the deliverable) and payment

by that date; prohibition, in

continuing outsourcing, of

refusal of receipt, reduction

of remuneration, returns,

beating down of prices, etc.;

accurate display of

recruitment information;

accommodation for

balancing work with

childcare and caregiving;

establishment of harassmentresponse

systems; 30 days'

advance notice of mid-term

termination. Jurisdiction: the

Japan Fair Trade

Commission, the Small and

Medium Enterprise Agency,

and the Ministry of Health,

Labour and Welfare

Cabinet

Secretariat

(policy portal),

Japan Fair Trade

Commission,

Small and

Medium

Enterprise

Agency

Japan —

expansion of

special

enrollment in

workers'

accident

compensation

insurance

Article 33 et seq. of

the Workers'

Accident

Compensation

Insurance Act

(special

enrollment). The

expansion is by a

ministerial order

partially amending

the Enforcement

Regulations of the

Act, etc.

Effective

November 1,

2024 (the

same day as

the Freelance

Act)

Freelancers engaged in

specified entrusted business

(BtoB outsourcing) become

eligible for special

enrollment (BtoC work is also

covered). In effect, virtually

all freelancers may enroll

voluntarily. Premiums are

borne entirely by the

individual; the Class II special

enrollment premium rate is

the basic daily benefit

amount × 365 × 3/1000 (0.3%).

The basic daily benefit

amount is chosen from 16

steps between 3,500 and

25,000 yen. Procedures run

through special enrollment

associations

Ministry of

Health, Labour

and Welfare

(leaflet;

premium rate

table)

United States

— earnings

test reform

Senior Citizens'

Freedom to Work

Signed April 7,

2000

Abolished the earnings test

for work at and after FRA

(full retirement age). It

U.S. Congress

(statute text),

SSA (official

Working Paper | Ageless Management in the AI Era 178

Act of 2000 (Public

Law 106-182)

remains before FRA: in years

before the year of reaching

FRA, $1 is withheld for every

$2 above the lower exempt

amount; in the year of

reaching FRA (up to the

month before the month of

attainment), $1 for every $3

above the upper exempt

amount. The 2026 exempt

amounts are $24,480 per year

(lower) and $65,160 per year

(upper). Withheld amounts

are effectively recovered

through benefit

recomputation after FRA

exempt-amount

table)

United States

— ADEA

Age Discrimination

in Employment Act

of 1967 (29 U.S.C.

§621 et seq.)

Enacted in

1967. The

1986

amendment

removed the

upper age

limit of

protection

Prohibits age discrimination

in employment (hiring,

discharge, pay, promotion,

etc.) against workers aged 40

and over. With the removal of

the upper age limit,

mandatory retirement is

unlawful in most occupations

(with limited exceptions such

as pilots)

EEOC

Germany —

abolition of

the earnings

limit

Eighth Act

Amending Book IV

of the Social Code

(Achtes Gesetz zur

Änderung des

Vierten Buches

Sozialgesetzbuch

— 8. SGB IV-ÄndG)

Effective

January 1,

2023

Completely abolished the

earnings limit

(Hinzuverdienstgrenze)

during receipt of early oldage

pensions (a permanent

measure applying to all

recipients). Before abolition,

the 2022 limit was 46,060

euros per year (the COVID

special level). For disability

pensions, not abolition but a

transition to a dynamic wagelinked

limit

Official FAQ of

the German

Pension

Insurance (DRV)

EU —

Employment

Equality

Directive

Council Directive

2000/78/EC (the

general framework

directive for equal

treatment in

employment and

occupation)

Adopted

November 27,

2000

Prohibits direct and indirect

discrimination in

employment and occupation

on grounds of religion or

belief, disability, age, or

sexual orientation. Article 6

permits justification of

differences of treatment on

grounds of age, conditional

on legitimate aims such as

employment policy and on

appropriate and necessary

means (an exception for age

EUR-Lex,

European

Commission

Working Paper | Ageless Management in the AI Era 179

alone; the basis provision for

the case law of the Court of

Justice of the EU on the

permissibility of mandatory

retirement)

EU — Young

Workers

Directive

Council Directive

94/33/EC (directive

on the protection

of young people at

work)

Adopted June

22, 1994

Covers persons under 18 and

prohibits in principle the

labor of children (under 15 or

in compulsory education).

Exceptions: cultural, artistic,

sporting, and advertising

activities (permit-based);

combined work/training-type

engagement for those 14 and

over; light work for those 14

and over; light work for

limited weekly hours for

those 13 and over. Regulates

the prohibition of dangerous

work for young people,

working time, night work,

and rest

EUR-Lex

(summary), EUOSHA

ILO —

Convention

No. 138

Minimum Age

Convention, 1973

(No. 138)

Adopted in

1973. Ratified

by Japan in

2000

The minimum age shall be

not less than the age of

completion of compulsory

schooling and, in any case,

not less than 15 (Article 2(3)).

The developing-country

exception was initially 14

(Article 2(4)). Light work at

ages 13–15 may be permitted

by national law (Article 7;

developing-country exception

12–14). Hazardous work at 18

(Article 3; an exception at 16

conditional on protection and

training = Article 3(3))

ILO (convention

text; ILO Office

in Japan)

ILO —

Convention

No. 182

Worst Forms of

Child Labour

Convention, 1999

(No. 182)

Adopted in

1999.

Universal

ratification by

all member

states

achieved on

August 4,

2020. Ratified

by Japan in

2001

Obligates immediate

measures for the prohibition

and elimination of the "worst

forms of child labour" for

those under 18 (slavery,

forced labor, and trafficking;

child soldiers; sexual

exploitation; illicit activities;

hazardous work). With

Tonga's ratification (the

187th), the first universal

ratification in ILO history.

The United States has not

ratified No. 138 but has

ratified No. 182

ILO (official

announcement)

Working Paper | Ageless Management in the AI Era 180

Singapore —

RRA

Retirement and Reemployment

Act

63/68 from

July 1, 2022;

64/69 from

July 1, 2026

A two-tier structure of a

statutory retirement age

(prohibition of compelled

retirement below that age)

and a re-employment-offer

obligation age. The

government has stated a

policy of raising these to

65/70 by 2030. An obligation

to offer re-employment, not

an obligation to retain

employment

Law firms and

HR professional

media (final

confirmation

against MOM

primary sources

remains

outstanding)

Korea —

retirementage

extension

act

Act on Prohibition

of Age

Discrimination in

Employment and

Elderly

Employment

Promotion

(amended 2013;

the amendment

passed in June

2013)

Effective 2016

for

workplaces

with 300 or

more

employees

and public

institutions;

2017 for those

with fewer

than 300

Obligates setting the

mandatory retirement age at

60 or above. The same act

also prohibits employment

discrimination on grounds of

age. Extension of the

statutory retirement age to 65

is under discussion between

labor and management and

in the National Assembly

JILPT, JILAF

Note: Entries reflect what was checked as of August 2026. Amounts, ages, and the like may change with revisions to

each institution. It is stated explicitly that, for the 2026 amendment of Singapore's RRA alone, final confirmation

against primary sources (the competent ministry) remains outstanding. This table is a description of institutions, not

legal advice.

References

These references include the grey literature and industry surveys (survey reports by NGOs,

membership organizations, and companies, etc.) whose evidence grade is stated explicitly in

the main text. Unrefereed preprints and technical reports (including working papers) are

marked as such at the end of each entry. Japanese statutes and Japanese-language official

materials are collected under the subheading "Primary Legal Sources and Official Materials"

at the end.

AARP Research (Perron, R.). (2026). Foresight 50+ Omnibus Survey, Wave 3: AI and workers age 50-plus

(fielded March 12–16, 2026; 1,015 U.S. workers aged 50 and over). Washington, DC: AARP. (grey

literature; membership-organization survey)

Accenture, Disability:IN, & AAPD. (2018). Getting to Equal: The Disability Inclusion Advantage. Accenture.

(grey literature; corporate survey)

Aigner, D. J., & Cain, G. G. (1977). Statistical theories of discrimination in labor markets. Industrial and

Labor Relations Review, 30(2), 175–187.

Akerlof, G. A. (1970). The market for "lemons": Quality uncertainty and the market mechanism. Quarterly

Journal of Economics, 84(3), 488–500. doi:10.2307/1879431

Alter, A. L., & Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation.

Personality and Social Psychology Review, 13(3), 219–235. doi:10.1177/1088868309341564

Working Paper | Ageless Management in the AI Era 181

Ayalon, L. (2026). Intergenerational relations in the workforce in the age of artificial intelligence: Where

do we go from here? International Psychogeriatrics, Article 100235. doi:10.1016/j.inpsyc.2026.100235

Backes-Gellner, U., & Veen, S. (2013). Positive effects of ageing and age diversity in innovative companies –

large-scale empirical evidence on company productivity. Human Resource Management Journal, 23(3),

279–295. doi:10.1111/1748-8583.12011

Baltes, P. B., & Baltes, M. M. (1990). Psychological perspectives on successful aging: The model of selective

optimization with compensation. In P. B. Baltes & M. M. Baltes (Eds.), Successful Aging: Perspectives

from the Behavioral Sciences (pp. 1–34). Cambridge University Press. doi:10.1017/

CBO9780511665684.003

Barney, J. (1991). Firm resources and sustained competitive advantage. Journal of Management, 17(1), 99–

120.

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcฤฑ, Ö., & Mariman, R. (2025). Generative AI without

guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National

Academy of Sciences, 122(26), e2422633122. doi:10.1073/pnas.2422633122 (Correction: doi:10.1073/pnas.

2518204122 — correction of author affiliation only)

Bick, A., Blandin, A., & Deming, D. J. (2024). The rapid adoption of generative AI. NBER Working Paper No.

32966. Cambridge, MA: National Bureau of Economic Research. (unrefereed working paper; updated

figures published by the Federal Reserve Bank of St. Louis, On the Economy, November 2025)

Bonsang, E., Adam, S., & Perelman, S. (2012). Does retirement affect cognitive functioning? Journal of

Health Economics, 31(3), 490–501. doi:10.1016/j.jhealeco.2012.03.005

Börsch-Supan, A., & Weiss, M. (2016). Productivity and age: Evidence from work teams at the assembly

line. The Journal of the Economics of Ageing, 7, 30–42. doi:10.1016/j.jeoa.2015.12.001

Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at work. Quarterly Journal of Economics,

140(2), 889–942. doi:10.1093/qje/qjae044

Budzyล„, K., et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy:

A multicentre, observational study. The Lancet Gastroenterology & Hepatology, 10(10). doi:10.1016/

S2468-1253(25)00133-5

Burmeister, A., Wang, M., & Hirschi, A. (2020). Understanding the motivational benefits of knowledge

transfer for older and younger workers in age-diverse coworker dyads: An actor–partner

interdependence model. Journal of Applied Psychology, 105(7), 748–759. doi:10.1037/apl0000466

Carlson, M. C., et al. (2008). Exploring the effects of an "everyday" activity program on executive function

and memory in older adults: Experience Corps. The Gerontologist, 48(6), 793–801.

Carlson, M. C., et al. (2009). Evidence for neurocognitive plasticity in at-risk older adults: The Experience

Corps program. Journals of Gerontology: Series A, Medical Sciences, 64A(12), 1275–1282.

Carlson, M. C., et al. (2015). Impact of the Baltimore Experience Corps Trial on cortical and hippocampal

volumes. Alzheimer's & Dementia, 11(11), 1340–1348.

Carstensen, L. L., Isaacowitz, D. M., & Charles, S. T. (1999). Taking time seriously: A theory of

socioemotional selectivity. American Psychologist, 54(3), 165–181. doi:10.1037/0003-066X.54.3.165

Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of

Educational Psychology, 54(1), 1–22. doi:10.1037/h0046743

Clark, A., & Chalmers, D. J. (1998). The extended mind. Analysis, 58(1), 7–19. doi:10.1093/analys/58.1.7

Coe, N. B., von Gaudecker, H.-M., Lindeboom, M., & Maurer, J. (2012). The effect of retirement on cognitive

functioning. Health Economics, 21(8), 913–927. doi:10.1002/hec.1771

Dawson, W. D., Smith, E., Booi, L., Mosse, M., Lavretsky, H., Reynolds, C. F., III, ... Eyre, H. A. (2022). Investing

in late-life brain capital. Innovation in Aging, 6(3), igac016. doi:10.1093/geroni/igac016

Working Paper | Ageless Management in the AI Era 182

Dell'Acqua, F. (2022–2023). Falling Asleep at the Wheel: Human/AI Collaboration in a Field Experiment on

HR Recruiters. Working paper, Harvard Business School / Laboratory for Innovation Science.

(unrefereed working paper; not published in a refereed journal, checked against a mirror-distributed

version)

Dell'Acqua, F., McFowland, E., III, Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L.,

Candelon, F., & Lakhani, K. (2023). Navigating the jagged technological frontier: Field experimental

evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School

Working Paper 24-013 / SSRN 4573321. (unrefereed working paper; the figures in this paper are from

the WP version; the Organization Science version, doi:10.1287/orsc.2025.21838, was confirmed for

existence only)

Dror, I. E., Charlton, D., & Péron, A. E. (2006). Contextual information renders experts vulnerable to making

erroneous identifications. Forensic Science International, 156(1), 74–78. doi:10.1016/j.forsciint.

2005.10.017

Dufouil, C., Pereira, E., Chêne, G., Glymour, M. M., Alpérovitch, A., Saubusse, E., Risse-Fleury, M., Heuls, B.,

Salord, J.-C., Brieu, M.-A., & Forette, F. (2014). Older age at retirement is associated with decreased risk

of dementia. European Journal of Epidemiology, 29(5), 353–361. doi:10.1007/s10654-014-9906-3

Eibich, P. (2015). Understanding the effect of retirement on health: Mechanisms and heterogeneity. Journal

of Health Economics, 43, 1–12.

El Morr, C., Kundi, B., Mobeen, F., Taleghani, S., El-Lahib, Y., & Gorman, R. (2024). AI and disability: A

systematic scoping review. Health Informatics Journal, 30(3). doi:10.1177/14604582241285743

Fried, L. P., et al. (2004). A social model for health promotion for an aging population: Initial evidence on

the Experience Corps model. Journal of Urban Health, 81(1), 64–78.

Fujiwara, Y., et al. (2023). The relationship between working status in old age and cause-specific disability

in Japanese community-dwelling older adults with or without frailty: A 3.6-year prospective study.

Geriatrics & Gerontology International.

Generation. (2024). The AI Divide: Survey of workers aged 45 and over and employers in France, Ireland,

Spain, the United Kingdom, and the United States (released October 8, 2024). generation.org. (grey

literature; NGO-commissioned survey)

Gerpott, F. H., Lehmann-Willenbrock, N., & Voelpel, S. C. (2017). A phase model of intergenerational

learning in organizations. Academy of Management Learning & Education, 16(2), 193–216. doi:10.5465/

amle.2015.0185

Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect

mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127.

doi:10.1136/amiajnl-2011-000089

Grossmann, I., Na, J., Varnum, M. E. W., Park, D. C., Kitayama, S., & Nisbett, R. E. (2010). Reasoning about

social conflicts improves into old age. Proceedings of the National Academy of Sciences, 107, 7246–7250.

doi:10.1073/pnas.1001715107

Hartshorne, J. K., & Germine, L. T. (2015). When does cognitive functioning peak? The asynchronous rise

and fall of different cognitive abilities across the life span. Psychological Science, 26(4), 433–443. doi:

10.1177/0956797614567339

Hautus, M. J., Macmillan, N. A., & Creelman, C. D. (2022). Detection Theory: A User's Guide (3rd ed.). New

York: Routledge.

Hong, L., & Page, S. E. (2004). Groups of diverse problem solvers can outperform groups of high-ability

problem solvers. Proceedings of the National Academy of Sciences, 101(46), 16385–16389.

Horn, J. L., & Cattell, R. B. (1967). Age differences in fluid and crystallized intelligence. Acta Psychologica,

26(2), 107–129. doi:10.1016/0001-6918(67)90011-X

Working Paper | Ageless Management in the AI Era 183

Insler, M. (2014). The health consequences of retirement. Journal of Human Resources, 49(1), 195–233. doi:

10.3368/jhr.49.1.195

Joshi, A., & Roh, H. (2009). The role of context in work team diversity research: A meta-analytic review.

Academy of Management Journal, 52(3), 599–627. doi:10.5465/amj.2009.41331491

Kadowaki, N. (2026a). Future Value Theory(ๆœชๆฅไพกๅ€ค็†่ซ–). VURA Working Paper Series No.1. ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃใƒ”

ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026b). Enterprise Redefinition(ไผๆฅญๅ†ๅฎš็พฉ). VURA Working Paper Series No.2. ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃ

ใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026c). Enterprise Redefinition Observed. VURA Working Paper Series No.3. ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃใƒ”

ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026d). From Job Description to Purpose Description. VURA Working Paper Series No.4.

ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026e). Brain Capital Management(่„ณ่ณ‡ๆœฌ็ตŒๅ–ถ). VURA Working Paper Series No.5. ใƒ“ใƒฅใƒผใƒฉ

ใ‚ญใƒฃใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026f). Self-Defined Society(่‡ชๅทฑๅฎš็พฉๅž‹็คพไผš). VURA Working Paper Series No.6. ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃ

ใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026g). Redefinition Capitalism(ๅ†ๅฎš็พฉ่ณ‡ๆœฌไธป็พฉ). VURA Working Paper Series No.7. ใƒ“ใƒฅใƒผใƒฉ

ใ‚ญใƒฃใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kadowaki, N. (2026h). Human on the Loop(ไบบใฎใ‚ฌใƒใƒŠใƒณใ‚นใฏใชใœ็ ด็ถปใ™ใ‚‹ใฎใ‹). VURA Working Paper Series

No.8. ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ.

Kaše, R., Saksida, T., & Miheliฤ, K. K. (2019). Skill development in reverse mentoring: Motivational

processes of mentors and learners. Human Resource Management, 58(1), 57–69. doi:10.1002/hrm.21932

Kirsh, B., Stergiou-Kita, M., Gewurtz, R., Dawson, D., Krupa, T., Lysaght, R., & Shaw, L. (2009). From margins

to mainstream: What do we know about work integration for persons with brain injury, mental illness

and intellectual disability? WORK, 32(4). doi:10.3233/WOR-2009-0851

Krzeminska, A., Austin, R. D., Bruyère, S. M., & Hedley, D. (2019). The advantages and challenges of

neurodiversity employment in organizations. Journal of Management & Organization, 25(4), 453–463.

doi:10.1017/jmo.2019.58

Li, Y., Gong, Y., Burmeister, A., Wang, M., Alterman, V., Alonso, A., & Robinson, S. (2021). Leveraging age

diversity for organizational performance: An intellectual capital perspective. Journal of Applied

Psychology, 106(1), 71–91.

Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle:

How Language Models Use Long Contexts. Transactions of the Association for Computational

Linguistics, 12, 157–173. doi:10.1162/tacl_a_00638

Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of

the American Medical Informatics Association, 24(2), 423–431. doi:10.1093/jamia/ocw105

Macmillan, N. A., & Creelman, C. D. (2005). Detection Theory: A User's Guide (2nd ed.). Mahwah, NJ:

Lawrence Erlbaum Associates.

Marinaci, T., Russo, C., Savarese, G., Stornaiuolo, G., Faiella, F., Carpinelli, L., Navarra, M., Marsico, G., &

Mollo, M. (2023). An inclusive workplace approach to disability through assistive technologies: A

systematic review and thematic analysis of the literature. Societies, 13(11), 231. doi:10.3390/

soc13110231

Meng, A., Nexø, M. A., & Borg, V. (2017). The impact of retirement on age related cognitive decline – a

systematic review. BMC Geriatrics, 17, Article 160. doi:10.1186/s12877-017-0556-7

National Academies of Sciences, Engineering, and Medicine. (2022). Global Roadmap for Healthy

Longevity. Washington, DC: The National Academies Press. (institutional report)

Working Paper | Ageless Management in the AI Era 184

Nielsen, J. (2023). Generative AI enhances old users' intellectual performance through wise winnowing.

UXTigers (published June 28, 2023; updated January 1, 2026). (essay; no empirical data)

Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial

intelligence. Science, 381(6654), 187–192. doi:10.1126/science.adh2586

Omri, A., Omri, H., & Afi, H. (2025). Exploring the impact of AI on unemployment for people with

disabilities: Do educational attainment and governance matter? Frontiers in Public Health, 13, 1559101.

doi:10.3389/fpubh.2025.1559101

Parker, M., Bucknall, M., Jagger, C., & Wilkie, R. (2020). Population-based estimates of healthy working life

expectancy in England at age 50 years: Analysis of data from the English Longitudinal Study of Ageing.

The Lancet Public Health, 5(7), e395–e403.

Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity:

Evidence from GitHub Copilot. arXiv:2302.06590. (unrefereed preprint; vendor-affiliated research)

Phelps, E. S. (1972). The statistical theory of racism and sexism. American Economic Review, 62(4), 659–

661.

Pizzinelli, C., & Tavares, M. M. (2026). Artificial Intelligence and Older Workers: Opportunity, Risk, and

Policy Tradeoffs. Pension Research Council, Wharton School (May 21, 2026, blog essay; related working

paper: AI and the Future of Work in an Aging Economy, SSRN 5345347). (policy essay; grey literature)

Reber, R., & Schwarz, N. (1999). Effects of perceptual fluency on judgments of truth. Consciousness and

Cognition, 8(3), 338–342. doi:10.1006/ccog.1999.0386

Rohwedder, S., & Willis, R. J. (2010). Mental retirement. Journal of Economic Perspectives, 24(1), 119–138.

doi:10.1257/jep.24.1.119

Salthouse, T. A. (2006). Mental exercise and mental aging: Evaluating the validity of the "use it or lose it"

hypothesis. Perspectives on Psychological Science, 1(1), 68–87. doi:10.1111/j.1745-6916.2006.00005.x

Salthouse, T. A. (2009). When does age-related cognitive decline begin? Neurobiology of Aging, 30(4), 507–

514. doi:10.1016/j.neurobiolaging.2008.09.023

Schaie, K. W. (2009). "When does age-related cognitive decline begin?" Salthouse again reifies the "crosssectional

fallacy". Neurobiology of Aging, 30(4), 528–533. doi:10.1016/j.neurobiolaging.2008.12.012

Schneid, M., Isidor, R., Steinmetz, H., & Kabst, R. (2016). Age diversity and team outcomes: A quantitative

review. Journal of Managerial Psychology, 31(1), 2–17. doi:10.1108/JMP-07-2012-0228

Schooler, C. (2007). Use it—and keep it, longer, probably: A reply to Salthouse (2006). Perspectives on

Psychological Science. doi:10.1111/j.1745-6916.2007.00026.x

Sen, A. (1999). Development as Freedom. New York: Alfred A. Knopf.

Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when

trained on recursively generated data. Nature, 631(8022), 755–759. doi:10.1038/s41586-024-07566-y

Smith, P., & Smith, L. (2021). Artificial intelligence and disability: Too much promise, yet too little

substance? AI and Ethics, 1, 81–86. doi:10.1007/s43681-020-00004-5

Sommers, S. R. (2006). On racial diversity and group decision making: Identifying multiple effects of racial

composition on jury deliberations. Journal of Personality and Social Psychology, 90(4), 597–612.

Spence, M. (1973). Job market signaling. Quarterly Journal of Economics, 87(3), 355–374. doi:

10.2307/1882010

Staff, R. T., Murray, A. D., Deary, I. J., & Whalley, L. J. (2004). What provides cerebral reserve? Brain, 127(5),

1191–1199. doi:10.1093/brain/awh144

Stern, Y. (2002). What is cognitive reserve? Theory and research application of the reserve concept. Journal

of the International Neuropsychological Society, 8, 448–460.

Working Paper | Ageless Management in the AI Era 185

Stern, Y. (2012). Cognitive reserve in ageing and Alzheimer's disease. Lancet Neurology, 11, 1006–1012. doi:

10.1016/S1474-4422(12)70191-6

Strathern, M. (1997). 'Improving ratings': Audit in the British University system. European Review, 5(3),

305–321. doi:10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4

Takeuchi, H., Ide, K., Wang, H., Tamura, M., & Kondo, K. (2024). The association of agricultural and nonagricultural

work on the healthy ageing of older adults in Japan: A 6-year longitudinal study from the

Japan Gerontological Evaluation Study. Preventive Medicine Reports, 49, 102949.

Teece, D. J., Pisano, G., & Shuen, A. (1997). Dynamic capabilities and strategic management. Strategic

Management Journal, 18(7), 509–533.

Thompson, A. (2014). Does diversity trump ability? An example of the misuse of mathematics in the social

sciences. Notices of the American Mathematical Society, 61(9), 1024–1030.

Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A

systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. doi:10.1038/

s41562-024-02024-1

van Ours, J. C. (2022). How retirement affects mental health, cognitive skills and mortality; An overview of

recent empirical evidence. De Economist, 170(3), 375–400. doi:10.1007/s10645-022-09410-y

Wallrich, L., Opara, V., Wesoล‚owska, M., Barnoth, D., & Yousefi, S. (2024). The relationship between team

diversity and team performance: Reconciling promise and reality through a comprehensive metaanalysis

registered report. Journal of Business and Psychology, 39, 1303–1354. doi:10.1007/

s10869-024-09977-0

Wegge, J., Jungmann, F., Liebermann, S., Shemla, M., Ries, B. C., Diestel, S., & Schmidt, K.-H. (2012). What

makes age diverse teams effective? Results from a six-year research program. Work, 41(Suppl 1), 5145–

5151. doi:10.3233/WOR-2012-0084-5145

WHO. (2020). Decade of Healthy Ageing: Plan of Action 2021–2030. Geneva: World Health Organization

(declaration adopted by the UN General Assembly on December 14, 2020). (international organization

document)

Xue, B., Cadar, D., Fleischmann, M., Stansfeld, S., Carr, E., Kivimäki, M., McMunn, A., & Head, J. (2018). Effect

of retirement on cognitive function: The Whitehall II cohort study. European Journal of Epidemiology,

33, 989–1001.

Zélity, B. (2023). Age diversity and aggregate productivity. Journal of Population Economics, 36(3), 1863–

1899. doi:10.1007/s00148-022-00911-3

Primary Legal Sources and Official Materials

Act on Stabilization of Employment of Elderly Persons, etc. (Kลnenreisha-tล no Koyล no Antei-tล ni

kansuru Hลritsu; Act No. 68 of 1971). Ministry of Health, Labour and Welfare statutes database. https://

www.mhlw.go.jp/web/t_doc?dataId=75049000&dataType=0&pageNo=1 (primary legal source)

Labor Standards Act (Rลdล Kijun Hล; Act No. 49 of 1947), Chapter 6 "Minors" (Articles 56–63). Hyogo

Labour Bureau commentary. https://jsite.mhlw.go.jp/hyogo-roudoukyoku/hourei_seido_tetsuzuki/

roudoukijun_keiyaku/nensyousya.html (primary legal source and administrative commentary)

Act on Ensuring Proper Transactions Involving Specified Entrusted Business Operators (Tokutei Jutaku

Jigyลsha ni kakaru Torihiki no Tekiseika-tล ni kansuru Hลritsu; Act No. 25 of 2023). Cabinet Secretariat

policy portal. https://www.cas.go.jp/jp/seisaku/atarashii_sihonsyugi/freelance/index.html (primary legal

and administrative source)

Overview of the Act Partially Amending the National Pension Act, etc. for Strengthening the Functions of

the Pension System in Light of Social and Economic Changes (Act No. 74 of 2025). Ministry of Health,

Labour and Welfare. https://www.mhlw.go.jp/content/12401000/001523466.pdf (primary administrative

source)

Working Paper | Ageless Management in the AI Era 186

Ministry of Health, Labour and Welfare. Overview of the amendment to the Act on Stabilization of

Employment of Elderly Persons (securing work opportunities up to age 70) [Kลnenreisha koyล antei hล

no kaisei (70-sai made no shลซgyล kikai kakuho) no gaiyล]. https://www.mhlw.go.jp/stf/seisakunitsuite/

bunya/koyou_roudou/koyou/koureisha/topics/tp120903-1_00001.html (primary administrative source)

Ministry of Health, Labour and Welfare. Leaflet on the expansion of special enrollment in workers'

accident compensation insurance (specified entrusted business operators). https://www.mhlw.go.jp/

content/001250166.pdf (primary administrative source)

Ministry of Health, Labour and Welfare. Table of Class II special enrollment insurance premium rates

(effective November 1, 2024). https://www.mhlw.go.jp/content/

tokubetsukanyuuhokenryouritsu_R0504.pdf (primary administrative source)

Ministry of Health, Labour and Welfare (Labour Standards Bureau). Administrative circular on the

employee status of students in internships (Circular Kihatsu No. 636 of September 18, 1997). Nagano

Labour Bureau materials. https://jsite.mhlw.go.jp/nagano-roudoukyoku/library/nagano-roudoukyoku/

_new-hp/2hourei_seido/roudoukijun/internship291006.pdf (administrative circular; Labour Bureau

commentary)

Report of the Labor Standards Act Study Group, "On the Criteria for Determining 'Worker' Status under the

Labor Standards Act" (Rลdล Kijun Hล Kenkyลซkai hลkoku; December 19, 1985). In: Ministry of Health,

Labour and Welfare, Labour Standards Bureau, "Reference Materials on the Determination of

Employee Status under the Labor Standards Act" (as of October 2024). https://www.mhlw.go.jp/content/

001462701.pdf (administrative source)

Japan Pension Service. Special page on the revision of the in-work old-age pension offset (zaishoku rลrei

nenkin) (effective April 2026). https://www.nenkin.go.jp/tokusetsu/zairoukaisei.html (primary

administrative source)

Japan Pension Service. Pensions while working (calculation method of the in-work old-age pension offset).

https://www.nenkin.go.jp/service/jukyu/seido/roureinenkin/zaishoku/20150401-01.html (primary

administrative source)

Japan Institute for Labour Policy and Training (JILPT). Korea: Overview of the retirement-age extension

act (2013). https://www.jil.go.jp/foreign/jihou/2013_6/korea_02.html (public research institute source)

Age Discrimination in Employment Act of 1967 (ADEA), 29 U.S.C. §621 et seq. EEOC. https://www.eeoc.gov/

statutes/age-discrimination-employment-act-1967 (primary legal source)

Council Directive 94/33/EC of 22 June 1994 on the protection of young people at work. EUR-Lex. https://eurlex.

europa.eu/legal-content/EN/LSU/?uri=celex:31994L0033 (primary legal source)

Council Directive 2000/78/EC of 27 November 2000 establishing a general framework for equal treatment

in employment and occupation. EUR-Lex. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?

uri=CELEX%3A32000L0078 (primary legal source)

Deutsche Rentenversicherung. FAQ: Änderungen beim Hinzuverdienst (8. SGB IV-ÄndG, effective January

1, 2023). https://www.deutsche-rentenversicherung.de/DRV/DE/Rente/Allgemeine-Informationen/

Wissenswertes-zur-Rente/FAQs/Rente/Hinzuverdienst_und_Einkommensanrechnung/

aenderungen_hinzuverdienst_liste.html (primary administrative source)

ILO. Minimum Age Convention, 1973 (No. 138). Convention text (hosted by OHCHR). https://www.ohchr.org/

en/instruments-mechanisms/instruments/minimum-age-convention-1973-no-138 (primary treaty)

ILO. Worst Forms of Child Labour Convention, 1999 (No. 182) — universal ratification (August 4, 2020).

https://www.ilo.org/resource/news/ilo-child-labour-convention-achieves-universal-ratification (primary

source; official ILO announcement)

L&E Global. Singapore: Retirement age and re-employment age to be raised on 1 July 2026. https://

leglobal.law/2026/03/24/singapore-retirement-age-and-re-employment-age-to-be-raised-on-1-july-2026-

Working Paper | Ageless Management in the AI Era 187

and-other-related-changes/ (professional commentary; final confirmation against MOM primary

sources remains outstanding)

Senior Citizens' Freedom to Work Act of 2000, Public Law 106-182. https://www.congress.gov/106/plaws/

publ182/PLAW-106publ182.htm (primary legal source)

Social Security Administration. Exempt Amounts Under the Earnings Test. https://www.ssa.gov/oact/cola/

rtea.html (primary administrative source)

Working Paper | Ageless Management in the AI Era 188

© 2026 by Vura Capital Innovation Holdings Inc.
โ€‹ใƒ“ใƒฅใƒผใƒฉใ‚ญใƒฃใƒ”ใ‚ฟใƒซใ‚คใƒŽใƒ™ใƒผใ‚ทใƒงใƒณใƒ›ใƒผใƒซใƒ‡ใ‚ฃใƒณใ‚ฐใ‚นๆ ชๅผไผš็คพ

bottom of page