Ageless Management
This paper develops Ageless Management, a theory of how organizations can redesign the value of human experience in the age of generative AI and longer working lives. It argues that as AI reduces the marginal cost of tasks associated with fluid-intelligence-type processing, the relative value of accumulated domain knowledge, contextual judgment, metacognition, and experience-based pattern recognition may increase. The paper proposes that the value of experience may shift from a production premium toward an audit premium: experienced individuals may create value by detecting errors, contextual inconsistencies, and hidden risks in AI-generated outputs. It integrates three mechanisms—cognitive complementarity between generations, the audit value of experience, and the lifelong extension of Brain Capital—and introduces the concept of generational decorrelation.
Rather than claiming that age itself creates superior judgment, the paper formulates these mechanisms as falsifiable hypotheses and proposes an empirical research agenda. Ageless Management ultimately reframes demographic aging as a question of how organizations can combine AI and diverse human cognitive capabilities to improve judgment, oversight, and long-term value creation.
No. 9
-
Title: Ageless Management: A Theory of Cognitive Complementarity, the Audit Value of Experience, and the Lifelong Extension of Brain Capital in the Age of AI
-
Version: 1
-
Publish date: August 22, 2026
-
PDF: https://zenodo.org/records/22055296/files/Ageless_Management_WP_v1.3_EN.pdf?download=1
-
SSRN:
-
Author: Naoki Kadowaki
-
Publisher๏ผVURA Capital Innovation Holdings Inc.
V U R A W O R K I N G P A P E R S E R I E S — N o . 9
Ageless Management
A Theory of Cognitive Complementarity, the Audit Value of Experience,
and the Lifelong Extension of Brain Capital in the Age of AI
ใจใผใธใฌใน็ตๅถ โ AIๆไปฃใซใใใ่ช็ฅ็็ธ่ฃๆงใ็ต้จใฎ็ฃๆปไพกๅคใใใใฆ่ณ่ณๆฌใฎ็ๆถฏๆกๅผตใฎ็่ซ
Naoki Kadowaki
VURA Capital Innovation Holdings, Inc.
August 2026
Version 1.3 (August 2026). Revised in response to three rounds of internal review. This working paper is an
unrefereed draft; comments are welcome. Contact: VURA Capital Innovation Holdings, Inc. The views
expressed are the author's own and do not represent the official position of the affiliated organization.
Abstract
This paper presents a theory of Ageless Management: a management regime that removes the
variable of chronological age from decisions about role allocation, evaluation, participation,
and exit, and instead allocates roles dynamically on the basis of measured cognitive
characteristics, accumulated domain experience, health status, and the individual's own
volition. Whereas conventional senior-employment and diversity initiatives have rested on a
paradigm of accommodation and compensation, this paper envisions a management model
that, through bidirectional complementarity between heterogeneous cognitive abilities
mediated by AI, converts multigenerational and diverse talent into co-creating agents of a
managerial resource: brain capital. The paper's principal axis is the axis of age and experience
linking super-seniors (aged 60 to their 90s) and youth; the extension to socially marginalized
groups is positioned as a secondary context of generalization, applicable only insofar as the
logic of the cognitive bottleneck and AI complementation applies. The theory's point of
departure is the following asymmetry. Generative AI lowers the marginal cost of fluid
intelligence (Gf)-type tasks and raises the relative marginal value of crystallized intelligence
(Gc)-type tasks (contextual interpretation and interpersonal judgment, and above all
evaluative and audit-type tasks)—although this rise is conditional on verification actually
detecting errors. Under this shift, the "cognitive bottleneck"—whereby difficulty in performing
the Gf-type components of a job bundle has forced exit from the job as a whole—can be
removed. At the same time, the supervision of AI output retains a residual that cannot be
delegated to AI (Human on the Loop: HOTL), and the supply side of supervision as a scarce
resource comes into question. This paper conceptualizes this scarce input operationally as
"experiential audit capacity"—the capacity to detect contextual errors, practical risks, and
ethical risks in AI output—measured by detection performance on verification tasks. The
claim that this capacity is formed not by chronological age but by the interaction of domainspecific
knowledge and operational schemata (the products of long-term domain experience)
with generalized crystallized intelligence and metacognition is placed not in the definition but
as a refutable proposition, and constitutes the core of the theory. This formation has boundary
conditions: when AI output conforms to the auditor's own received views, the synergy of
confirmation bias and automation bias means experience can instead impede detection, and
in domains where technological change is rapid and the half-life of domain knowledge is
short, the audit effectiveness of experience can decline. Further, the paper formalizes
"generational decorrelation"—the tendency of supervisor pools that have lived through
different era environments and technological generations to be less likely to share blind spots
in judgment—as an organizational supply source of oversight value. Its realization is explicitly
conditioned on three requirements: a channel condition, that auditors have unmediated
access to raw output; an organizational condition, that power gradients are flattened so that
objections reach decision-making; and a model condition, that the AI output under audit is not
monopolized by a single foundation model and cross-verification by models from different
developers is used in parallel. The claims are consistently conditional. The three social
problems of population aging, constrained youth participation, and the exclusion of socially
marginalized groups convert into untapped supply sources of brain capital (stock K ×
utilization rate u) only when the conditions of AI-mediated complementarity and Brain Safety
Working Paper | Ageless Management in the AI Era 2
are met; the conversion is not automatic. Moreover, the outcomes of multigenerational staffing
are evaluated not solely by the quality of ideas but as net benefit—including the avoidance of
excessive risk and the reduction of rework, and net of the cost of decision time required for
verification. The paper presents eight definitions; twelve propositions, each with a refutation
condition—including the superiority of dynamic role allocation over fixed, chronological-agebased
role allocation (Proposition 12)—and three testable hypotheses (brain health,
multigenerational team outcomes, and the non-compressibility of experience). Most of the
propositions are untested: this paper is not a report of established facts but a proposal of
theory equipped with a map of verification.
Keywords: Ageless Management, crystallized intelligence, fluid intelligence, cognitive complementarity,
experiential audit capacity, generational decorrelation, brain capital, multigenerational ecosystem,
Human on the Loop, Brain Safety, healthy life expectancy, dynamic role allocation
JEL classification: J14, J24, J26, M12, M14, M54, O33, I31
1. Introduction
1.1 Demographic Structure and "the Domain of Longevity"
Contemporary corporate management stands in the midst of a demographic transformation
without precedent in human history. Rising life expectancy and declining fertility are trends
common to the advanced economies, and employment systems premised on the linear threestage
sequence of education, work, and retirement carry over a design conceived for an era
when life expectancy was substantially shorter than it is today. International policy
frameworks have made this transition an explicit theme: in December 2020 the United
Nations General Assembly declared the start of the Decade of Healthy Ageing (2021–2030),
designating as action areas the development of environments that support older people's
functional ability and the combating of ageism (WHO 2020; international organization
document). An extension of life expectancy does not in itself mean an extension of healthy
life expectancy, but the continuing expansion of "the domain of longevity"—the segment of
life once designated as post-retirement—cannot be ignored as a change in the preconditions
of management.
The three-stage model is embedded in firms not as mere custom but as a bundle of
institutions. Simultaneous mass hiring of new graduates designs the transition from
education to work as a one-time gateway; seniority-based treatment uses years of tenure as a
proxy variable for ability; and mandatory retirement age (teinen) systems enforce the
transition from work to retirement uniformly by chronological age. These institutions all
share the implicit assumption that age is a sufficiently good predictor of ability and role.
Public discussion of multi-track, multi-stage lives has spread, but the work of reformulating
that demand as a theory of corporate role allocation—at the level of who is allocated what,
Working Paper | Ageless Management in the AI Era 3
under which contractual form, and with what protections—has not yet been adequately
carried out. This paper takes that gap as its subject.
Moreover, this implicit assumption is itself inconsistent with the findings of cognitive
science. Cognitive aging is not a uniform decline but is asynchronous across abilities.
According to cross-sectional data from large web samples, abilities such as processing speed
peak in early adulthood, whereas the peak of accumulation-based abilities such as
vocabulary is observed in the late 60s to early 70s, and there is substantial heterogeneity in
the peak ages of cognitive abilities (Hartshorne & Germine 2015; the cross-sectional design
and the inclusion of cohort effects are discussed in detail in Section 2). If, at a given age,
abilities that decline coexist with abilities that are preserved or still growing, then an
institution that allocates roles by the single variable of chronological age is discarding too
much information. And it is this paper's contention that the discarded information contains
precisely those cognitive assets that remain unrecovered as managerial resources.
This change should not be discussed solely as a question of the quantity of labor supply.
This paper's concern is a question of quality: the cognitive assets accumulated in the
expanded "domain of longevity"—long-term domain experience, crystallized intelligence (Gc),
and metacognition—are not reaching work under the current mechanisms of role allocation.
Mandatory retirement systems and age-based treatment allocate roles by the proxy variable
of chronological age rather than by measured individual cognitive characteristics. As
discussed below, the distributions of cognitive characteristics across age groups overlap
substantially, and individual differences can exceed age differences. Uniform allocation by
chronological age can be rational only when the cost of measuring individual characteristics
is high and no alternative allocation criterion exists. The diffusion of generative AI is
undermining this very premise—that is this paper's point of departure.
One thing, however, should be made clear at the outset. This paper does not claim that
older people do not decline. The average age-related decline in fluid intelligence (Gf) is a
robust finding of cognitive science, and this paper accepts it as a premise (Section 2). Nor does
the paper assert that working makes people healthy. The relationship between work and
brain health faces difficulties of causal identification, and this paper treats it as a hypothesis
to be tested (H1) (Sections 3 and 9). What this paper interrogates is not the fact of aging itself
but a design problem of the allocation mechanism: what aspect of aging forces exit from
roles, and whether that forcing factor can be removed by AI.
The reason for setting the level of analysis at corporate management (micro) should also be
stated. Responses to population aging have so far been discussed mainly at the level of
institutional policy on pensions and mandatory retirement (macro) and at the level of
individual health and reskilling (individual). But the place where cognitive assets actually
meet tasks is the organization. Who is allocated which role, to which tasks AI is assigned, how
supervision is designed, and what contractual forms and protections participation is given—
each of these is a managerial decision, and institutional policy merely supplies the constraints
Working Paper | Ageless Management in the AI Era 4
outside them. Even if macro institutions change, unless the firm's theory of role allocation
changes, the cognitive assets of "the domain of longevity" will not reach work. Conversely,
once firm-level design is established, it can supply concrete requirement specifications to
institutional reform debates (Section 7).
1.2 The Limits of Conventional Senior-Employment and D&I Models
The senior re-employment and diversity and inclusion (D&I) initiatives undertaken at many
firms have, despite the legitimacy of their ideals, carried a structure that connects poorly to
value creation. Conventional senior employment has typically been framed in the context of
legal compliance and corporate social responsibility, and has tended to remain within a
compensatory, protective mindset—"offset age-related decline and assign routine work"—a
mindset of restoring a minus to zero. In this configuration, the individuals concerned feel
they are not expected to contribute, and the organization in turn perceives rigidifying labor
costs. The paradigm of accommodation and compensation rests on an accounting that books
its targets as costs, and to that extent it is not sustainable.
This impasse appears at three levels. At the level of motivation, the shrinking of roles saps
the individual's cognitive engagement and reproduces a self-perception of being written off.
At the financial level, continued employment unconnected to value creation is treated as a
cost center and becomes a target of cuts with every business cycle. At the strategic level, both
senior employment and D&I are marginalized as initiatives outside the core business and
never become design variables of the business model. It has also been pointed out that D&I
initiatives have fallen into an isomorphic impasse. So long as the targets are framed as beings
to be protected, the success of the initiative depends on the continuation of goodwill and
budget, not on the success of the business. What is needed is not an additional layer of moral
persuasion but the identification of pathways by which diverse cognitive assets connect to
outcomes, and the design of the conditions under which those pathways function.
This paper locates two structural causes at the root of the impasse. The first cause is that job
design has failed to resolve the "cognitive bottleneck." A job is not a demand for a single
ability but a bundle of Gf-type components (processing speed, working memory, learning of
novel procedures) and Gc-type components (accumulated knowledge, contextual
interpretation, interpersonal judgment). When performance of the Gf-type components
becomes difficult, exit from the job as a whole is forced even if high levels of Gc-type ability
are retained. Abilities that are possessed do not reach work—this is what this paper calls the
cognitive bottleneck (formally, Definition 3 in Section 5), and conventional routine-work reemployment
has not removed this bottleneck but merely circumvented the problem by
moving people to low-load jobs that do not touch it. When generative AI sharply lowers the
marginal cost of Gf-type components, this bottleneck itself becomes removable—this is one of
the central theoretical claims of this paper (previewing Propositions 1–2).
Working Paper | Ageless Management in the AI Era 5
The second cause is that the unconditional claim that diversity generates value in itself has
not been empirically supported. Facing this point squarely, without papering over it, is a
condition of the integrity of this paper's argument. Meta-analytic results on the relationship
between age diversity and team outcomes are consistent in hovering near a zero average
effect. Joshi & Roh (2009), in a meta-analysis of 39 studies and 8,757 teams, estimated the
direct effect of age diversity on performance at r = −.06, and Schneid, Isidor, Steinmetz &
Kabst (2016), in a quantitative review of 74 studies, concluded that the overall relationship
between age diversity and team outcomes is nonsignificant (the sole exception being
turnover, r = .11, whose direction is if anything unfavorable). In the most recent and largest
registered-report meta-analysis, Wallrich et al. (2024) (615 reports, 2,638 effect sizes), the
effect of demographic diversity as a whole was r = .014, effectively zero.
Yet the same literature also identifies the conditions under which effects appear. Backes-
Gellner & Veen (2013), using large-scale linked German data, reported that age diversity has a
positive effect on firm productivity when, and only when, the firm is engaged in creative
rather than routine tasks. Wegge et al. (2012), from a research program covering three
industries and more than 745 teams, identified high task complexity, a climate low in age
discrimination, and positive appraisal of diversity as success conditions, and Wallrich et al.
(2024) likewise confirmed that the diversity–performance relationship becomes more positive
when task complexity and dependence on creative divergence are high. In other words, what
the evidence has rejected is the unconditional celebration of diversity, not conditional
complementarity. This paper theorizes the content of those conditions as the state in which AI
substitutes for and complements Gf-type components so that heterogeneous cognitive assets
become connectable across individuals—AI-mediated complementarity—and formalizes it as
a moderator of diversity's performance effect (previewing Proposition 6). The empirical
headwind against diversity is, for this paper, not a refutation but a demand to identify the
missing mediating mechanism.
On the other hand, evidence exists suggesting that homogeneity has its costs as well. In
mock-jury experiments, diverse groups have been reported to exchange a wider range of
information than homogeneous groups, with majority-group members themselves making
fewer factual errors (Sommers 2006). This, however, is laboratory research manipulating
racial diversity, and replication with age diversity has not been confirmed—this paper treats
it strictly as indirect evidence. The reason this point gains weight in the age of AI is that the
humans supervising AI output tend to be composed of homogeneous groups sharing the same
training and information environment as the AI's own era. Supervisors who learned from the
same teaching materials and grew up in the same technological environment are also likely
to share the errors they overlook. Given WP8's formalization that the value of a supervision
channel depends on the decorrelation of errors (Kadowaki 2026h), people who have lived
through different era environments, technological generations, and failure cases can carry
managerial significance not merely as objects of fairness considerations but as supply sources
Working Paper | Ageless Management in the AI Era 6
of supervisory independence. This is the preview of what this paper calls generational
decorrelation, and whether it holds is a matter for testing (Proposition 5, Hypothesis H3).
Further, data from the early diffusion phase of generative AI show a distribution that reads
as a headwind for older age groups. According to Bick, Blandin & Deming (2024), based on a
representative U.S. survey, workplace generative AI usage rates as of late 2024 were 34.5%
among those aged 18–29 and 34.6% among those aged 30–39, but 16.7% among those aged 50–
64—those in their 50s and above at roughly half the level of younger workers. In a fivecountry
survey by the employment-support NGO Generation (2024; grey literature), 90% of
U.S. hiring managers said they would consider candidates under 35 for AI-related roles, while
only 32% would consider those over 60; and in an AARP survey (2026; grey literature), only
12% of U.S. workers aged 50 and over had received AI training while 49% said they wanted to
learn—a 37-point gap between motivation and opportunity. These disparities are not
evidence of ability differences. Differences in usage rates include differences in opportunity,
training, and job composition, and the skew in hiring intentions is not even a measured
finding of discrimination from an audit study. Left unaddressed, however, AI can act to widen
the divide along age lines. The same technology can become, depending on design, either a
device for removing the cognitive bottleneck or an amplifier of new exclusion—identifying
the managerial design variables that determine which way it tips is the practical motivation
of this paper.
There is one more headwind: compression pressure on the value of experience itself.
Analyzing the introduction of a generative AI assistant to 5,172 customer-support agents,
Brynjolfsson, Li & Raymond (2025) reported that resolutions per hour rose 15% on average,
but the gains were concentrated among less experienced, less skilled workers (about +30%),
while the most experienced group saw only a small speed improvement and a slight decline
in quality. In a field experiment with 758 BCG consultants (Dell'Acqua et al. 2023; figures from
the working paper version), improvement among lower performers likewise exceeded that
among higher performers on tasks within AI's capability frontier. If AI compresses surfacelevel
differences in skill and experience, the reading that the veteran's premium will
disappear looks natural. But the same experiment also showed that on tasks outside AI's
capability frontier, the probability of reaching a correct answer fell by 19 percentage points in
the AI-using group. What is being compressed is the difference in the production of
deliverables; there is no evidence that the difference in the ability to verify the validity of AI
output against context has been compressed. On the contrary, the more unverified AI output
flows into organizations, the higher the relative value of verification capacity rises. Reading
the headwind data carefully leads not to a negation of this paper's claim but directly to the
question of what is compressed and what is not (Section 4, previewing Proposition 3).
Nor does the vulnerability of the AI transition lie only at the upper end of the age
distribution. An asymmetry has been noted whereby youth, who lose the very entry point for
accumulating experience as entry-level jobs are substituted by AI, are especially vulnerable
(Ayalon 2026). At the upper end, accumulated experience fails to reach work; at the lower
Working Paper | Ageless Management in the AI Era 7
end, the opportunity to accumulate experience itself withers—the ability to address both ends
of the problem simultaneously is the reason this paper adopts the framework of a
multigenerational ecosystem theory rather than a senior-employment theory.
To summarize. The limits of the conventional models derive not from a shortage of
goodwill but from a shortage of theory. Employment extension that leaves the cognitive
bottleneck in place does not let ability reach work. Unconditional celebration of diversity
carries no persuasive force against the near-zero average effects of the meta-analyses. And
the disparity data of the generative AI transition are an ambivalent signal, foretelling
widening exclusion if neglected and connected complementarity if designed for. This paper
rereads each of these headwinds as a demand to specify the conditions, and fixes those
conditions in the form of definitions, propositions, and refutation conditions. This is the
paper's undertaking: to carry out the shift from the paradigm of accommodation and
compensation to a paradigm of co-creation as theory, not slogan.
Evidence-grade note: among the figures in this section on usage rates, hiring intentions, and training
opportunity, Bick, Blandin & Deming (2024) is an NBER working paper (unrefereed; representative
survey), while Generation (2024) and AARP (2026) are surveys by an NGO and a membership
organization (grey literature). All are descriptive statistics based on self-report, not measurements of
ability or productivity.
1.3 Research Questions
From the above problem awareness, this paper poses the following central question. Under
what conditions does the complementation of cognitive abilities by generative AI render role
allocation based on chronological age unnecessary, and enable the cognitive assets of
multigenerational and diverse talent to be converted into the managerial resource of brain
capital? This question decomposes into three subquestions.
RQ1 (removal of the bottleneck): Does AI complementation of Gf-type components
remove the cognitive bottleneck and let retained Gc-type abilities reach work? And when
it does, how does the relative marginal value of Gc-type tasks change (Propositions 1–2)?
RQ2 (the supply side of supervision): For the non-delegable residual of supervising AI
output, does experiential audit capacity—the capacity to detect contextual errors,
practical risks, and ethical risks in AI output—constitute a scarce supply source? Is that
capacity formed by the interaction of long-term domain experience with Gc-type ability
and metacognition? And can generationally heterogeneous supervisor pools
organizationally supply the decorrelation of errors (Propositions 3–5, Hypothesis H3)?
RQ3 (conditional resource conversion): Under what conditions—AI-mediated
complementarity, Brain Safety, and institutional design not confined to the employment
boundary—does the multigenerational ecosystem increase brain capital (stock K ×
utilization rate u) bidirectionally? And what happens when those conditions are absent
(Propositions 6–11, Hypotheses H1–H2)?
•
•
•
Working Paper | Ageless Management in the AI Era 8
The three subquestions stand in a cumulative relationship. If RQ1 does not hold, the job exit
of older workers is the consequence of general ability decline, and this paper's undertaking
loses its foundation. If RQ1 holds but RQ2 does not, supervision in the age of AI becomes a
generic skill unrelated to the extent of experience, and no value specific to multigenerational
staffing arises. If RQ1 and RQ2 hold but the conditional design of RQ3 is absent, expanded
participation becomes indistinguishable from an expansion of unprotected labor supply. That
is, this paper's theory is a three-story structure—a theory of ability (RQ1), a theory of
supervision (RQ2), and a theory of institutions (RQ3)—and the propositions are arranged so
that each story can be refuted independently.
A limitation on the scope of the questions should also be attached. What this paper
addresses is the domain of knowledge work in which the evaluation, supervision, and
contextual judgment of AI output determine outcomes. No generalization is claimed to jobs
whose core is physical skill, or to domains where AI adoption itself does not advance.
Moreover, the term "super-seniors (aged 60 to their 90s)" is used throughout this paper as a
description of an age interval and does not imply that the people within that interval are
homogeneous—on the contrary, the very magnitude of individual differences is the ground
for rejecting allocation by chronological age (previewing Proposition 12).
This paper answers these questions in the mode of a conceptual paper with measurement
proposals. That is, it fixes the theory in testable form through verbatim definitions (8) and
propositions with refutation conditions (12), and draws a map of verification through
hypotheses with explicit identification strategies (3). Most of the paper's propositions are
untested, and the paper does not itself deliver empirical resolution. The value of a theory, on
this view, lies not in declaring it correct but in specifying in advance what evidence would
lead to its rejection.
Terminological discipline is also made explicit in the introduction. The structure in which
humans supervise the operation of AI is written throughout this paper as Human on the Loop
(HOTL), and the expression Human-in-the-Loop is used verbatim only when quoting primary
sources. In this paper the term HOTL is a description of function and is not intended to confer
any doctrinal authority. "Super-seniors (aged 60 to their 90s)" is used as a description of an
age interval, and "socially marginalized groups" as a general term for people who have
experienced exclusion from institutions and markets; neither carries a connotation of
protected status. The concepts requiring definition—Ageless Management, Gf-type/Gc-type
tasks, the cognitive bottleneck, experiential audit capacity, generational decorrelation, AImediated
complementarity, the multigenerational ecosystem, Brain Safety—are fixed as
verbatim definitions in Section 5; the introduction confines itself to previewing their main
points.
Working Paper | Ageless Management in the AI Era 9
1.4 Positioning of This Paper and Adjacent Research
This paper is the ninth in the VURA Working Paper Series and sits in the third layer of the
series' three-layer architecture—Layer 1, era structure (macro): Redefinition Capitalism
(Kadowaki 2026g); Layer 2, social structure (meso): the Self-Defined Society (Kadowaki 2026f);
Layer 3, corporate management (micro): the body of management theories (Figure 1). The
paper inherits its framework from three papers. First, Future Value Theory, which inverts the
origin of value from past cash flows to future value (Kadowaki 2026a). Second, Brain Capital
Management, which conceives brain capital as "stock (K) × utilization rate (u)" and places its
protection and disclosure at the foundation of management (Kadowaki 2026e). Third, the
Human on the Loop (HOTL) theory, which formalizes the value of AI supervision as
independence × detection probability and generalizes the value of a supervision channel to
the "decorrelation of errors" (Kadowaki 2026h). In addition, the framework of enterprise
redefinition and its observation (Kadowaki 2026b, 2026c) and the role-design theory shifting
from the distribution of work to the distribution of purposes and questions (Kadowaki 2026d)
are used for positional reference only.
The content of the inheritance should be specified. From BCM (Kadowaki 2026e), this paper
inherits the measurement framework that conceives brain capital as the product of stock (K)
and utilization rate (u), and the discipline of distinguishing the assessment of states from the
execution of measures. This paper extends that framework from the employees of a single
firm to the participants of a multigenerational ecosystem not confined to the employment
boundary—curbing the depreciation of K among super-seniors, forming K early among youth
and socially marginalized groups, and raising u across the organization (previewing
Proposition 8). From FVT (Kadowaki 2026a), the paper inherits the evaluative inversion that
places the origin of value in future value creation rather than past track records; this is the
premise for bringing onto the evaluation agenda the cognitive assets of youth with few years
of track record and of participants with gaps in their résumés. However, the pathway by
which diversity of thought propagates through discontinuous innovation to ESG evaluations
and the cost of capital is treated in this paper not as a verified fact but as a pathway
hypothesis (Section 6).
Consistency with WP8 (Kadowaki 2026h) is, above all, the lifeline of this paper. WP8 showed
that human supervision of AI often fails—automation bias, alarm fatigue, constraints on
supervisory capacity. Accordingly, this paper does not claim that humans are good at
supervision, nor does it treat the unverified statement that older people are good at AI
supervision as an established fact. The paper's claim takes the following form. Since a
residual of supervision that cannot be delegated to AI exists, the supply side of supervision as
a scarce resource must be interrogated, and experiential audit capacity is a candidate scarce
input. And generational decorrelation is a supply-side response to the decorrelation condition
that WP8 set out. Even if independence is secured, however, detection probability is not
thereby guaranteed. That is precisely why this paper submits this claim itself as a hypothesis
for testing (H3). Note that self-citations within the series are handled in three categories
Working Paper | Ageless Management in the AI Era 10
—"inherited," "referenced," and "tested in this paper"—and claims lacking independent
supporting evidence outside the series are not made to appear established through chains of
self-citation (Section 10). In particular, the extension of HOTL's decorrelation condition to the
generational axis (Proposition 5) and the relationship between BCM's discussion of learning
protection and the non-compressibility of experience (H3) are newly claimed by this paper
and are untested.
On top of this inheritance, the temporal robustness of the theory is declared in advance.
What this paper's propositions rest on is not the transient performance gap of the current
model generation—"today's LLMs are poor at contextual verification." They rest on the
structure of verification independence formalized by WP8 (Kadowaki 2026h, Propositions 10
and 11): whatever the intelligent system, a verifier whose errors are correlated with its own
training distribution cannot self-supply statistical independence in detecting its own
systematic blind spots. The independence of verification is not supplied by improvements in
capability. Because this structure does not depend on the performance profile of any
particular model generation, advances in AI self-correction and self-verification do not, in
themselves, render this paper's propositions obsolete. What changes is the relative
composition of the supply sources of decorrelation—humans with heterogeneous experience,
and models of different lineages—and this paper's framework persists as that allocation
problem. This declaration does not, however, mean that the human contribution within that
allocation is invariant. The possibility that the contribution shrinks, and the means of
detecting it, are addressed head-on in Section 10.5.
Figure 1 The series' three-layer architecture. This paper is No. 9 in Layer 3—corporate management
(micro). Note: the placement is architectural and is not a claim of logical dependence.
The single-point novelty of this paper lies in the following. Whereas conventional senioremployment
and D&I theory has been a paradigm of accommodation and compensation (cost/
Layer 1 Epochal structure (macro) Redefinition Capitalism (RCap)
โฆ Redefinition Capitalism (2026g)
Layer 2 Social structure (meso) Self-Defined Society (SDS)
โฅ Self-Defined Society (2026f)
Layer 3 Enterprise management (micrEon)terprise redefinition and the theory system
โ Future Value Theory (2026a) — value theory
โก Enterprise Redefinition (2026b) — transformation framework
โข Enterprise Redefinition Observed (2026c) — evidence
โฃ From JD to Purpose Description (2026d) — role design
โค Brain Capital Management (2026e) — human foundation
โง Human on the Loop (2026h) — oversight structure
โจ Ageless Management — this paper: extension along the age axis (extends โค; bridges โ and โง)
Working Paper | Ageless Management in the AI Era 11
CSR), this paper presents a management model that converts multigenerational and diverse
talent into co-creating agents of the managerial resource of brain capital, through
bidirectional complementarity between heterogeneous cognitive abilities mediated by AI. A
terminological note is in order. When this paper says bidirectional complementarity, it refers
to the bidirectionality specified in Definition 6 (Section 5)—that the benefits of
complementation (substitution of Gf-type components) and the supply of audit (provision of
Gc-type verification) flow mutually among participants. The "bidirectional capital formation"
of Proposition 8 (that brain capital increases in all participating strata) is a distinct claim, and
the two are used separately. The theoretical contribution condenses into three operations.
First, the operation of re-attributing the supervisory value of older workers from "age" to "the
interaction of long-term domain experience × Gc × metacognition (= experiential audit
capacity)." Age is merely a function of the time that makes accumulation possible, and that is
exactly why this paper is "ageless (independent of age)" rather than a celebration of seniority.
Second, the operation of formalizing generational heterogeneity as an organizational supply
source of the "decorrelation of errors." Third, the operation of converting the three social
problems of population aging, constrained youth participation, and the exclusion of
marginalized groups into untapped supply sources of brain capital, conditional on AImediated
complementarity and Brain Safety. Conversely, the paper also makes explicit the
points on which it claims no novelty. The idea of conceiving AI as a prosthesis for Gf-type
abilities belongs to the lineage of assistive-technology and prosthetics research (Section 2),
and the idea of taking marginalized groups as objects of inclusive design has precedents in
inclusive design. This paper's contribution lies not in the invention of the individual insights
but in the composition that integrates them into a refutable management theory.
Here, the paper explicitly declares the primary–secondary relationship of its scope. The
principal axis of this paper is the axis of age and experience linking super-seniors (aged 60 to
their 90s) and youth. The core of the definitions and propositions—the formation of
experiential audit capacity (Proposition 4), generational decorrelation (Proposition 5)—is
formulated on this axis, and the verification designs (H1–H3) also target this axis. By contrast,
the extension to socially marginalized groups is a secondary context of generalization,
applicable only insofar as the logic of the cognitive bottleneck (Definition 3)—the structure
whereby difficulty in performing some components of a job bundle prevents retained
abilities from reaching work—and AI complementation applies. What must be stressed is that
this generalization does not claim identity of mechanism. Age-related change in Gf-type
components, the constraints on experience accumulation among youth still in development,
and sensory, motor, and cognitive accessibility constraints are heterogeneous as mechanisms,
and transplanting the argument of the aging axis unmodified onto the developmental axis or
the accessibility axis would be an overgeneralization that erases the mechanisms specific to
each axis and the contexts of the people concerned. What this paper claims is limited to the
commonality of the allocation structure—that difficulty in performing some components of a
bundle forces exclusion from the whole bundle—and the applicability of the logic that AI
Working Paper | Ageless Management in the AI Era 12
complementation can act on that structure. The particulars of marginalized-group
participation grounded in the mechanisms, institutions, and participatory research specific to
each axis remain a task beyond this paper.
In addition to the primary–secondary relationship of scope, the meta-structure of the
theory's lineage is also declared. This paper traverses several bodies of literature—cognitive
science (Section 2), health epidemiology (Section 3), and institutional theory (Section 7)—but
these are not parallel pillars. The principal axis of this paper is the theory of firm competitive
advantage. The paper connects to the lineage of the resource-based view, which holds that a
managerial resource becomes a source of sustained competitive advantage when it is
valuable, rare, inimitable, and non-substitutable (Barney 1991), and of dynamic capabilities
theory, which explains competitive advantage by the capacity to integrate and reconfigure
such resources and competences under changing environments (Teece, Pisano & Shuen 1997),
in the following way. Experiential audit capacity (Definition 4) is a path-dependent product
that can be accumulated only over time through long-term domain experience (Propositions
3 and 4); it cannot be procured instantly on the market, and, if Proposition 3 is correct, cannot
be replicated by AI—in that sense it is a candidate for a rare and inimitable supervisory
resource. And the composition that organizes this resource—the supervisory portfolio that
designs the decorrelation of errors (Section 6), dynamic role allocation (Proposition 12)—sits
at the level of dynamic capabilities that reconfigure resources under change. FVT's (Kadowaki
2026a) evaluative inversion toward future value also connects with this paper on this lineage
of competitive advantage. By contrast, institutional theory (Section 7) and the health
discussion (Section 3) are positioned as the institutional context and the complementary
assets that make the attainment of competitive advantage possible. The protective institutions
for non-employment participation are the institutional context that enables connection to the
resource of experiential audit capacity, and brain health and Brain Safety (Section 8) are the
complementary assets that prevent the resource's depreciation and sustain its utilization. The
descriptions in Sections 3 and 7 should therefore be read not as free-standing health theory
or policy theory but as the specification of the conditions for competitive advantage. This
structure is reconfirmed in the conclusion (Section 11).
The paper's normative anchor is also made explicit. When this paper says "brain capital," it
is not a resource concept that decomposes human beings into parts of brain function and
selects and reuses only the parts with market value. Such a reading—the commodification of
human beings, or a eugenic reading that sorts people by cognitive characteristics—this paper
explicitly rejects. The paper's concept of capital connects to Sen's (1999) capability approach,
which conceives development as the expansion of the substantive freedoms to live the lives
people have reason to value, and the protection and extension of brain capital are positioned
as the foundation of cognitive dignity and human flourishing. That is, measurement and
allocation are not devices for ranking people but means of returning to individuals the
freedom of participation that the crude proxy variable of chronological age has taken from
them. This normative premise functions not as decoration but as a design constraint. That
Working Paper | Ageless Management in the AI Era 13
Brain Safety (Section 8) includes protection from exploitation among its requirements, and
that Section 10 discusses the misuse of measurement and its diversion into selection as risks
of this paper itself, are consequences of this anchor. Even so, the danger that measurement
converts into a new selection device does not disappear, and this paper takes on that danger
as self-criticism in Section 10.
The differences from adjacent research are as follows. Dawson et al. (2022) advocated
investment in "Late-Life Brain Capital" and called for a value shift integrating the knowledge
and experience of older people into the economy and society. This is the direct predecessor
framework of this paper, but it remains an invited-article advocacy piece and lacks
definitions, propositions, refutation conditions, and verification designs at the level of a
management regime. This paper operationalizes that advocacy into a theory that firms can
implement and test. Ayalon (2026) discusses intergenerational relations in the workplace in
the age of AI and points out the asymmetry that youth may be especially vulnerable owing to
the substitution of entry-level jobs. Where that article discusses, from the standpoint of
ageism and social policy, how AI generates intergenerational tension, this paper formalizes,
from the standpoint of management and organization, the design conditions under which
multigenerational staffing generates value—the two stand in a relationship of
complementarity, not opposition. Nielsen (2023) discussed a division of labor in which older
knowledge workers use experience to winnow the options AI generates ("wise winnowing"),
but this is an essay without empirical data, and this paper translates that intuition into
testable propositions (Propositions 3–5) and an experimental design (H3). A publication with
the similar title "Ageless Collaboration" also exists, but this paper does not rely on it. Finally,
the research gap is made explicit. Whether expertise improves AI supervision is itself
contested—there are reports that experts, too, overlook AI's errors—and, a fortiori, no study
operationalizing age or tenure to measure AI-supervision performance could be found within
the scope of this paper's search (as of August 21, 2026). This paper's propositions give this gap
a testable form.
1.5 Structure of This Paper
The structure of this paper is as follows. Section 2 organizes the cognitive-science
foundations: the Gf/Gc distinction, the heterogeneity of ability-specific peak ages, the lineage
of selective optimization with compensation (SOC) and the extended mind, and the
controversy over cognitive reserve—presented as controversy. Section 3 examines the
evidence on work and brain health with evidence grades, identifying the difficulties of causal
identification and the moderating variable of "quality of work." Section 4 organizes the
evidence on what generative AI compresses (differences in production) and what it does not
(differences in verification), and carries out the conceptual separation of domain experience
from chronological age. Section 5 is the theoretical core of the paper, presenting eight
definitions and twelve propositions together with refutation conditions. Section 6 discusses
the design of the multigenerational ecosystem: the role reconfiguration of the four players,
Working Paper | Ageless Management in the AI Era 14
the operation of dynamic role allocation, and the organizational design of generational
decorrelation. Section 7 compares internationally the institutional hurdles to employmentbased
and non-employment-based participation, treating as a proposition the risk of
degeneration into exploitation brought about by gaps in institutional protection. Section 8
designs Brain Safety in both directions: protection from cognitive load and protection from
exploitation. Section 9 presents the map of measurement and verification, detailing the
identification strategies and experimental designs for Hypotheses H1–H3. Section 10 discusses
the paper's limitations and self-criticism—the risk of new stereotyping, the misuse of
measurement, survivorship bias, and conflicts of interest—and Section 11 concludes.
Appendix A contains the experimental protocol for Hypothesis H3, and Appendix B the
details of the institutional comparison.
A reading guide is attached. Readers who wish to confirm only the skeleton of the theory
may read the definitions and propositions of Section 5 and the map of verification of Section 9
first. Practitioners interested in implementation will find Section 6's design discussion and
Section 8's Brain Safety guidelines relevant. The limitations and self-criticism of Section 10,
however, are part of this paper's claims, and it is hoped that they will be consulted alongside
the propositions whenever these are cited or used. Until it is tested, this paper's theory is a
proposal, and depending on the results of testing it may be rejected—presenting it with that
possibility left open is, on this view, the first safeguard against the term Ageless Management
turning into a device of new age stereotyping.
2. Cognitive-Scientific Foundations
This section organizes the findings of cognitive science that constitute the theoretical
premises of Ageless Management. The purpose is fourfold. First, it introduces the distinction
between fluid intelligence (Gf) and crystallized intelligence (Gc) accurately from the original
sources, and confirms that this distinction is a relative weighting on a continuum, not a
binary classification (Section 2.1). Second, it presents the empirical finding that the peak ages
of cognitive abilities differ greatly across abilities, while at the same time making explicit the
limitations of the cross-sectional designs on which that finding rests (Section 2.2). Third, it
introduces SOC theory, which frames adaptation to aging as "selective optimization with
compensation," and the "extended mind" framework, which does not confine cognition to the
inside of the skull, and positions cognitive complementation by AI within this lineage of
compensation (Section 2.3). Fourth, it organizes the controversy surrounding the concept of
cognitive reserve and the "use it or lose it" hypothesis, together with the evidence on both
sides (Section 2.4). The work of this section is the foundation for the subsequent theory
construction (Section 5), and it places priority on drawing a clear line between the parts
where the evidence is established and the parts that remain contested.
Working Paper | Ageless Management in the AI Era 15
2.1 Fluid and Crystallized Intelligence — Cattell-Horn Theory
The standard starting point for discussing age-related change in intelligence is the distinction
between fluid intelligence (Gf) and crystallized intelligence (Gc). Cattell (1963) formulated,
through a critical experiment, the theory that intelligence is not reducible to a single general
factor but is composed of two general factors of different natures. Fluid intelligence (Gf) is the
reasoning capacity for solving novel problems, referring to cognitive operations that depend
relatively little on prior knowledge — working memory, information-processing speed, and
the learning of novel procedures. Crystallized intelligence (Gc) is the knowledge and skill
accumulated through education, occupation, and life experience, manifesting as vocabulary,
general knowledge, contextual interpretation, and interpersonal judgment.
The original source that demonstrated the age differences of these two factors with crosssectional
data is Horn & Cattell (1967). That study reported a divergence pattern in which
fluid intelligence (Gf) shows a declining trend with age throughout adulthood, while
crystallized intelligence (Gc) shows a trend of maintenance or increase. This divergence is the
starting point of the entire argument of this paper. That is, age-related cognitive change is not
a uniform "decline" but a structural change in which declining components and maintained
or growing components coexist. The singular statement "cognitive function declines with age"
is a summary that paints over this structure with an average, and in the process of
summarizing it discards the information most important for management — which
components remain and which components are lost. From the standpoint of role design, the
identification of the components that are lost and the components that remain is precisely the
starting point, and the Gf/Gc distinction provides the minimal vocabulary for that purpose.
Note that the study is based on cross-sectional data, and it is not appropriate to attribute
specific onset ages of decline or effect sizes to that paper (the limitations of cross-sectional
designs are detailed in Section 2.2).
Two conceptual cautions should be made explicit here. First, Gf and Gc are statistical
constructs extracted by factor analysis, and it is rare for a real task or job to depend on only
one of them. Drafting a document requires both vocabulary (Gc) and working memory (Gf),
and negotiation requires both contextual interpretation (Gc) and immediate information
processing (Gf). Accordingly, the distinction between "Gf-type tasks / Gc-type tasks" that this
paper introduces in Section 5 (Definition 2) is a relative weighting on a continuum — on
which component the task's performance primarily depends — and not a binary
classification. Blurring this point leads to a new simplification — "just give older workers Gctype
jobs" — which is itself a variant of the age stereotyping this paper criticizes.
Second, the Gf/Gc distinction does not imply a typology of individuals in which "the young =
Gf, the old = Gc." For both factors, individual differences are no smaller than age differences,
and the distributions of age groups overlap widely. People in their 70s who retain high fluid
intelligence (Gf), and people in their 30s with rich crystallized intelligence (Gc), are entirely
ordinary. What this section treats are average tendencies of age groups, and the validity of
Working Paper | Ageless Management in the AI Era 16
using age as the basis for placing individuals is precisely what this paper examines negatively
in Section 5 (Proposition 12).
The reason this paper adopts the Gf/Gc distinction as the basis of its theory lies in its
parsimony and in the directness of its mapping onto the functional characteristics of AI.
Regarding the factor structure of intelligence, more multilayered models have been
developed after the Cattell-Horn lineage, but for this paper's purpose — constructing a
framework that identifies which cognitive components of tasks AI can cheaply substitute for
or complement and which parts remain on the human side — the two-component distinction
between the bundle of processing speed, working memory, and novel learning (Gf) and the
bundle of accumulated knowledge, contextual interpretation, and interpersonal judgment
(Gc) provides necessary and sufficient resolution. What generative AI primarily substitutes
for is Gf-type load — search, summarization, documentation, and the execution of routine
procedures — and this correspondence is the premise of Proposition 1 (the marginal value
shift) in Section 5.
An observation follows immediately from this framework. A real job is not a single task but
a bundle of tasks in which Gf-type components and Gc-type components are tied together.
And under conventional employment institutions, this bundle has been required to be
performed in its entirety by the same individual. As a consequence, when the performance of
a Gf-type component within the bundle — for example, mastering the operating procedures
of a new business system, or processing large volumes of material at high speed — becomes
difficult, exit from the job as a whole is forced, however high the level at which the rest of the
bundle can be performed. This structure, in which retained crystallized intelligence (Gc) is
lost without ever reaching the work, is what this paper formalizes in Section 5 as the
cognitive bottleneck (Definition 3). The divergence pattern confirmed in this section — the
coexistence of declining components and maintained or growing components — is the
empirical underlay of that formalization. Conversely, if the Gf/Gc divergence did not exist, or
if every ability declined with the same slope, this paper's central claim — that selective
complementation by AI changes the employability of older workers — would not hold. In that
sense, the findings of this section are the precondition on which the viability of this paper's
theory turns.
Furthermore, for the subsequent argument (Proposition 4 in Section 5), one more
distinction is introduced within the accumulation-type abilities: the distinction between
generalized crystallized intelligence — knowledge and skills such as vocabulary, reading
comprehension, and general education that are exercised across broad domains detached
from the context of acquisition — and domain-specific knowledge — facts, procedures, and
operational schemata that have meaning within a particular industry, job, or practice. Both
are products of learning and experience, and within the framework of Cattell-Horn theory
they are grouped on the same side as acquired abilities, but they differ in transferability.
Generalized Gc transfers across domains, whereas domain knowledge is bound to the domain
in which it was formed, and its value drops sharply outside that domain. What the
Working Paper | Ageless Management in the AI Era 17
vocabulary and general-knowledge scales of intelligence tests capture is mainly the former,
and what research on occupational expertise has targeted is mainly the latter. This distinction
matters for this paper. What long-term occupational experience accumulates is mainly the
latter — knowledge and operational schemata concerning the cases, failures, and tacit
constraints of the domain in question — and that is a variable distinct from vocabulary-test
scores. When Proposition 4 in Section 5 attributes the formation of experiential audit capacity
to "the interaction between domain-specific knowledge and operational schemata,
generalized crystallized intelligence, and metacognition," this conceptual distinction between
the two components is presupposed.
There is also a construct closely related to, but not identical with, crystallized intelligence
(Gc): metacognition — the capacity to monitor and control, from a bird's-eye view, one's own
cognitive states and the limits of one's knowledge. The experiential audit capacity that this
paper defines in Section 5 (Definition 4) is a capacity operationally defined by detection
performance on tasks of verifying AI output, and it is not equated with Gc alone. This paper
attributes its formation to the interaction of long-term domain experience with Gc-type
abilities and metacognition, but this is not part of the definition; it is an empirical claim to be
tested (Proposition 4). Possessing much knowledge and knowing where one's knowledge runs
out are distinct abilities, and in the context of verifying AI output, this paper holds that the
latter — the judgment of where to stop, what to doubt, and whom to consult — plays the
decisive role. This distinction is also consistent with the wisdom research discussed below
(Section 2.2), which includes "recognition of the limits of one's knowledge" among the core
dimensions of wise reasoning.
2.2 Peak Ages by Ability — There Is No Single Prime
The Gf/Gc divergence pattern was shown at finer resolution by Hartshorne & Germine (2015).
That study comprehensively analyzed two lines of evidence — cognitive-task data from
48,537 web-based participants and normative data from standardized intelligence and
memory tests — and concluded that "there is considerable heterogeneity in the peak ages of
cognitive abilities: some abilities peak around high-school graduation, some plateau in early
adulthood and begin declining in the 30s, and others do not peak until the 40s or later." The
ability-by-ability guide values are given in Table 1. Processing speed peaks at about age 18
and visual working memory at about age 25, while reading others' emotional states plateaus
from ages 40 to 60, and vocabulary peaks in the late 60s to early 70s. Even within the same
person, decades separate the point at which the first ability passes its peak from the point at
which the last ability reaches its own. Especially noteworthy is that the reading of others'
emotional states — a component of social cognition — remains on a plateau from ages 40 to
60, the middle to later stretch of a working life. This suggests that the cognitive basis of roles
requiring interpersonal judgment — negotiation, coordination, mentoring, and the contextual
evaluation of AI output on which this paper focuses — is maintained until far later than the
peak of processing speed.
Working Paper | Ageless Management in the AI Era 18
Table 1 Peak ages by cognitive ability (guide values from cross-sectional data)
Ability (task) Approximate peak age Source
Processing speed (digit-symbol
coding)
About age 18 Hartshorne &
Germine (2015)
Visual working memory About age 25 Hartshorne &
Germine (2015)
Working memory for numbers
(digit span)
Early to mid-30s Hartshorne &
Germine (2015)
Emotion reading (Reading the
Mind in the Eyes)
Plateau from ages 40 to 60 Hartshorne &
Germine (2015)
Vocabulary Late 60s to early 70s Hartshorne &
Germine (2015)
Wise reasoning (social-conflict
tasks)
Group aged 60+ scored higher than young
and middle-aged groups
Grossmann et al.
(2010)
Note: All rows are based on cross-sectional designs comparing age groups and do not directly show individual
developmental trajectories. Peak ages are guide values to be interpreted with latitude. Grossmann et al. (2010) is a
comparison across age groups rather than an estimate of peak age, and thus differs in kind from the other rows.
Figure 2 The asynchrony of peak ages across abilities. Note: Schematic. Based on Hartshorne &
Germine (2015) and related sources. Because of the cross-sectional design, note the confounding with
cohort effects.
The implication of this finding is that the question "when does cognitive function peak?" is
itself ill-posed. What Hartshorne & Germine (2015) showed is that there is no single age at
which all abilities reach their summit simultaneously — no cognitive prime — and that at
nearly every age, abilities that are still rising, abilities on a plateau, and abilities in decline
coexist. Figure 2 depicts this asynchrony schematically. This structure invalidates one-
20 30 40 50 60 70 80
Age (years)
Relative performance (schematic)
Processing speed (peak ≈ age 18) Working memory (peak ≈ age 25–30)
Emotion reading (plateau, 40s–60s) Vocabulary / crystallized knowledge (late 60s–early 70s)
Working Paper | Ageless Management in the AI Era 19
dimensional rankings along the axis of age — both "the younger, the more capable" and "the
older, the more seasoned." The processing speed of an 18-year-old and the vocabulary of a 70-
year-old are the relative heights of different abilities at different points in time, not ranks on a
single scale. Translated into the language of organizational design, members of different ages
should be described not as having "more or less of the same ability" but as having "different
ability profiles," and the heterogeneity of profiles is a resource that can be deployed
complementarily according to the cognitive demands of bundled tasks. The validity and
conditions of this translation — in particular, when mediation by AI becomes necessary — are
the subject of Sections 5 and 6.
In the domain of social judgment as well, components that improve with age have been
reported. Grossmann et al. (2010) had approximately 250 adults (sample size pending
verification against the original; three groups — young, middle-aged, and 60+) read materials
depicting social conflicts and then had their reasoning about subsequent developments blindrated
on dimensions of wisdom (recognition of others' perspectives, recognition of the limits
of one's own knowledge, emphasis on compromise and multiple resolutions, and so on). The
group aged 60 and over scored higher on wise reasoning than the young and middle-aged
groups, and this advantage held after controlling for fluid intelligence (Gf) and social class.
However, what this result measures is rated scores on a particular social-reasoning task, and
it does not support the generalization that "older people are wise." Extrapolation to
management judgment is an inference across differences in task characteristics, and this
paper treats it at the level of suggestion.
Next, the limitations of the cross-sectional designs on which these figures rest must be
made explicit. A cross-sectional study compares different age groups at a single point in time.
Thus the statement "vocabulary peaks in the late 60s to early 70s" means that the group at
that age at the time of measurement had the highest mean score; it does not directly show
that an individual's vocabulary keeps growing to that age. Cohort effects — intergenerational
differences in educational attainment and intellectual environment, the so-called Flynn effect
— inevitably contaminate age differences. Indeed, Hartshorne & Germine (2015) themselves
found that the peak of vocabulary had shifted about 15 years later than in the Wechsler
standardization data of the 1970s–80s, and cited changes in the modern environment of
intellectual stimulation as a factor. The estimated peak ages themselves move with cohorts.
Furthermore, web-based samples may carry self-selection biases with respect to education
and digital literacy.
From the standpoint of this paper, the existence of cohort effects is at once a limitation and
an implication. That age differences in peak ages and ability levels include generational
differences in educational and intellectual environments means there is no guarantee that
the values observed for today's cohorts in their 60s and 70s will apply unchanged to the
cohorts who will be in their 60s and 70s twenty years from now. Future super-seniors (aged
60 to their 90s), with longer years of education and deeper exposure to digital environments,
may enter old age with cognitive profiles different from today's same-age cohorts. Embedding
Working Paper | Ageless Management in the AI Era 20
age into institutions as a fixed indicator of ability means taking on, in addition to the
divergence between group means and individuals, a second source of error: cohort drift. This
too is one of the reasons allocation criteria should be moved from age to measured
characteristics.
The divergence between cross-sectional and longitudinal estimates has sharpened into a
public controversy over the onset age of decline. Salthouse (2009), on the basis of data
including 2,350 cross-sectional and 729 longitudinal participants, argued that "some aspects
of age-related cognitive decline begin in healthy, educated adults in their 20s and 30s." In that
study, the peaks of 12 cognitive variables lay in the range of ages 22–27, and significant
declines were detected at ages 27–42. Performance differences between ages 18 and 60 reach
about 1 standard deviation on speed tasks and 0.6–0.7 standard deviations on reasoning and
memory variables. At the same time, the same study also reported that knowledge-based
measures such as vocabulary and general knowledge rise consistently until at least age 60 —
the Gf/Gc divergence pattern is reconfirmed here as well. This paper notes, however, that the
article was an invited controversy paper, and a critical comment by Schaie, who led the
Seattle Longitudinal Study, was published alongside it in the same journal (Schaie 2009). The
rebuttal is that by longitudinal estimates the onset of decline comes after midlife, and that
cross-sectional estimates make the onset look excessively early because of cohort effects. That
is, the controversy structure itself — "decline begins in the 20s (cross-sectional)" versus "it
begins after midlife (longitudinal)" — is the current state of the evidence, and this paper does
not adopt either side as settled doctrine.
What matters for this paper's argument is that however this controversy is resolved, the
core findings of this section — the heterogeneity of peak ages across abilities, and the longterm
maintenance and growth of knowledge-based abilities — are unshaken. In both crosssectional
and longitudinal evidence, the divergence between the trajectories of processingspeed-
type abilities and knowledge-accumulation-type abilities is consistently observed. The
dispute concerns the onset timing and slope of decline, not the existence of the divergence.
In deriving managerial implications, two further reservations must be layered on. First,
average trajectories do not justify the placement of individuals. The figures above concerning
peak ages are all means of age groups; individual differences within each age group are large,
and the distributions across groups overlap widely. The difference of about 1 standard
deviation on speed tasks between ages 18 and 60 reported by Salthouse (2009) is large as a
difference of group means, but not large enough to erase the overlap of distributions.
Inference from mean differences to the treatment of individuals is the very structure of
statistical discrimination. Second, "peak ages" based on cross-sectional data must not be
translated literally into an individual's career design. The cross-sectional finding that
"vocabulary grows until the late 60s" does not guarantee that any particular individual's
vocabulary will keep growing to that age. The reason this paper argues, from Section 5
onward, for allocation based on measured characteristics rather than age is precisely to
absorb these two reservations — the divergence between group means and individuals, and
Working Paper | Ageless Management in the AI Era 21
the divergence between cross-sectional and longitudinal evidence — on the side of
institutional design. The variable of age is a coarse proxy that appears valid only when all of
these divergences are ignored.
2.3 SOC Theory and the Extended Mind — Placing AI in the Lineage of
"Compensation"
On how individuals adapt to age-related change in cognitive resources, developmental
psychology has offered a picture different from "passive acceptance of decline." SOC theory
(selective optimization with compensation), presented by Baltes & Baltes (1990), frames
successful aging as the cooperation of three processes: (1) Selection — narrowing goals and
domains of activity; (2) Optimization — concentrating remaining resources on the narrowed
domains to maintain mastery; (3) Compensation — making up for lost means with substitute
means. In a celebrated example that Baltes and colleagues used repeatedly, the elderly pianist
Rubinstein recounted that he narrowed his repertoire (selection), concentrated his practice
on it (optimization), and deliberately played the passages just before fast ones more slowly so
that the fast passages would sound faster by contrast (compensation). At the stage of the 1990
theoretical chapter, SOC theory was a presentation of a framework, and the empirical
demonstration of its effects was left to subsequent research; but in formalizing adaptation to
aging as a design problem of resource allocation and substitution of means, it is the direct
scaffold of this paper's theory construction.
The "selection" process is underwritten on the motivational side by the socioemotional
selectivity theory of Carstensen, Isaacowitz & Charles (1999). According to that theory, the
selection of social goals is determined not by chronological age itself but by the "perception of
remaining time." When time is perceived as plentiful, knowledge-acquisition goals take
priority; when time is perceived as limited, goals of emotional meaning and emotion
regulation take priority. What matters is that the theory asserts not a "decline" of motivation
with age but a "reallocation," and moreover that time perception is plastic: young people
under time constraints show the same goal shifts as older people. This finding — that the
driver is time horizon, not age — is consistent with this paper's direction of removing
chronological age from the seat of explanatory variable. As an implication for role design, for
participants whose time horizon prioritizes goals of emotional meaning and emotion
regulation, meaning-fulfilling roles such as developing successors and assuring the quality of
judgment may have high motivational fit. But this is a suggestion from theory, not a
justification for assigning roles to particular age brackets — time horizon does not correlate
perfectly with age, and it is plastic.
The means of compensation are not confined to the interior of the individual. The
"extended mind" thesis of Clark & Chalmers (1998) presented an active externalism on which
cognitive processes are not closed within the skull and external tools and records can
function as constituents of the cognitive system. In the famous thought experiment, Otto, who
has a memory impairment, uses a notebook as a substitute for memory. If the notebook is
Working Paper | Ageless Management in the AI Era 22
reliably and constantly consulted and its contents are automatically trusted, it is a bearer of
belief on a par with in-head memory (the parity principle). What must be stressed here is that
this is a philosophical argument, not empirical research. This paper does not use this
framework as grounds for the empirical proposition that "older adults' cognition can be
supplemented by external tools." It is used as the original source of a conceptual framework
that does not confine the locus of cognitive ability to the individual skull, and as justification
for a perspective that includes external resources, AI among them, within the object of design
as parts of the cognitive system. This perspective carries a practical implication for the unit of
ability assessment. What role-allocation judgment should ask is not "what can this individual
do unaided?" but "what can this individual do as a system with reliably available external
resources?" In the same way that an institution that measured a spectacle-wearer's eyesight
without the spectacles to judge fitness to drive would be irrational, measuring performance
capacity under unaided conditions in a work environment where AI complementation is
standard is a mistaking of the object of measurement. However, as discussed later (Section 4),
the divergence between aided performance and unaided ability is itself also a risk to be
managed.
At the intersection of these two lineages — the "compensation" of SOC theory and the
"externalization" of the extended mind — stands technology as prosthesis. Spectacles and
hearing aids are prostheses of the senses; notebooks and calendars are prostheses of memory.
The fields of Assistive Technology, HCI, and disability studies have accumulated decades of
research on cognitive prosthetics for cognitive impairment and aging. The idea of using
generative AI as a prosthesis for Gf-type components — search, summarization,
documentation, and the execution of novel procedures — clearly belongs to this lineage, and
this paper claims no novelty for the idea itself. Technologies of cognitive accessibility —
memory-aid systems, reminders, screen readers, voice interfaces — were researched and
implemented before the advent of AI as means of technically lifting the exclusion from
activity grounded in cognitive constraints. Likewise, the practice of inclusive design, which
treats disability and old age not as objects of special accommodation but as initial conditions
of design, also precedes this paper, and the perspective of treating the participation of socially
marginalized groups as a design problem is not this paper's invention either. This paper's
theoretical contribution does not lie there. The novelty this paper claims, as shown in Section
5, lies in bidirectional complementarity (Definition 6) — heterogeneous cognitive assets
connected across individuals through the mediation of the AI prosthesis — and in formalizing
it as an allocation problem of management resources. Whereas a prosthesis is a one-way
relation that fills an individual's deficit, what this paper treats is an organizational structure
in which the heterogeneous abilities freed by prostheses audit and complement one another.
Summarized from the standpoint of SOC theory, this paper's undertaking is as follows.
Traditionally, selection, optimization, and compensation were adaptive strategies that
individuals carried out within their own resource constraints. When AI provides
compensation for Gf-type components cheaply, the unit of selection and optimization
Working Paper | Ageless Management in the AI Era 23
becomes extensible from the individual to the organization. That is, room emerges to design
organizationally — on the basis of measured characteristics rather than the coarse proxy of
chronological age — who selects which roles and which abilities to optimize for. This is the
cognitive-scientific grounding of dynamic role allocation in Section 5 (Definition 1 and
Proposition 12).
At the same time, it should be foreshadowed that AI as compensation has a distinctive
failure mode that traditional prostheses largely lacked. Spectacles do not form false images,
but generative AI generates hallucinations, outputting plausible errors in fluent form.
Compensation of Gf-type components by AI therefore generates a new cognitive load — the
verification of output — a load that, this paper argues, requires Gc-type abilities and domain
experience. The structure in which compensation is not an unconditional solution but
demands another scarce resource, oversight, is treated in earnest in Section 4 (the evidence
on failures of AI oversight) and Section 5 (experiential audit capacity, Definition 4). What
should be confirmed at the level of this section is that this structure can be described without
contradiction inside SOC theory. The introduction of a means of compensation demands new
selection and optimization — who should attain mastery in verification — and that allocation
problem is precisely the subject of this paper.
2.4 Cognitive Reserve and the Controversy over "Use It or Lose It"
In discussions of Ageless Management, the claim that "continuing to work helps maintain
cognitive function" is often placed as a premise. This paper does not adopt that premise
without examination. This section distinguishes two related concepts — cognitive reserve and
the "use it or lose it" hypothesis — and then organizes the current state of the evidence
together with the arguments on both sides. The two are often spoken of interchangeably, but
the structure of their claims differs. Cognitive reserve is a claim of "buffering of level" — that
accumulated intellectual assets buffer the impact of brain pathology and age-related change;
use it or lose it is a claim of "alteration of slope" — that continued intellectual activity changes
the rate of decline itself. The argumentative move of supporting the latter with evidence for
the former, while leaving this distinction blurred, circulates widely in practitioner discourse.
Cognitive reserve is a concept introduced to explain individual differences in the ability to
maintain cognitive function even in the face of age-related brain change or pathology. Stern
(2002) presented a framework distinguishing passive "brain reserve," referring to hardwarelike
individual differences such as brain volume, from active "cognitive reserve," the
efficiency and flexibility of processing. Years of education, occupational attainment, and
engagement in intellectual activity are used as proxy indicators of reserve, and the
framework has permeated the epidemiology of aging and dementia (Stern 2012). However, as
Stern himself repeatedly cautions, the evidence for reserve consists of observational findings
based on proxy indicators, and the causal assertion that "cognitive reserve prevents
dementia" cannot be drawn from the current evidence. In addition, the limitations of the
proxy indicators themselves require attention. Years of education and occupational
Working Paper | Ageless Management in the AI Era 24
attainment are strongly entangled with socioeconomic status, upbringing environment, and
childhood intelligence, and even if an association between these indicators and late-life
cognitive function is observed, it is difficult to disentangle from observational data alone
whether it is an effect of accumulated intellectual load or a reflection of confounding initial
conditions.
An example of observational evidence consistent with the cognitive-reserve hypothesis is
Staff, Murray, Deary & Whalley (2004). Using 92 members of the Aberdeen 1921 birth cohort,
with intelligence-test scores at age 11 and cognitive function at age 79, together with brain
MRI in old age, the study examined whether three candidate proxies of reserve (years of
education, intracranial volume, and occupational attainment) predicted late-life cognitive
function even after controlling for childhood intelligence and age-related brain change. The
results were that education explained 5–6% of the variance in late-life memory, and
occupational attainment about 5% of memory and 6–8% of reasoning, while intracranial
volume was not a significant predictor; the authors concluded that this was consistent with
an active model in which lifetime intellectual load accumulates reserve. This paper cites this
study as observational evidence supporting the cognitive-reserve hypothesis. What must be
stressed is that this study is not a demonstration of the "use it or lose it" hypothesis. The
sample is an observational study of 92 people and cannot establish causation; what was
measured was the association between the proxy indicators of education and occupation and
late-life cognitive level, not whether continued intellectual activity changes the rate of
decline. Moreover, the variance explained is small, at 5–8%.
What made this distinction between "level" and "slope" explicit is Salthouse's (2006) critical
examination of the "use it or lose it" hypothesis. Salthouse pointed out that the existence of a
correlation between intellectual activity and cognitive performance (more active people score
higher) and the claim that intellectual activity changes the rate of age-related decline itself
are separate claims, and that it is the latter that requires testing. Examining the available
evidence on that basis, he concluded that "at present there is little scientific evidence that
differences in engagement in intellectually stimulating activities alter the rate of mental
aging." This critique weighs heavily on this paper. The popular claim that "keeping working
protects cognitive function" is sustained by an unreflective slide from the correlation between
activity and level (which is observed) to an activity-induced change in slope (which is not
established).
However, summarizing Salthouse (2006) as "use it or lose it has been refuted" is equally
mistaken. Salthouse's conclusion is insufficiency of evidence, not disproof, and he
acknowledges the correlation between activity and performance itself. Furthermore,
Schooler (2007) presented a rebuttal in the same journal, and the question remains an open
controversy on the same pages. Salthouse himself appends a practical recommendation: even
though evidence that decline can be slowed is unestablished, there is also no evidence of
harm, so people should behave as if the hypothesis were true and continue intellectually
stimulating activities. To summarize the current state of the evidence: (1) correlations
Working Paper | Ageless Management in the AI Era 25
between intellectual activity, education, and occupation and cognitive level have been
observed repeatedly; (2) causal evidence that these change the slope of decline is lacking; (3)
the controversy is ongoing.
This controversy gives direct discipline to this paper's hypothesis design. Following
Salthouse's (2006) distinction, to claim brain-health effects of Ageless Management requires
three things: (1) showing a difference in the slope of decline, not the correlation that workers
have higher cognitive levels; (2) distinguishing health selection — the healthier keep working
— from reverse causation — cognitive decline causes retirement; and (3) identifying the
mediation — through which pathway any effect runs. How far the empirical literature on
retirement and cognitive function since the so-called Mental Retirement study has answered
this identification problem is examined in Section 3, and this paper's own testing design
(Hypothesis H1 in Section 9) makes explicit, in line with this discipline, an identification
strategy using exogenous variation and a mediation analysis of cognitive engagement.
This paper does not adjudicate this controversy. Not adjudicating it is the starting point of
this paper's theoretical construction. That is, this paper does not treat "work protects brain
health" as an established premise; it examines the empirical evidence on work and brain
health with grades in Section 3, formalizes the claim as a mediation hypothesis (Proposition
7) — if an effect exists, its pathway runs not through hours of work but through the
maintenance of cognitive engagement — and reduces it to the testable Hypothesis H1 (Section
9). What this paper inherits from the cognitive-reserve framework is not a causal assertion
but a structural perspective — that accumulated intellectual assets can function as a buffer
against age-related change — and this connects to the stock (K) concept of brain capital (Brain
Capital) (Sections 5 and 8).
Finally, one more finding about the shape of the trajectory of decline should be added, as it
gives a boundary to this paper's theory. The main analytic focus of Salthouse (2009), treated
in Section 2.2, is adults aged 18–60, but the paper also refers, supplementarily, to evidence
that the magnitude of age-related decline accelerates at older ages. In a cross-sectional sample
of about 800 adults aged 61–96 from the author's laboratory, the per-year gradient of decline
exceeded the estimates for adults under 60 — a gradient about 2 times as steep on speed
variables and nearly 4 times as steep on memory variables. This is a comparison across age
groups based on cross-sectional data — the reservation of Section 2.2 about the
contamination of cohort effects applies as it stands — and it does not directly show
nonlinearity in individual trajectories. But the finding suggests that the decline of fluid
intelligence (Gf) in old age cannot be extrapolated indefinitely as "a gentle constant slope" —
that is, decline may not remain linear but may accelerate. The implications for this paper's
theory are two. First, compensation of Gf-type components by AI (Section 2.3) may have a
floor — a state in which attentional resources themselves are depleted cannot be filled by
external complementation, and Proposition 2 in Section 5 builds this floor in explicitly as a
boundary condition of complementation. Second, making this boundary explicit is not a
Working Paper | Ageless Management in the AI Era 26
justification for exclusion in old age but an honest demarcation of the theory's scope of
application. The implications of the boundary are discussed again in Section 10.
The upshot of this section can be summarized in a single point. This section has confirmed
the Gf/Gc divergence (Section 2.1), the heterogeneity of peak ages across abilities and its crosssectional
limitations (Section 2.2), the theoretical lineage of compensation (Section 2.3), and
the unresolved controversy over activity and cognitive maintenance (Section 2.4). The
conclusion running through them is that the evidence of cognitive science shows there is no
single cognitive prime at which all abilities reach their summit simultaneously. Processingspeed-
type abilities and knowledge-accumulation-type abilities trace different trajectories;
peak ages are scattered across decades depending on the ability; individual differences
within age groups and the overlap of distributions across groups are large; and crosssectional
and longitudinal estimates remain opposed on the onset timing of decline. Under
this evidentiary situation, an institution that decides role allocation and exit by the single
variable of chronological age lacks cognitive-scientific grounding. The rational alternative is
to allocate roles dynamically on the basis of measured cognitive characteristics, accumulated
domain experience, health status, and the person's own intent. This is the section's set-up for
the definition of Ageless Management formalized in Section 5 (Definition 1), and
compensation of Gf-type components by AI (Section 2.3) is examined from the next section
onward as the technical condition that makes this allocation shift feasible.
3. Work and Brain Health: The Evidence
This section examines the empirical foundation on which this paper's theory may rest — the
evidence on the relationship between work, retirement, cognitive function, and health. The
contour of the conclusion can be stated in advance. The evidence in this field is not
monolithic. While multiple estimates exist suggesting that cognitive decline accelerates after
retirement, estimates of equal methodological standing exist suggesting that retirement
instead improves health. This paper therefore does not adopt the proposition "work protects
the brain" as a premise. The purpose of this section is to survey precisely the reach and limits
of the evidence, and to identify which questions are closed and which remain open. As
discussed below, the open question is not "does work protect the brain?" but "what kind of
work could protect it?" — and this paper's theory (Section 5) is designed as an answer to this
open question.
3.1 The Evidence Since Mental Retirement
The idea of "use it or lose it" is the most widely circulated piece of common wisdom about
cognitive aging. What must be noted here is that this is the name of a hypothesis, not the
name of an established finding (recall the controversy over cognitive reserve discussed in
Section 2.4). Work has been regarded as a natural domain of application for this hypothesis:
for most people, work is the largest single activity that simultaneously supplies intellectual
Working Paper | Ageless Management in the AI Era 27
stimulation, social contact, time structure, and a sense of role. If the hypothesis holds for
work, then retirement is a systematic loss of cognitive stimulation and should hasten
cognitive decline. This subsection surveys the body of research that has tested this prediction,
attending to the strength of the identification strategies and the direction of the results. To
state it in advance: the evidence is split between directions that support the hypothesis and
directions that do not, and the pattern of that split itself carries important information for
this paper's theory.
The starting point of the research lineage that treats the relationship between work and
cognitive function with the identification strategies of economics is the "Mental Retirement"
paper of Rohwedder & Willis (2010). They compared internationally the 2004 waves of the US
HRS (about 20,000 people), the English ELSA (about 9,000), and SHARE covering 11 European
countries (1,000–3,000 per country), using immediate and delayed recall of 10 words (0–20
points) as the cognitive measure. The share of people aged 60–64 not in paid work varies
greatly by country, from about 30% in the United States and Denmark to 80–90% in France
and Austria. This difference was produced by each country's pension, tax, and disabilitybenefit
institutions, and can be regarded as determined independently of individual cognitive
ability. Using this institutional variation as an instrumental variable, they estimated a causal
effect in which early retirement lowers memory scores by about 4.7 points (5.7 points in the
specification controlling age in one-year increments). The authors summarize that "early
retirement appears to have a significant negative and quantitatively important causal effect
on the cognitive ability of people in their early 60s."
This effect size, however, cannot be taken at face value. The standard deviation of the
cognitive scores in the sample is about 3.3, so a 4.7-point drop corresponds to more than 1.4
times that. Subsequent research has not replicated this magnitude, and some estimates —
such as Coe et al. (2012), discussed below — could not confirm even an effect in the same
direction. Being a country-level comparison, it cannot fully remove country-specific
confounders such as educational systems, test administration, and cultural differences. That
is, Rohwedder & Willis (2010) was the watershed that brought the problem of causal
identification into this field, and at the same time it is a study whose effect size should be
cited only with the reservation that "later research has produced smaller estimates, and some
estimates of zero."
There is other evidence in the same direction. Bonsang, Adam & Perelman (2012), applying
instrumental variables based on Social Security eligibility ages and individual fixed effects to
the US HRS panel, reported that retirement has a significant negative effect on cognitive
function, and that the effect appears not immediately after retirement but with a delay.
Dufouil et al. (2014) analyzed linked claims and pension data on 429,803 retired selfemployed
workers in France, and reported that each additional year of age at retirement was
associated with a dementia hazard ratio of 0.968 (95% CI 0.962–0.973), equivalent to a risk
reduction of about 3.2% per year. However, the authors themselves state explicitly that
further evidence is needed to assess whether this association is causal, and that reverse
Working Paper | Ageless Management in the AI Era 28
causation — prodromal cognitive decline leading to earlier retirement — cannot be ruled out.
This paper likewise does not cite it as a causal effect.
As longitudinal evidence tracking the same individuals across the transition into
retirement, there is Xue et al. (2018) from the British civil-service cohort Whitehall II.
Following 3,433 people for up to 14 years each before and after retirement (up to 28 years in
total), verbal memory declined about 38% faster after retirement than before, after adjusting
for age-related decline. Abstract reasoning and verbal fluency, by contrast, showed no
significant change across retirement; the effect was specific to memory. A secondary finding
is also suggestive: while employed, higher employment grade was associated with slower
memory decline, but after retirement this protective association disappeared, and the rate of
decline became equivalent regardless of grade. Note that "38% faster" is a relative value; the
decline in absolute terms is gradual. Moreover, the sample is limited to white-collar civil
servants, and the endogeneity of retirement timing remains, so this result too cannot be
asserted as causal.
What about the level of systematic review? Meng, Nexø & Borg (2017) systematically
reviewed longitudinal studies on retirement and age-related cognitive decline, and noted that
only 7 studies met the inclusion criteria (1 more added in an updated search), of which 4
derived from the same cohort (the US HRS) and thus had low independence. The results
divide by domain. For fluid intelligence (Gf) the evidence is conflicting: two high-quality
studies pointed toward retirement slowing decline, while one of moderate quality pointed
toward acceleration. Only two studies addressed crystallized intelligence (Gc), yielding no
more than weak evidence that decline accelerates only for those retiring from jobs high in
interpersonal complexity. The authors' overall assessment is that no conclusion can be drawn
that retirement uniformly accelerates cognitive decline; the evidence is mixed and there are
large research gaps. It would be an error to cite this review as grounds that "retirement has
been established as bad for cognition," and equally an error to cite it as grounds that
"retirement has been established as harmless."
Randomized controlled trial (RCT) evidence that goes beyond the limits of observational
research exists not for work itself but for structured productive social engagement.
Experience Corps is a US program in which older adults serve 15 hours per week in
elementary schools providing reading support, library support, and classroom support; Fried
et al. (2004) presented its design philosophy as a social model of health promotion. Carlson et
al. (2008), in exploratory analyses of a pilot RCT (149 people), reported that the participation
group improved in executive function and memory relative to the waitlist control group, with
improvements of 44–51% among those with impaired executive function at baseline (while
the same stratum in the control group declined). Carlson et al. (2009), in a preliminary fMRI
study of 17 women aged 65 and over, reported that participation was accompanied by
improved cognitive function and significant changes in prefrontal activity patterns.
Furthermore, the imaging substudy of the Baltimore Experience Corps Trial (Carlson et al.
2015) followed 111 people (58 intervention, 53 control; mean age 67.2; predominantly African
Working Paper | Ageless Management in the AI Era 29
Americans from low-income areas) for 24 months and reported that, whereas annual brain
atrophy of 0.8–2% normally occurs after age 65, cortical and hippocampal volumes in men in
the intervention group increased by 0.7–1.6% over two years. Women in the intervention
group showed only slight increases that did not reach statistical significance, and women in
the control group declined by about 1% over 24 months. Participants with larger volume
increases also showed larger improvements on memory tests. Because these are RCTs,
"improved" can be written — but two reservations are required. First, the participants were a
restricted low-income, urban population, and generalization requires caution. Second, this is
the effect of 15 hours per week of structured volunteering, not of employed labor, and it is not
direct evidence of the effect of continued employment. This point is taken up again in Section
3.3.
The evidence above is organized in Table 2. Three notes on how to read the table. First, the
evidence grades classify the strength of identification strategies, not a ranking of the
credibility of results. Given that instrumental-variable studies of the same A− grade estimate
opposite signs, a high grade does not guarantee agreement of conclusions. Second, the
outcomes differ across studies. Memory scores, dementia registration, brain volume, and selfreported
health are not mutually substitutable measures, and effects on "cognitive function"
and effects on "health" must be read separately. Third, the restrictions of the study
populations (civil servants only, the self-employed only, low-income urban areas only, 14
Japanese municipalities only) are not footnote-level details but direct determinants of
external validity. Table 2 should be read not as a list of answers to the single question "work
and brain health," but as a map of estimates under different populations, measures, and
identification strategies.
Table 2 Principal evidence on work, retirement, cognitive function, and health
Study Design Main result
Evidence
grade
Rohwedder &
Willis (2010)
International comparison of
HRS, ELSA, SHARE.
Institutional differences in
pensions etc. as
instrumental variables
Estimated that early retirement lowers
memory scores (out of 20) by about 4.7
points. Effect size exceeds 1.4 times the
standard deviation (about 3.3); repeatedly
criticized as too large
A−
Bonsang et
al. (2012)
US HRS panel. Social
Security eligibility-age IV +
individual fixed effects
Retirement has a significant negative
effect on cognitive function. Effect
reported to appear with a delay
A−
Coe et al.
(2012)
US HRS men. Employer
early-retirement "windows"
used as IV
Concluded that the negative simple
correlation cannot be called causal. No
clear relationship between retirement
duration and cognition for white-collar
workers; for blue-collar workers, a
positive relationship instead
A−
Insler (2014) Retirement significantly positive for
health. Mediated by improved health
A−
Working Paper | Ageless Management in the AI Era 30
US HRS. Subjective
probability of continued
work used as IV
behaviors such as reduced smoking and
increased exercise
Eibich (2015) German SOEP. Regression
discontinuity using
eligibility-age
discontinuities
Retirement improves health in the long
run and reduces healthcare utilization.
Mechanisms: relief from stress, sleep,
exercise
A−
Dufouil et al.
(2014)
429,803 retired selfemployed
workers in
France. Observational study
of linked claims and
pension data
Each additional year of retirement age
associated with dementia hazard ratio
0.968 (95% CI 0.962–0.973). Authors
themselves state causality is unestablished
B
Xue et al.
(2018)
Whitehall II, 3,433 people.
Longitudinal follow-up up
to 28 years around
retirement
Verbal memory declines about 38% faster
after retirement (relative value).
Reasoning and fluency unchanged.
Protective association of employment
grade during employment disappears
after retirement
B
Meng et al.
(2017)
Systematic review of 7+1
longitudinal studies (no
meta-analysis)
Evidence mixed. Gf conflicting; for Gc only
weak evidence of accelerated decline after
retirement from jobs high in interpersonal
complexity
S
van Ours
(2022)
Narrative review of recent
causal-inference research
On average, mental health improves with
retirement, cognitive skills decline,
mortality unaffected. Effects highly
heterogeneous by occupation,
voluntariness, etc.
N
Carlson et al.
(2008)
Exploratory analysis of
Experience Corps pilot RCT
(149 people)
Participation group improved in executive
function and memory relative to controls.
44–51% improvement in the baselineimpaired
stratum
A (small)
Carlson et al.
(2015)
Baltimore Experience Corps
Trial imaging substudy (111
people, 24 months)
Cortical and hippocampal volumes in
intervention-group men increased 0.7–
1.6% over 2 years. No significant
difference in women. Restricted
population
A
Parker et al.
(2020)
Multistate life-table
estimation, UK ELSA, 15,284
people
Healthy working life expectancy at age 50
about 9 years. Men 10.9 years, women 8.3
years; regional gap about 4.5 years
B
(descriptive)
Takeuchi et
al. (2024)
JAGES, 48,221 people, 6-year
longitudinal study (Japan)
Relative to retirees, agricultural workers
show favorable associations such as
dementia OR 0.45 and mortality OR 0.68.
Healthy worker effect not removed
B
Note: Evidence grades follow the classification of the evidence notes. A = RCT or valid natural experiment /
instrumental variables (A− with identification reservations), B = longitudinal observational study (causation cannot be
claimed), S = systematic review, N = narrative review. Note that even within the same grade the signs of the estimates
do not agree. All are peer-reviewed publications.
Working Paper | Ageless Management in the AI Era 31
3.2 The Difficulty of Causal Identification — Health Selection, Reverse
Causation, and the Voluntariness of Retirement
Before interpreting the evidence in Table 2, the central identification problem of this field
must be made explicit: "does one work because one is healthy, or is one healthy because one
works?" The association "workers are healthy" in observational data is contaminated by at
least two mechanisms. The first is health selection (the healthy worker effect). Because
healthier people keep working longer, a positive association between work and health arises
even without any effect of work. The second is reverse causation. The prodromal phase of
dementia is held to extend over 10 years or more, and if prodromal decline hastens
retirement, then what looks like "cognitive decline after retirement" is in fact "retirement
caused by cognitive decline." It is for this reason that Dufouil et al. (2014), even after
reporting in sensitivity analyses that the association remained when retirements immediately
preceding onset were excluded, still did not claim causality. A third difficulty is the
voluntariness of retirement. Retirement chosen by oneself and forced retirement due to
deteriorating health or dismissal may carry different implications for subsequent health.
The problem of voluntariness deserves a little more elaboration. As van Ours (2022)
organizes it, the effects of retirement are highly heterogeneous depending on whether
retirement is voluntary or involuntary. This heterogeneity is an obstacle to identification and,
at the same time, itself a substantive finding. Voluntary retirement usually occurs in a state in
which the transition to post-retirement life has been prepared. Forced retirement occurs as
the simultaneous loss of role, income, and social contact. If the two are lumped into the single
treatment "retirement" and an average effect is estimated, which sign emerges depends on
the composition of the two within the sample. What matters from the standpoint of Ageless
Management is that uniform age-based exit compulsion, such as mandatory retirement
(teinen) systems, is precisely an apparatus that institutionally mass-produces this "forced
retirement." If the effect of retirement depends on voluntariness, the question to ask is not
whether retirement is good or bad, but who decides the timing and form of exit. This point is
taken up again in the institutional comparison of Section 7 (international differences in
mandatory retirement and pension systems).
The standard response to these problems is the quasi-experiment: using institutional
variation determined independently of individual health and cognition — pension-eligibility
ages or the offer of early-retirement incentives — as instrumental variables (IV) or regression
discontinuities (RD). But the important point is that the reality of this field is that conclusions
can flip depending on the choice of IV. Coe et al. (2012) used as an instrument the earlyretirement
"windows" that employers are legally required to offer on a non-discriminatory
basis unrelated to cognitive ability, and concluded that the association between retirement
and cognitive decline seen in simple correlations cannot be called causal. Moreover, for
white-collar workers there was no clear relationship between retirement duration and
cognition, while for blue-collar workers a positive relationship between retirement duration
Working Paper | Ageless Management in the AI Era 32
and cognition was estimated instead. This is an estimate squarely at odds with Rohwedder &
Willis (2010) and Bonsang et al. (2012), and this paper presents the two as a pair of evidence
of equal standing.
Why do methods of equal standing produce opposite answers? One reading is that different
instruments identify effects in different subpopulations. The people who decide to retire in
response to institutional variation in pension-eligibility ages and the people who respond to
an employer's early-retirement incentive are not the same population in occupation, health
status, or voluntariness of retirement. What the instrumental-variables method estimates is a
local effect among those who responded to the institutional variation in question, not a
universal "effect of retirement." If so, the disagreement among estimates can be read less as
evidence of methodological defect than as evidence that the effect of retirement itself differs
by population and context. Indeed, the occupation-specific results of Coe et al. (2012) — no
relationship for white-collar, positive for blue-collar — show that structure invisible so long
as one asks about average effects appears the moment one asks about heterogeneity. This
reading foreshadows the argument of Section 3.3.
Evidence on the opposite side has also accumulated for health outcomes other than
cognition. Insler (2014), using as an instrument respondents' previously reported subjective
probability of working past age 62/65, reported that retirement is significantly positive for
health, and that reductions in smoking and increases in exercise enabled by increased free
time mediate the positive effect. Eibich (2015), using the discontinuity in financial incentives
generated by eligibility ages in the German SOEP, estimated that retirement improves health
in the long run and reduces healthcare utilization. As mechanisms he cites relief from workrelated
stress, increased sleep, and increased frequency of exercise. In the reading of van
Ours (2022), who organizes the recent causal-inference literature, the empirical results vary
greatly across studies, but on average, mental health improves with retirement, cognitive
skills decline the longer the retirement duration, and mortality is largely unaffected. And the
effects are highly heterogeneous by individual attributes, occupation, institutions, and the
voluntariness of retirement — so much so that the author himself describes the field as prone
to "not seeing the forest for the trees." The policy implication van Ours draws is neither
uniform extension of working lives nor uniform support for early retirement, but the
expansion of individual flexibility of choice over the timing of retirement. This implication
points in the same direction as this paper's Definition 1, which rejects uniform role allocation
by chronological age and makes the person's own intent a constituent of allocation decisions.
One note is also in order on how to read the numbers. Many of the effects cited in this
section are relative values. The "38% faster decline" of Xue et al. (2018) is a relative
comparison of rates of decline, and the decline in absolute terms is gradual. The "about 3.2%
per year" of Dufouil et al. (2014) is a conversion of a hazard ratio, and given the baseline
dementia prevalence (2.65% in that cohort), the difference in absolute risk is small. Relative
values are useful for conveying the existence of an effect, but practical decision-making
requires absolute values, and this paper does not circulate relative expressions on their own.
Working Paper | Ageless Management in the AI Era 33
This paper does not paper over this evidentiary situation. The most honest summary of the
current state is as follows. Mental and physical health can improve with retirement
(especially from high-strain, physical work). For cognitive function, multiple estimates
suggest accelerated decline after retirement, but estimates of no effect or improvement exist
with methods of equal standing, and the sign changes with occupation and the voluntariness
of retirement. Therefore it cannot be written monolithically that "work is good for health."
This paper takes up this controversy as Hypothesis H1. That is, rather than presupposing a
brain-health effect of work, it presents the effect as a testable hypothesis with specified
conditions, and designs it — identification strategy included — in Section 9.
Hypothesis H1 (Brain-Health Hypothesis)
Super-seniors engaged in Gc-exercising, AI-complemented roles show a lower rate of
cognitive decline (MoCA, etc.) than same-age peers who have exited work, mediated by
the maintenance of cognitive engagement. Identification strategy: a quasi-experiment
using exogenous variation from pension-system reforms or mandatory-retirement rules
as instrumental variables, or exploiting exogeneity in reasons for working. Treatment of
health selection and reverse causation to be specified explicitly.
3.3 The Quality of Work as a Moderating Variable
So long as the body of evidence in Table 2 is read as a "contest of average-effect estimates,"
what one obtains is deadlock. But rereading it with the heterogeneity of effects as the subject,
a consistent structure comes into view. The variable that flips the sign is not whether one is
working, but what kind of work it is, and why one left. In Coe et al. (2012), cognition moved in
the direction of improvement after retirement for blue-collar workers. Meng et al. (2017)
found weak evidence of accelerated Gc decline only among those who retired from jobs high
in interpersonal complexity. In Xue et al. (2018), the protective association of employment
grade during employment disappeared after retirement. The mechanism in Eibich (2015) was
relief from stress. These are consistent with the reading that the meaning work has for
cognition depends strongly on the cognitive content and load of the job. Exit from depleting
work can improve health; exit from cognitively rich roles can be the loss of stimulation that
had been maintained.
This paper formalizes this reading as the moderating variable "quality of work." The key
construct is cognitive engagement — active involvement in roles containing Gc-type
components such as evaluation, contextual interpretation, and interpersonal judgment. If a
brain-health effect of work exists, what carries it is not the length of working hours or the
existence of an employment contract but the maintenance of this cognitive engagement —
that is this paper's theoretical wager. That is: if a brain-health effect of work exists, it is
mediated not by the length of working hours but by the maintenance of cognitive
engagement through occupying Gc-type roles. This claim is formally stated in Section 5 as
Working Paper | Ageless Management in the AI Era 34
Proposition 7 (the cognitive-engagement pathway), with a refutation condition attached. The
point here is that Proposition 7 stands as a candidate hypothesis that explains the deadlock of
Sections 3.1–3.2. The sign of the average effect of work fails to settle, on this reading, because
the category "work" mixes cognitively rich roles with depleting ones. However, this mediating
structure is itself untested, and the need for testing by mediation analysis is stated explicitly
in the refutation condition of Proposition 7.
This formalization connects to the cognitive-scientific foundations of Section 2. As seen
there, fluid intelligence (Gf) tends to decline throughout adulthood, while crystallized
intelligence (Gc) can be maintained or rise into old age. Defining cognitive engagement as
engagement with Gc-type components carries a double implication for work in old age. First,
roles that use maintained abilities permit continued engagement and can remain sources of
stimulation. Second, roles that overtax declining abilities are depletion before they are
stimulation, and become sources of exit pressure. The single piece of Gc-side evidence found
by Meng et al. (2017) — weak evidence that decline accelerates only after retirement from
jobs high in interpersonal complexity — is at least not inconsistent with this reading, since
jobs high in interpersonal complexity are the typical case of jobs weighted toward the Gc-type
components of contextual interpretation and interpersonal judgment. However, this is a posthoc
consistency check, not a test. Whether this interpretation is right is precisely what
Hypothesis H1 must ask through mediation analysis.
Seen from this standpoint, the significance of Experience Corps is more than "an RCT of
older volunteers." Experience Corps is a program that explicitly designed the quality of the
role. Within the structured time of 15 hours per week, participants took on a role — reading
support for schoolchildren — that demands interpersonal judgment and contextual
interpretation and whose outcomes accrue to others. Fried et al. (2004) called this a "social
model" of health promotion because the intervention instrument was the provision of a
meaningful social role, not an individual prescription of exercise or cognitive training. That
its RCT arms showed, albeit in restricted populations, improvements in executive function
and memory (Carlson et al. 2008) and increases in brain volume among men (Carlson et al.
2015) suggests that the designed quality of a role can act on cognition in old age. To repeat,
this is not evidence about employed labor. Extrapolation to employment requires a bridge
through the higher-order concept of "productive social engagement," and the strength of that
bridge is itself an object of testing. But for this paper's theory this distinction is, if anything,
convenient. What this paper seeks to design is not the extension of employment contracts but
the allocation of roles within a multigenerational ecosystem (Definition 7) not limited to the
boundary of employment — and Experience Corps is precisely a case of a designed, nonemployment
role.
3.4 Healthy Working Life Expectancy — Health as a Precondition
Finally, this section confirms the constraint that determines the work-health relationship
from the opposite side: the evidence on "the number of years one can work in health in the
Working Paper | Ageless Management in the AI Era 35
first place." Parker, Bucknall, Jagger & Wilkie (2020), using multistate life-table estimation on
the English ELSA (2002–2013, 15,284 people aged 50 and over, linked to NHS mortality
records), estimated healthy working life expectancy (HWLE) at age 50 — the expected
number of years spent both healthy and in work — at about 9 years. The implication of this
figure lies less in the mean than in the distribution. Men 10.9 years versus women 8.3 years.
The North East about 4.5 years shorter than the South East; 6.8 years in the most deprived
areas. Manual workers about 1 year shorter than non-manual workers. And HWLE from age
50 falls short of the years remaining until the state pension age. That is, "working in health
until pension age" is not guaranteed even on average, and the shortfall is systematically
skewed by sex, region, occupation, and income. This is a descriptive estimate, and the figures
move with the definitions of health and work, but the implication for policy and management
design is clear. Any framework that discusses the extension of working lives must build in the
fact that health is its precondition and is unequally distributed.
At the same time, it must be noted that healthy working life expectancy is not a fixed
quantity. The estimate of Parker et al. (2020) is a description that takes current working
environments, job design, and health distributions as given, and if any of these change, the
figures can change too. The fact that manual workers' HWLE is about 1 year shorter suggests
that it is the physically and cognitively depleting components of jobs that determine the
number of workable years. If so, changing the composition of the job bundle —
complementing the depleting components with technology and shifting the center of gravity
of roles toward components that use maintained abilities — becomes an intervention
hypothesis that could act on healthy working life expectancy itself. This is the public-health
restatement of the AI complementation of Gf-type components (Proposition 2) developed
from Section 4 onward. However, whether role redesign actually extends HWLE is untested,
and this paper presents it only as a design goal.
What of the Japanese evidence? From the Japan Gerontological Evaluation Study (JAGES),
Takeuchi et al. (2024) followed 48,221 people aged 65 and over in 14 municipalities for 6 years
and reported that, compared with retirees, agricultural workers showed favorable
associations: dementia odds ratio 0.45, care-need risk ratio 0.64, severe care-need risk ratio
0.65, healthy-life-expectancy-loss risk ratio 0.69, and mortality odds ratio 0.68. Nonagricultural
workers also showed similar risk reductions across all outcomes, and the neverworked
group had higher risk than retirees. The authors cite the high physical activity of
rural areas among candidate mechanisms. However, this study has no instrumental variable,
and the authors themselves note the healthy worker effect, respondent bias in selfadministered
surveys, and the restriction to a Japanese population. This result is therefore
kept as a description of association — "favorable associations between continued work and
health outcomes are observed in large Japanese longitudinal data as well" — and is not cited
as causal, since the selection whereby healthier people keep working could explain much of
the association. In addition, Japan also has a prospective study of the association between
work and care-need onset by frailty status (Fujiwara et al. 2023), and the association between
Working Paper | Ageless Management in the AI Era 36
work and late-life health has been observed repeatedly in Japanese data as well. But all of
these are observational studies, and the identification problem is of the same form as in the
Anglo-American literature.
The evidence on healthy working life expectancy imposes two disciplines on this paper.
First, Ageless Management cannot presuppose that "everyone can work indefinitely." The
years one can work in health are finite and unequally distributed along socioeconomic lines.
A design that does not render invisible, once again, those who cannot work or can no longer
work is an internal condition of this paper's theory (this survivorship-bias problem is
revisited as self-critique in Section 10). Second, since pathways exist by which continued
work itself depletes health (recall the mechanisms in Insler 2014 and Eibich 2015), an
occupational safety and health standard that protects participants' brain capital from
depreciation — Brain Safety (Definition 8) — is not an appendage of the theory but an
essential component. This point is developed in Section 8.
3.5 Conclusion of This Section — Identifying the Open Questions
The examination in this section can be summarized as follows. First, "work protects the
brain" cannot be asserted on current evidence. IV estimates and longitudinal evidence
suggesting post-retirement cognitive decline (Rohwedder & Willis 2010; Bonsang et al. 2012;
Xue et al. 2018) coexist with estimates of equal standing suggesting no effect or health
improvement (Coe et al. 2012; Insler 2014; Eibich 2015), and the conclusion of the systematic
review (Meng et al. 2017) is likewise "mixed." Second, this deadlock is nevertheless not
uninformative. The variables that divide the sign of the effect are occupation, the
voluntariness of retirement, and the cognitive content of the work — suggesting that only
"continued work conditional on voluntariness and quality of work," not "uniform extension
of working lives," can be consistent with the evidence. Third, the RCT that designed the
quality of the role (Experience Corps) showed that designed productive social engagement
can act on cognitive function and brain structure in restricted populations.
Fourth, the years one can work in health are themselves finite and unequally distributed
(Parker et al. 2020), and although favorable associations between work and health outcomes
are observed in large Japanese longitudinal data as well (Takeuchi et al. 2024), causal
evidence purged of health selection does not exist.
Accordingly, what is open is not the average-effect question "does work protect the brain?"
but the design question "what kind of work could protect it, and for whom?" This paper's
theory is constructed as an answer to this question. Section 4 examines the evidence on how
AI changes the cognitive content of work, and Section 5 presents Proposition 7, with cognitive
engagement as the mediating pathway, along with the set of propositions including dynamic
allocation to Gc-type roles. And the brain-health effect of work itself is registered not as an
assertion but as Hypothesis H1 on the map of verification in Section 9. This is as far as the
Working Paper | Ageless Management in the AI Era 37
evidence permits — and as far as the evidence permits, this section goes: that is the
conclusion of this section.
4. AI and Experience: What Is Compressed and What Is Not
The preceding sections confirmed that age-related change in cognitive abilities is not
monolithic (Section 2) and that the relationship between work and brain health depends on
the moderating variable of quality of work (Section 3). As the third pillar of this paper's
theory, this section organizes the empirical research on how generative AI is reorganizing the
economic value of "experience." There are three questions. First, what aspects of experience
does AI compress? Second, what does it not compress? Third, which abilities does dependence
on AI depreciate?
To anticipate the conclusion, the picture drawn by the evidence is as follows. For the
production of deliverables within AI's zone of competence, those with less experience gain
the most, and experience differences are compressed (Section 4.1). Outside AI's zone of
competence, by contrast, users' performance actually deteriorates, and human oversight of AI
output can fail systematically (Section 4.2). Furthermore, assisted performance gains do not
guarantee unassisted ability, and everyday dependence on AI can depreciate unaided skill
(Section 4.3). These three findings suggest the relative scarcification of the functions of
evaluation, verification, and oversight, rather than production. However, the "expertise"
measured in these empirical studies is domain experience and skill, not chronological age.
This conceptual separation is the pivot of this section (Section 4.4); it leads to the
reinterpretation of the headwind data surrounding older workers and AI (Section 4.5), and to
the empirical foundation — and the demarcation of the limits — of the definitions and
propositions presented in Section 5.
4.1 Empirical Evidence on Experience Compression
Empirical evidence on the effects of generative AI on job performance has accumulated
rapidly since 2023. Among the most credible findings is the compression of differences in
experience and ability. Brynjolfsson, Li & Raymond (2025) conducted a quasi-experiment
exploiting the staggered rollout of a generative AI conversational assistant among 5,172
agents at a customer-support firm. Average productivity, measured as resolutions per hour,
rose by 15%. The effect was not uniform, however. Workers with less experience and skill
showed improvements of roughly 30% across all productivity measures, while the most
experienced and highest-skilled workers saw only small gains in speed, with a slight decline
in conversation quality.
The degree of compression can be expressed in the language of the experience curve. AIusing
agents with two months of tenure performed on par with non-using agents with more
than six months of tenure. The authors interpret this as a shortening of the experience curve
by about four months. The mechanism has also been identified. The AI was trained on the
Working Paper | Ageless Management in the AI Era 38
conversation texts of top performers, and it presents as recommendations their tacit
behavioral patterns — clarifying questions, active listening, avoidance of escalation, and
adjustment of tone. After deployment, the communication patterns of low-skill workers
converged toward those of high-skill workers. What is being compressed here, in substance,
is the extraction and transfer of experts' tacit knowledge.
The study also reports two ancillary findings. First, even during AI system outages, workers
with longer AI experience — especially those who had followed the recommendations
faithfully — maintained higher productivity than their own pre-deployment levels. Part of AI
use can take hold as learning rather than mere dependence. This point is revisited in Section
4.3, paired with Budzyล et al. (2025). Second, after deployment, customer sentiment improved
and turnover declined (especially the retention of new hires). Note, however, that the study
covers routine tasks in a single firm and a single occupation, and age is not among the axes of
analysis.
Noy & Zhang (2023), in a preregistered online experiment with 453 college-educated
professionals, tested the effects on occupation-related writing tasks. The ChatGPT group took
40% less time and received quality ratings 18% higher. For this paper, the core is the change in
the distribution. In the control group, the correlation between performance on the first and
second tasks was 0.49, whereas in the treatment group it fell to 0.25. This is compression in
the sense that initial performance differences shrink by roughly half under treatment, and
participants with lower initial performance benefited more. In a follow-up survey two weeks
later (82% response rate), 33% of the treatment group had continued using the tool in their
actual work; among those with no prior experience, the figures were 26% in the treatment
group versus 9% in the control group (p=0.048). As a limitation, the tasks were one-off and
brief, and learning and long-term skill formation were not measured.
As a large-scale experiment in a near-field setting, Dell'Acqua et al. (2023) conducted a
preregistered field experiment using GPT-4 with 758 consultants at the Boston Consulting
Group (about 7% of the firm's individual-contributor level). On the 18 tasks designed to fall
within AI's capabilities ("inside the frontier"), tasks completed rose by +12.2%, speed by
+25.1%, and quality by more than 40% relative to the control group. The compression pattern
was replicated. Performers with below-average baseline performance improved by +43%
relative to their own baseline, and above-average performers by +17%. Gains are larger
toward the bottom, but the top gains as well; the summary "no effect at the top" is wrong. The
study's conceptual contribution lies in its formalization of the "jagged technological frontier":
the boundary between what AI does well and does poorly is hard to see in advance, and
discerning that boundary itself becomes a new human skill. This formalization leads directly
to the problem of oversight in the next subsection.
The same pattern appears in software development. Peng et al. (2023) randomly assigned
95 professional developers to treatment and control and reported that use of GitHub Copilot
shortened completion time on an HTTP-server implementation task to 71.17 minutes versus
Working Paper | Ageless Management in the AI Era 39
160.89 minutes (−55.8%). Heterogeneity analyses suggest larger benefits for developers with
fewer years of experience, developers with heavier coding loads, and older developers (aged
25–44). The study is an unrefereed vendor-affiliated study, with the limitations of a low
completion rate and a single task, but it replicates the compression-type result on the
experience axis while standing as one of the few data points suggesting that, on the age axis,
"the younger, the greater the gain" may not hold; it is referenced again in Section 4.4.
What is robust across these four studies is the pattern that, on tasks within AI's zone of
competence, the lower-performing and less experienced gain the most. Translated into the
language of economics, this is downward pressure on the experience premium in deliverable
production. The market value of the production-side differences that experience has
incrementally conferred — deliverable quality, working speed, procedural knowledge —
begins to decline once AI can substitute for and transfer them in units of several months.
Here the content of the compressed abilities must be examined precisely. What is being
compressed is not only Gf-type components. The behavioral patterns whose transfer
Brynjolfsson, Li & Raymond (2025) confirmed — clarifying questions, active listening,
avoidance of escalation, adjustment of tone — include Gc-type components belonging to
interpersonal judgment in the classification of Definition 2 (Section 5). That is, the mapping of
Section 2 — "what generative AI primarily substitutes for is the Gf-type load" — must be
qualified as a mapping in the context of production. In the context of production, so long as
they can be formalized as training data, even the interpersonal and contextual Gc-type
components become objects of transfer. The axis along which the boundary of compression is
drawn is therefore not the type of ability (Gf-type or Gc-type) but the distinction of function —
differences in production, or differences in verification. Proposition 1 of Section 5 (the
marginal value shift) formalizes, as the obverse of this downward pressure, the rise in the
relative marginal value of Gc-type tasks — above all evaluation- and audit-type tasks (those
that exercise the contextual interpretation and interpersonal judgment of Definition 2 in
verification). But what that "uncompressed function" is, and who possesses it, cannot be
derived from the compression evidence. That is the subject of the following two subsections.
Table 3 Compression of experience and ability differences by generative AI: key empirical studies
Study Task and context Main effects
Group benefiting
most Evidence grade
Brynjolfsson,
Li & Raymond
(2025)
Customer-support
work. 5,172 agents;
quasi-experiment
with staggered
rollout
Resolutions +15%. AIusing
agents with 2
months' tenure on par
with non-users with
over 6 months' tenure
(experience curve
shortened by about 4
months)
Workers with less
experience and
skill (about +30%).
Top tier: only small
speed gains, with a
slight decline in
quality
A (peerreviewed;
largescale
field quasiexperiment)
Noy & Zhang
(2023)
Occupation-related
writing. 453 collegeeducated
Time −40%; quality
+18%. Cross-task
performance
A (peerreviewed,
but a
Working Paper | Ageless Management in the AI Era 40
professionals;
preregistered online
experiment
correlation fell from
0.49 to 0.25
(compression of initial
differences)
Participants with
lower performance
on the first task
laboratory-style
task)
Dell'Acqua et
al. (2023)
18 consulting tasks
(inside the frontier).
758 BCG consultants;
GPT-4
Completions +12.2%;
speed +25.1%; quality
over +40%
Below-average
performers (+43%).
Above-average
performers also
gained +17%
B+ (workingpaper
preprint;
large-scale
preregistered
field
experiment)
Peng et al.
(2023)
HTTP-server
implementation task.
95 professional
developers; GitHub
Copilot; RCT
Completion time 71.17
vs. 160.89 minutes
(−55.8%)
Developers with
fewer years of
experience,
developers with
heavier coding
loads, and older
developers (aged
25–44)
B (unrefereed;
vendor-affiliated
study; single
task)
Note: All effects are relative to the control condition or to performers' own baselines. The figures for Noy & Zhang
(2023) follow the version published in Science. The figures for Dell'Acqua et al. (2023) follow the working-paper
version (an Organization Science version appears to have been published in 2025, but this paper uses the verified
working-paper figures). Evidence grades are this paper's rating (A = peer-reviewed; B = preprint or technical report; C
= grey literature; D = essay; +/− indicate relative rating within a grade).
In summarizing these results, the subject of compression must be delimited precisely. What
is compressed is the experience difference in deliverable production, under
contemporaneous assistance, within AI's zone of competence. Judging the zone of
competence, verifying the output, and internalizing the learning are each separate problems,
and, as the following two subsections show, the compression evidence does not carry over to
them. This delimitation is the empirical basis of the first half of Proposition 3 of Section 5 (the
non-compressibility of experience), namely that "contemporaneous AI assistance compresses
differences in production (deliverable quality, working speed, procedural knowledge)." The
second half — that it does not compress differences in verification (differences in experiential
audit capacity) — is a theoretical claim that at present lacks direct evidence of the same
standard, and its status is clarified in Section 4.4.
4.2 Oversight Failure
The compression story has a reverse side. Dell'Acqua et al. (2023) also ran tasks designed so
that AI would be prone to error ("outside the frontier"). On these tasks, the probability that
the AI-using group reached the correct answer was 19 percentage points lower than the
control group. It is not that quality fell by 19%; the probability of reaching the correct answer
fell by 19 points. The same population that gained over +40% in quality inside the frontier
clearly deteriorated outside the boundary. Because the boundary is hard to see in advance,
discerning how much to entrust to AI — that is, oversight of AI output — emerges as the
human-side function that separates outcomes.
Working Paper | Ageless Management in the AI Era 41
Can humans, then, perform that oversight well? Dell'Acqua's (2022–2023) solo-authored
working paper "Falling Asleep at the Wheel" offers cautionary evidence on this point. 181
professional HR recruiters each evaluated 44 résumés (7,964 in total; 5,184 after attention
checks), with the quality of the AI provided randomly assigned: a "near-perfect AI" disclosed
as roughly 99% accurate, a "high-quality AI" at 85%, a "low-quality AI" at 75%, and a no-AI
control. The central finding is that evaluators given the low-quality AI were more accurate
than those given the high-quality AI (accuracy improvement over control on a 10-point scale:
+0.314 for the low-quality AI versus +0.103 for the high-quality AI). The low-quality-AI group
examined each résumé for about 8.8–10 seconds longer and invested significantly more effort
(p<0.01). When the AI is understood to be excellent, humans do not raise their cognitive effort
— they "fall asleep at the wheel." This relationship is non-monotonic, however. The condition
disclosed as "near-perfect" performed best (+0.787), and the high-quality-AI group also
improved over the control. The accurate summary is therefore not "the better the AI, the
worse the human" but "human vigilance slackens most under an imperfect but highperforming
AI." Note that, at the time of verification, the study was a working paper not
published in a refereed journal, and this paper's figures are based on a mirrored distribution
copy (evidence grade B−). Even with this caveat, the finding that oversight effort is
endogenously determined by the perceived quality of the AI is consistent with the WP8
framework discussed below.
This finding attaches a condition to the compression evidence of Section 4.1. In
Brynjolfsson, Li & Raymond (2025), faithful adherence to AI recommendations was associated
with high performance because the tasks were routine and within AI's zone of competence.
The same adherence behavior flips to the −19-percentage-point side outside the jagged
frontier, and among evaluators whose oversight effort has slackened, it appears as degraded
detection accuracy. Whether "following the AI pays," in other words, depends on the task's
position, and the task's position is hard to see in advance. Precisely for this reason, the
function of doubting and verifying output cannot be switched off even in environments
where adherence performs well.
Human oversight failure is observed beyond individual studies, at the level of metaanalysis.
Vaccaro, Almaatouq & Malone (2024), through a systematic review and metaanalysis
of 106 experimental studies and 370 effect sizes, reported that human–AI
combinations on average performed below the best of human alone or AI alone (Hedges'
g=−0.23, 95%CI −0.39 to −0.07). The assumption that adding a human improves outcomes does
not, on average, hold.
WP8 of this series (Kadowaki 2026h) theorized this body of oversight-failure evidence
within the framework of "Human on the Loop (HOTL)." Five points of its skeleton are needed
for this section's argument. First, oversight is a consumed resource. Supervisors' attention
and cognitive effort are finite, and one cannot assume they are supplied constantly and free
of charge. Second, the solvency condition. The cognitive resources, time, and cost that can be
devoted to oversight have an upper bound, and oversight can be sustained only within that
Working Paper | Ageless Management in the AI Era 42
range. No solution exists that thickens oversight without limit. Third, alarm fatigue. The more
outputs and alerts there are to check, and the more false alarms are mixed in, the more
sluggish supervisors' responses become. Fourth, automation bias. Humans have a systematic
tendency to over-follow a system's suggestions, and the Falling Asleep at the Wheel findings
can be read as the effort-allocation version of this. Fifth, the value of oversight is determined
by independence × detection probability. That the supervisor's errors are independent of the
AI's errors (error decorrelation) is a necessary but not a sufficient condition; it becomes value
only when multiplied by the probability of actually detecting errors. In addition, WP8 points
out that there is an upper bound on the number of targets one supervisor can effectively
oversee (span), and that supervisor vigilance is harder to sustain in environments where
errors occur only rarely (the rarity effect); the Falling Asleep at the Wheel finding that
improvement in AI performance itself makes oversight harder is consistent with the latter.
WP8 further organizes, from case analyses, the point that homogeneous supervisor groups
trained in the same era as the AI share blind spots and have difficulty satisfying the
independence condition. This point becomes the starting point of this paper's Proposition 5,
which positions generational heterogeneity as an organizational source of decorrelation, and
of the design arguments of Section 6.
Here the relationship between automation bias and experience requires a theoretical
treatment of the reverse mechanism. The oversight-failure evidence tends to invite a reading
on which experience is an immunity to this failure mode. According to systematic reviews,
however, automation bias is widely observed in professional decision making, including
clinical decision support (Goddard, Roudsari & Wyatt 2012), and arises most readily on tasks
with high verification complexity and heavy cognitive load (Lyell & Coiera 2017). The
verification of contextual and practical risks — precisely the setting into which experiential
audit capacity should be deployed — falls under this condition. On the other hand, that expert
judgment can be systematically distorted by prior expectations and contextual information
has been shown by pre-AI expertise research. Dror, Charlton & Péron (2006) reported that
when five fingerprint experts with an average of 17 years of experience were re-presented
with fingerprint pairs they themselves had previously identified as matches, together with
context suggesting a case of misidentification, four of the five reversed their own past
judgments. With the limitation of being a preliminary report on a small number of cases, it is
suggestive on the point that long professional experience confers no immunity to
confirmatory contextual distortion.
Superimposing these two lines of findings yields the following theoretical possibility. When
AI output is plausibly constructed and also fits the auditor's own past successes and industry
received wisdom, confirmation bias and automation bias can compound. The experienced
auditor's mental model processes convention-conforming output fluently, and that processing
ease is readily confused with a sense of "already verified." Because verification effort has
already been curtailed by automation bias, no additional cross-checking is triggered. The
consequence is that the more experienced the auditor, the more he or she may uncritically
Working Paper | Ageless Management in the AI Era 43
accept plausible AI output that fits his or her own mental model. Ironically, the
inexperienced, who lack the conventional schemata against which to check, are relatively less
liable to this particular trap. That is, the claim of the audit value of experience operates most
strongly when the error contradicts the auditor's received wisdom, and can break down
when the error conforms to it. However, no study directly measuring the compounding of
confirmation bias and automation bias in the auditing of AI output could be confirmed within
the scope of this paper's search, and this treatment is theoretical. Proposition 4 of Section 5
writes this possibility into the body of the proposition as boundary condition (i), and
Hypothesis H3, by embedding in the task both errors that conform to the auditor's industry
received wisdom (confirmation-conforming) and errors that contradict it (deviant), tests this
boundary simultaneously with the main effect of experience (Section 9, Appendix A).
Alongside the compounding with confirmation bias, a route by which the surface
properties of the object of verification paralyze the audit is also to be anticipated from classic
findings of cognitive psychology. According to research on processing fluency, the more
subjectively easy a stimulus is to process, the more people judge its content to be true. That
manipulation of perceptual fluency alone raises truth judgments of sentences has been
shown experimentally (Reber & Schwarz 1999), and the finding that fluency broadly elevates
judgments of truth, liking, and confidence has been organized by a systematic review (Alter &
Oppenheimer 2009). Generative AI output — grammatically well-formed, stylistically
polished, and low in processing resistance — sits at this high-fluency pole. The detection
trigger of experiential audit capacity is the cognitive dissonance generated by mismatch
between accumulated schemata and the output, but high-fluency output raises the very
threshold of that dissonance, so the detection mechanism can be bypassed without any deficit
in the auditor's capacity. Fluency bias, that is, is a second paralysis route for auditing,
operating independently of the confirmation-bias route that originates in the auditor's
internal priors. Proposition 4 of Section 5 writes this into the body of the proposition as
boundary condition (iii), and Hypothesis H3, by controlling the fluency of the task documents
(high fluency versus low fluency) as a factor or covariate, tests this boundary within the same
experiment as the main test (Section 9, Appendix A).
Let this paper's position be stated explicitly here. This paper does not claim that "humans
are good at oversight." As this section has shown, the evidence points rather to the failureproneness
of human oversight. This paper's claim is confined to a single point. As WP8
showed, so long as a non-delegable residual of oversight and verification exists — a residual
that cannot be delegated to AI — the supply side of oversight, a scarce and consumed
resource, must be interrogated. As a candidate for that scarce input, this paper theorizes
experiential audit capacity (Section 5, Definition 4). Whether experiential audit capacity
actually raises detection probability, however, is untested and is the object of Hypothesis H3
(Section 9). Super-seniors' participation in auditing is likewise subject to WP8's capacity
constraints, including the solvency condition and alarm fatigue. The design of the individuallevel
version of these constraints is Brain Safety in Section 8.
Working Paper | Ageless Management in the AI Era 44
4.3 The Divergence Between Assisted and Unassisted Performance
The compression evidence (Section 4.1) measured performance under AI assistance. What
remains when the assistance is removed? The most direct answer to this question comes from
Bastani et al. (2025), a randomized controlled trial within mathematics classes covering
roughly 1,000 students (grades 9–11) at a large high school in Turkey. Students were assigned
to three groups: a control group with textbook and notes only, a GPT Base group using a raw
ChatGPT-type interface, and a GPT Tutor group using teacher-designed guardrailed prompts.
During assisted practice sessions, performance improved substantially: +48% over control for
GPT Base and +127% for GPT Tutor. On the unassisted exam with access removed, however,
the GPT Base group performed significantly worse than control by 17%, and the GPT Tutor
group was statistically indistinguishable from control. The guardrails eliminated the harm
but generated no gain. In the mechanism analyses, the messages of GPT Base students were
dominated by the "what is the answer" type, which the authors call the "crutch" effect.
Answer-copying use bypasses conceptual learning and collapses in unaided settings. A
correction to the paper has been published, but its content is solely a fix to the authors'
affiliation information, with no change to the result figures above (correction content verified
on August 21, 2026). It should be kept in mind that extrapolation from high-school
mathematics to workplace skill formation remains an analogy.
For occupational skill, the same type of concern is shown by Budzyล et al. (2025), a
retrospective observational study at four endoscopy centers in Poland comparing the
adenoma detection rate (ADR) of standard non-AI colonoscopies before and after the
introduction of AI-assisted colonoscopy (1,443 procedures in total: 795 before introduction,
648 after). ADR fell from 28.4% (226/795) before introduction to 22.4% (145/648) after, an
absolute difference of −6.0 percentage points (95%CI −10.5 to −1.6, p=0.0089), and AI exposure
was an independent factor in the ADR decline (odds ratio 0.69, 95%CI 0.53–0.89). Being a
retrospective observational study, it cannot establish causation and goes no further than
"suggestion." Yet it is the first clinical suggestion that everyday dependence on AI can
depreciate, on a timescale of months, the unaided skill that can be exercised without AI. The
detection skill of expert endoscopists had been considered a paradigm of the kind of
perceptual and contextual skill that AI cannot compress. Even that "uncompressible ability"
can depreciate if it is not exercised.
At first glance, this finding contradicts the outage analysis of Brynjolfsson, Li & Raymond
(2025) (Section 4.1), in which AI use takes hold as learning. The two can, however, be read
integratively as a difference in task structure and AI design. The endoscopy AI substitutes for
the detection of lesions itself; the human detection skill lies dormant unexercised and
depreciates. The customer-support AI presents suggested responses, and adoption and
execution remain with the human. Because exercise continues, learning through imitation
can occur. That is, "does AI aid learning or dissolve skill" is not an either-or rule of thumb but
depends on the design variable of which of the human's cognitive processes the AI takes over
Working Paper | Ageless Management in the AI Era 45
and which it leaves on the human side — such is this paper's reading. This treatment is itself
hypothetical; no direct comparative experiment exists.
WP5 of this series (Kadowaki 2026e), within its account of the formation of brain capital
(Brain Capital), proposed the institutional securing of "protected unassisted practice," that is,
practice opportunities from which AI assistance is deliberately removed. The guardrail-group
results of Bastani et al. (2025) and the depreciation suggestion of Budzyล et al. (2025) are
empirical backing for the concern behind that proposal, and at the same time demand an
extension of its scope of application. Depreciation risk is not confined to young people in the
course of learning. What must be stressed in the context of Ageless Management is that the
same logic operates on super-seniors' crystallized intelligence (Gc) and on the experiential
audit capacity that stands upon it. If audit and verification roles are designed to depend on AI
pre-screening, the auditors' own detection skills depreciate, and the source of oversight value
can be undermined. The role design of the multigenerational ecosystem must therefore build
unassisted exercise opportunities into seniors' Gc-type roles as well (Section 8). Whether
protected unassisted practice is effective for maintaining experiential audit capacity,
however, is untested, and this paper hands it over to the measurement framework of Section
9 as a design hypothesis to be tested.
It should be added that this logic of depreciation applies to this paper's resource-conversion
thesis itself. The "held but idle Gc-type abilities and experiential audit capacity" assumed by
Propositions 2 and 11 (Section 5) are, for as long as a period without exercise continues after
exit, exposed to depreciation and obsolescence by the logic of this subsection. That is, the Gc
and detection abilities of super-seniors long after exit are not a reservoir sleeping intact.
Domain-contextual knowledge becomes obsolete as institutions, technologies, and markets
change, and detection skill lacking exercise has no reason to escape the depreciation pathway
that Budzyล et al. (2025) suggested even for professionals in active practice. The resourceconversion
claim is therefore conditioned on time elapsed since exit. Those who have
recently exited and those long after exit cannot be treated alike, and for the latter, estimating
depreciation with years since exit as a variable — including the possibility of recovery
through re-exercise — becomes an independent empirical task (see the commentary on
Proposition 2 and Section 10.4).
The divergence between assisted and unassisted performance also has implications for the
institutions of measurement and evaluation. What an organization can observe is, in many
cases, only the deliverables produced under AI assistance. Looking only at the practicesession
performance in Bastani et al. (2025) (+48%, +127%), one cannot see the deterioration
on the unassisted exam (−17%). Likewise, looking only at the paperwork of AI-assisted audit
work, the depreciation of the auditor's own detection skill remains invisible until an event
occurs in which the AI drops out. Measuring assisted performance and unassisted ability
separately is therefore not a theoretical taste but a requirement of risk management. That the
experimental design for Hypothesis H3 (Section 9, Appendix A) incorporates both the AIassisted
and unassisted conditions as factors, and measures the detection of surface errors
Working Paper | Ageless Management in the AI Era 46
and of contextual risks separately, is intended to capture this divergence at the level of
measurement design.
4.4 Conceptual Separation: Domain Experience and Chronological Age
Before connecting this section's evidence to the debate on older workers, the conceptual
distinction on which the success or failure of this paper's entire argument turns must be
made explicit. The "experts" in the prior empirical studies are operationalized by the level of
domain experience and skill, not by chronological age. The axis in Brynjolfsson, Li &
Raymond (2025) is tenure and skill measures. Noy & Zhang (2023) stratify by first-task
performance, and Dell'Acqua et al. (2023) by baseline performance. The published abstract of
Budzyล et al. (2025) includes no analysis by endoscopists' age or years of experience. That is,
there is almost nothing this section's evidence can say directly about "older workers."
The inference "AI helps the less experienced; therefore the experience premium of older
workers is lost (or preserved)" is accordingly doubly short-circuited, whichever direction it
takes. First, the correlation between experience and age varies greatly across occupations and
job-change histories. Second, older novices (career changers and returners) and young
experts (early specializers) exist in large numbers in reality. For this reason, this paper treats
the experience premium, the judgment-and-oversight premium, and chronological age as
three separate variables. As one of the few data points concerning age, Peng et al. (2023)
reports larger benefits for "older developers" alongside developers with fewer years of
experience, but this "older" is a matter within the 25–44 range, not data on those aged 50 and
over. It is a fragment showing that a simple "the younger, the greater the gain" gradient on the
age axis is not self-evident, but it cannot be extrapolated to the debate on older workers.
This distinction can be made intuitive by a thought experiment. A person who moves to a
different industry at 60 is high in chronological age but shallow in experience of that domain,
and belongs to the side this section's evidence identifies as "benefiting most." Conversely, an
expert in his or her 30s who has trained in a single domain since the teens is low in
chronological age but is a holder of long-term domain experience. If what determines the
competence of oversight and verification is experience, this expert in the 30s can possess high
experiential audit capacity, and the 60-year-old career changer does not. This paper's
framework accepts this consequence. Indeed, it is precisely because it accepts this
consequence that Proposition 4 can specify as its own refutation condition that "chronological
age independently predicts detection ability even after controlling for years of domain
experience." If age itself retains independent explanatory power, this paper's reattribution
thesis is rejected.
On that basis, this paper's claim is made explicit in two stages. The first stage is the
empirically established range. Heterogeneity of outcomes under AI use has been observed
along the axis of domain experience and skill. That discerning the frontier and verifying the
output are emerging as new human-side skills is also consistent with the evidence of Sections
Working Paper | Ageless Management in the AI Era 47
4.1–4.2. However, how far expertise improves that competence in verification and oversight is
itself still an open question. A report of a randomized labeling experiment said to show that
even experts miss AI's errors has appeared in PNAS Nexus (verified by title and journal only:
Why do experts miss AI's errors? Evidence from a randomized labeling experiment. PNAS
Nexus, 5(6), pgag146. The author names and content are unverified, and the bibliographic
details are under verification. This paper therefore confines itself to a footnote-level mention
of its existence and does not include it in the reference list). In addition, as organized in
Section 4.2, there is the reverse theoretical possibility that when errors conform to the
auditor's received wisdom, experience may instead impede detection through the
compounding of confirmation bias and automation bias. The audit advantage of expertise, if
it exists, is conditioned on the content dimension of the error. The second stage is this paper's
theoretical claim, and it is untested. Section 5 first defines the ability to detect the contextual
errors, practical risks, and ethical risks contained in AI output operationally — by detection
performance on verification tasks, not by attributes or rank (Definition 4, experiential audit
capacity). The definition itself does not include what forms this capacity. The claim that
individual differences in this detection ability are formed by the interaction of long-term
domain experience with Gc-type abilities and metacognition is presented, as an empirical
claim independent of the definition, in the form of Proposition 3 (the non-compressibility of
experience) and Proposition 4 (the formation of experiential audit capacity), with refutation
conditions, and is directly tested by the 2×2 factorial design of Hypothesis H3 (Section 9,
Appendix A).
Between these two stages lies an unfilled empirical gap. No study that operationalized age
or tenure and measured performance in AI oversight or verification could be found within
the scope of this paper's search. Not only the relationship between experience and AIoversight
ability, but direct evidence on the age axis is missing, and it is this absence that
grounds this paper's research-gap claim.
This subsection's conceptual distinction faces one more anticipated objection. If the
substance of experiential audit capacity lies in long-term contextual memory — the retention
of past failure cases, organizational history, and the history of customer relationships — then,
the objection runs, once expanding context windows and retrieval augmentation make it
possible to inject past data in bulk, Gc's contextual memory would likewise be substituted by
AI, and experiential audit capacity would move over to the compressed side. The response
has two stages. First, this objection mistakes the core of experiential audit capacity for the
retained volume of recorded past data. A substantial part of the judgment material at work in
auditing is never recorded in the first place — the subtleties of interpersonal relationships,
the emotional history of an organization, and value judgments never made explicit either lose
most of their information at the moment of documentation or are never documented at all.
Moreover, even where records exist, what the auditor is discerning is the semantic gap
between the described data and living reality — the mismatch whereby a description that is
consistent on paper diverges from actual conditions on the ground. However much the
Working Paper | Ageless Management in the AI Era 48
volume of data injected into the context is expanded, what the machine is given is only the
description side, and this gap is not closed. As a limitation at a different level from this, it has
also been reported that the use of injected long contexts is itself imperfect — language model
performance degrades significantly when the relevant information sits in the middle of a long
input context (lost in the middle: Liu et al. 2024). This performance limit, however, may be
alleviated by generational model turnover, and this paper does not place it among the main
pillars of its response.
Second, the more fundamental response lies at the level of the independence of
verification. Even if the substitution of contextual memory were to succeed completely — that
is, even if the same model read the whole of its own vast context and verified its own output
— that act supplies no independence of the interpreter. As WP8 (Kadowaki 2026h) formalized,
the value of oversight is independence × detection probability, and a verifier that shares its
error distribution with the generator exhibits systematic false negatives toward the
generator's own systematic blind spots. The same model reading its own context is, in
precisely this sense, a configuration in which generation and verification share an error
distribution, and no increase in the volume of contextual memory increases the
independence of verification. The volume of contextual memory and the independence of
verification are separate problems. What the expansion of context windows can compress is
therefore the function of retention and recall of memory in the context of production, not the
function of an independent interpreter in the context of verification — the boundary line
drawn at the end of Section 4.1, between differences in production and differences in
verification, is drawn in the same place against this objection as well.
Search record: as of August 21, 2026. Via WebSearch (in English), three queries — (1) study older
workers experience advantage supervising AI oversight expertise age empirical evidence experiment;
(2) "domain expertise" "AI oversight" OR "human oversight" experiment experts better detecting AI
errors age older professionals; (3) seniority tenure moderates ability to verify AI output errors
experiment "years of experience" LLM verification randomized — together with derivative searches
during the verification of individual references, were run, but no empirical study operationalizing age
or tenure and measuring AI-oversight performance could be confirmed. Adjacent studies (the expertiseaxis
miss experiment, the heterogeneity within the 25–44 band, etc.) are noted in the main text.
Finally, the bridging assumption under which this paper speaks of age is made explicit. Age
is merely a function of the time that makes possible the accumulation of long-term domain
experience, and the oversight value is attributed not to age but to experience. Even if holders
of experiential audit capacity are expected to be relatively numerous among super-seniors
(aged 60 to their 90s), that is a distributional tendency, and it does not mean that an
individual's ability may be inferred from his or her chronological age. It is precisely for this
reason that this paper's framework is "ageless" — independent of age — and not a celebration
of seniors (Section 1.4, Section 10).
Working Paper | Ageless Management in the AI Era 49
4.5 Headwind Data on Older Workers and AI, and Their Conversion
With the concepts separated, consider the data that do currently exist on the age axis. They
are mainly disparities in use, adoption, and opportunity, and their direction is a headwind.
The most credible is the representative repeated survey of Bick, Blandin & Deming (2024)
(NBER working paper, evidence grade B+). As of late 2024, roughly 40% of Americans aged 18–
64 used generative AI, an adoption speed exceeding that of the PC and the internet.
Workplace usage rates, however, show a clear age gradient: 34.5% at ages 18–29, 34.6% at 30–
39, and 29.5% at 40–49, against 16.7% at 50–64 — those in their 50s and above stand at
roughly half the level of those under 40. Adoption rates are moving rapidly: in the survey's
updated figures (August 2025 survey), overall adoption reached 54.6% (+10 percentage points
year over year) and workplace use 37.4%. Explicit statement of the survey date is
indispensable for adoption figures.
The grey-literature surveys are also directionally consistent. The five-country survey by the
employment-support NGO Generation (France, Ireland, Spain, the UK, and the US; published
October 2024) covers 2,610 entry- and mid-level workers aged 45 and over and 1,488
employers. In hiring for AI-related roles, 90% of US recruiters said they would "consider
candidates under 35," while only 32% would "consider those over 60" (86% versus 33% in
Europe). Also, only 15% of those aged 45 and over used generative AI at work. This 90%
versus 32%, however, is a statement of recruiters' intent, not an audit experiment measuring
hiring outcomes. In the AARP survey (conducted March 2026; 1,015 US workers aged 50 and
over), familiarity with workplace AI stood at 52%, use in daily work at 23%, and experience of
AI training at only 12%, while 49% wanted to learn — an observed 37-point "gap between
willingness and training opportunity." As a policy-oriented commentary, Pizzinelli & Tavares
(2026) point out an "asymmetry of opportunity and risk": older workers are relatively more
likely to hold jobs with high AI exposure and high complementarity and can benefit from AI
in jobs requiring experience, judgment, and interpersonal skills, but because their labormarket
mobility is low and the costs of job change and retraining are high, the blow is greater
should they end up on the substituted side.
Evidence-grade note: the Generation survey (C) and the AARP survey (C) in this subsection are grey
literature not subject to peer review (an NGO-commissioned survey and a membership-organization
survey), and the figures rest on intent and self-report. Pizzinelli & Tavares (2026) is a policy commentary
(C+), and citation is confined to conceptual points. None of these are measurements of ability or
outcomes. Bick, Blandin & Deming (2024) is an NBER working paper (B+), and publication in a refereed
journal is unconfirmed.
How to read these headwind data is the fork in the road for this paper. First, the age
gradient in usage rates is evidence neither of ability differences nor of benefit differences.
The disparities are confounded with differences in opportunity, training, and job
composition, and AARP's 37-point gap suggests a shortfall on the supply side (training
opportunity), not the demand side (willingness to learn). Second, none of the surveys
Working Paper | Ageless Management in the AI Era 50
measures the quality of use, that is, supervisory use versus substitutive use. Against this
section's body of evidence, this is a decisive omission. Third, from the standpoint of the
cognitive bottleneck (Section 5, Definition 3), this headwind can be explained as the
consequence of current job design and training allocation tacitly presupposing young,
execution-type (Gf-type-component-centered) AI use. If so, the prescription is neither to keep
older workers away from AI nor to hurry their assimilation into execution-type use, but a
conversion of role design: access to Gc-type, audit-type roles and the corresponding
reallocation of training and opportunity (Section 6). For this paper, the headwind data are not
a refutation but a given that demonstrates the need for conversion.
Moreover, this section's body of evidence shows that this headwind carries a danger of selffulfillment.
By the logic of Section 4.3, abilities that are not exercised depreciate. If older
workers continue to be kept away from opportunities for AI use, from audit-type roles, and
from training, then what was initially a mere opportunity gap can, with the passage of time,
convert into a real skill gap — for the discernment of AI's zone of competence and the
verification of its output are honed only within the experience of engaging with AI.
Conversely, a design that places holders of experience into audit-type roles and gives them
opportunities for exercise puts existing Gc-type assets to work and at the same time prevents
their depreciation. Whether to leave the headwind data as a given or to convert them through
role design is an organizational choice variable. It must be noted again, however, that the
effect of this conversion itself belongs to this paper's untested propositions (Propositions 2
and 11).
The direction of role conversion itself has been articulated in anticipatory form by
practitioners. Nielsen argues that "wise winnowing" — a division of labor in which generative
AI generates large numbers of options and elders holding accumulated evaluative capacity
select the best of them — can extend the productive careers of older knowledge workers.
This, however, is an essay by a prominent practitioner (evidence grade D+) and contains no
new empirical data whatsoever on AI and older users. This paper positions it not as evidence
but as a pioneering statement of a testable hypothesis. What this paper does from Section 5
onward is to formalize this intuition into the refutable form of the conceptual apparatus of
Definitions 4–6, Propositions 3–5, and Hypotheses H2–H3.
To summarize this section. (i) The compression of experience differences is empirically
established, but its subject is differences in deliverable production within AI's zone of
competence. (ii) Oversight is a scarce, consumed resource and can fail systematically.
Experience is not an immunity to this failure mode and, for convention-conforming errors,
can even become an impediment. This paper does not presuppose human skill at oversight; it
interrogates the supply side of oversight. (iii) Assisted performance does not guarantee
unassisted ability, and abilities that are not exercised — including Gc and experts' detection
skills — can depreciate. (iv) All of the above evidence lies on the experience-and-skill axis,
and direct evidence on the age axis is missing. To fill this gap, the next section presents the
Working Paper | Ageless Management in the AI Era 51
group of definitions centered on experiential audit capacity and the group of propositions
equipped with refutation conditions.
5. Theory: Definitions and Propositions
5.1 The Architecture of the Theory
Building on the evidence organized in the preceding sections, this section constructs the
theory of Ageless Management as eight definitions and twelve propositions. This paper is a
conceptual and measurement-proposal paper, and most of its propositions are untested. The
discipline of this section therefore rests on two commitments. First, every proposition carries
a refutation condition, so that the proposition itself states what observations would reject the
theory. Second, the commentary attached to each proposition, together with Table 7 (the map
of propositions), makes explicit which evidence from which section supports each
proposition and where this paper's untested claims begin. Not allowing established findings
to be conflated with this paper's theoretical wagers is the design policy of this section.
The core of the theory consists of the three theoretical operations announced in the
introduction (Section 1.4). The first operation is the reattribution from age to experience. This
paper does not adopt the popular argument that attributes the oversight value of older
workers to "age" — the argument that seniors are suited to auditing because they are
experienced. The oversight value is attributed to experiential audit capacity (Definition 4),
operationally defined by detection performance, and Proposition 4 attributes its formation to
the interaction between long-term domain experience and Gc-type abilities and
metacognition. Age is merely a function of the time that makes accumulation possible. It is
precisely this reattribution that makes this paper's framework "ageless" rather than a
celebration of seniors. As confirmed in Section 4.4, what prior empirical studies measured as
"expertise" was domain experience and skill, not chronological age, and this reattribution is
consistent with the very structure of the evidence. Definition 4 and Propositions 3 and 4 carry
this operation.
The second operation is to formalize generational heterogeneity as a source of "error
decorrelation." WP8 (Kadowaki 2026h) formalized the value of AI oversight as independence
× detection probability and generalized the value of independent oversight channels to "error
decorrelation." This paper positions the property that the overlap of judgmental blind spots is
small among individuals who have experienced different historical environments,
technological generations, and failure cases (Definition 5) as the organizational source of this
decorrelation. The starting point is the observation, drawn from WP8's case analyses, that
homogeneous supervisor groups sharing the same contemporaneous training and
information environment as the AI are prone to shared blind spots. This extension to the
generational axis, however, is a new claim of this paper that WP8 itself did not make, and it is
untested. Definition 5 and Proposition 5 carry this operation.
Working Paper | Ageless Management in the AI Era 52
The third operation is the conversion of social problems into resources. The three social
problems of population aging, constrained youth participation, and the exclusion of
marginalized groups are converted, conditional on AI-mediated complementarity (Definition
6) and Brain Safety (Definition 8), into untapped sources of brain capital (Brain Capital) =
stock (K) × utilization rate (u). What must be emphasized is the conditional side. This paper
does not claim unconditional conversion. The condition-dependence of employment effects
shown in Section 3, the risks of oversight failure and skill depreciation shown in Section 4,
and the empirical baseline that the average effect of age diversity is near zero (detailed in the
commentary on Proposition 6) all forbid the simple claim that "participation generates value."
Definitions 6–8 and Propositions 6–11 carry this operation.
The correspondence between evidence and propositions can be sketched in advance. The
cognitive-science foundations of Section 2 — the divergence of Gf and Gc, the heterogeneity of
peak ages across abilities, and the overlap of distributions across age groups — are the
empirical underpinning of Definitions 2 and 3, and they supply the premises of Proposition 2
and the grounds of Proposition 12. The evidence on work and brain health in Section 3 shows
the deadlock of average effects and the emergence of "quality of work" as a moderating
variable, motivating Proposition 7. The empirical work on AI and experience in Section 4
gives direct evidence for the first half of Proposition 1 and the first half (the compression
side) of Proposition 3. Up to this point is the territory supported by independent empirical
research. By contrast, the existence claim (a) and removal effect (b) of Proposition 2, the
second half of Proposition 3, the reattribution of Proposition 4, the generational decorrelation
of Proposition 5, the AI-mediated moderation of Proposition 6, the bidirectional capital
formation of Proposition 8, the transferability of Proposition 10, the resource conversion of
Proposition 11, and the allocation comparison of Proposition 12 are untested theoretical
propositions newly asserted by this paper, and the means of testing them are Hypotheses H1–
H3 in Section 9 (Table 7). This paper does not prove that the theory of this section is correct.
What this section does is fix the theory in a form in which its correctness can be adjudicated.
5.2 Definitions
The following eight definitions are held fixed throughout this paper. Immediately after each
definition, a note records why the definition was chosen and what it excludes.
Definition 1 (Ageless Management)
A management regime that removes the variable of chronological age from decisions on
role allocation, evaluation, participation, and exit, and that dynamically allocates roles
— within an ecosystem not limited to the boundary of employment — on the basis of
measured cognitive characteristics, accumulated domain experience, health status, and
the person's own intent.
Working Paper | Ageless Management in the AI Era 53
The core of this definition lies in pairing the removal of age with the specification of
replacement variables. As confirmed in Section 2, the peak ages of cognitive abilities are
scattered across decades depending on the ability, and the distributions of age groups overlap
widely. Under this evidentiary situation, chronological age is no more than a coarse proxy for
cognitive characteristics, and if the proxy is to be removed, the variables on which decisions
rest must be explicitly substituted. Of the four variables specified by Definition 1, cognitive
characteristics and domain experience are on the ability side, while health status and the
person's own intent are on the constraint-and-preference side; the explicit inclusion of the
latter two is meant to avoid presupposing that one can work or wants to work (a structural
response to the survivorship-bias critique in Section 10).
It is worth making explicit what this definition excludes. Extending the mandatory
retirement age (teinen) or raising the age ceiling of continued employment is an operation
that moves the threshold while retaining the variable of age, and is not Ageless Management.
Preferential treatment of a particular age bracket, such as "promoting the active engagement
of seniors," also falls outside the definition, because it uses age as an allocation variable. On
the other hand, because the definition demands measurement-based allocation, it
internalizes the danger of measurement misuse — the danger that the measurement of
cognitive characteristics turns into a new apparatus of selection and discrimination. This
danger is a cost the definition accepts, and it is addressed explicitly in the refutation
condition of Proposition 12 and the self-critique of Section 10.
Definition 2 (Gf-type tasks / Gc-type tasks)
Tasks whose performance depends primarily on fluid intelligence (processing speed,
working memory, and the learning of novel procedures) are called Gf-type tasks; tasks
whose performance depends primarily on crystallized intelligence (accumulated
knowledge, contextual interpretation, and interpersonal judgment) and on
metacognition are called Gc-type tasks. This classification is a relative weighting on a
continuum, not a binary.
This definition transfers the factor distinction of Cattell-Horn theory (Section 2.1) from
persons to the attributes of tasks. There are two reasons for the transfer. The first is the
mapping onto the functional characteristics of AI. What generative AI primarily substitutes
for is Gf-type cognitive load — search, summarization, documentation, and the execution of
routine procedures (Section 4) — and describing tasks in terms of Gf/Gc allows the domain AI
can substitute for and the functions remaining on the human side to be discussed in the same
vocabulary. The second is the avoidance of typing individuals. Definition 2 classifies tasks, not
people. The rereading "older people = Gc-type personnel" is a variant of the age stereotyping
this paper criticizes, and it is explicitly excluded, together with the definition's continuum
clause — not a binary. Real jobs are bundles of Gf-type and Gc-type components, and the
structure of this bundle is precisely the subject of the next definition, Definition 3.
Working Paper | Ageless Management in the AI Era 54
Definition 3 (Cognitive bottleneck)
Given that a job is a bundle of Gf-type and Gc-type components, the state in which
difficulty in performing the Gf-type components forces exit from the job as a whole, so
that the Gc-type abilities the person holds never reach the work.
Definition 3 is a concept for describing the job exit of older workers not as a wholesale loss of
ability but as a consequence of an institutional structure — how jobs are bundled. Given the
Gf/Gc divergence confirmed in Section 2 — the coexistence of declining components and
maintained or growing components — difficulty in performing the Gf-type components of a
bundle can force exit no matter how well the remainder of the bundle could be performed.
What is lost then is not only the individual's income. The organization, too, is relinquishing
an asset — the Gc-type abilities held — while leaving it idle. Definition 3 is a description of a
"state"; the claim that this state exists as an independent exit pathway is not the definition but
Proposition 2(a). Nor does this definition restrict the cause of the bottleneck to aging.
Circumstances that constrain the performance of Gf-type components — illness, disability,
career interruption for childcare or caregiving — exist regardless of age, and Definition 3
treats them as the same structure. It is this commonality of structure that allows Ageless
Management to be discussed in continuity with the participation of socially marginalized
groups (Proposition 11).
Definition 4 (Experiential Audit Capacity)
The capacity to detect contextual errors and practical and ethical risks contained in AI
output. This paper defines it operationally, not by attributes or position, but by detection
performance on verification tasks. (Note: what forms this capacity is not part of the
definition. The formation mechanism is the empirical claim of Proposition 4.)
Definition 4 is the central concept of this paper and embodies two constructional choices. The
first is purification into an operational definition. The definition specifies the detection
capacity solely by a measurable quantity — detection performance on verification tasks —
and does not build into the definition what forms the capacity: domain experience, Gc-type
abilities, metacognition, or chronological age. This division of labor is a structural choice to
protect the falsifiability of the theory. If the formation mechanism (the interaction of
experience × Gc × metacognition) or age-independence were written into the definition itself,
Proposition 4 would become analytically true the moment the definition was adopted, and its
empirical content would be emptied. Moreover, if a result emerged in which chronological
age predicted detection capacity even after controlling for years of experience, an escape into
the definition — "what was measured was not the experiential audit capacity of Definition 4"
— would become available as a way of evading refutation. By purifying the definition into an
operational one and making the formation mechanism the empirical claim of Proposition 4,
Working Paper | Ageless Management in the AI Era 55
this escape route is sealed, and the variance decomposition of Hypothesis H3 — the
interaction term of experience × Gc × metacognition, and the partial effect of age after
controlling for experience — becomes, as it stands, the test of Proposition 4.
The second is the restriction of the detection targets. The definition restricts the targets of
detection to contextual errors and practical and ethical risks, and does not include surfacelevel
errors (errors of form, procedure, and document quality). The empirical work in Section
4 shows that differences in production are compressed by AI, and if the definition included
detection power in the domain being compressed, the concept would fail to capture the value
specific to experience. This restriction leads directly to the design of the dependent variables
of Hypothesis H3 (the separation of detection rates for surface-level errors and for contextual
risks). Note that nothing can be derived from the definition about whether this operationally
defined detection capacity is systematically higher among holders of long-term domain
experience, or whether it is independent of chronological age. Those are the empirical claims
of Propositions 3 and 4, and they are untested. Moreover, following the logic of Section 4.3,
experiential audit capacity is itself an asset that can depreciate if not exercised, and its
protection is a design object of Brain Safety in Section 8.
Definition 5 (Generational Decorrelation)
The property that, among individuals who have experienced different historical
environments, technological generations, and failure cases, the correlation of the error
distributions of judgment (blind spots) is lower than within a single generation.
Definition 5 defines generational heterogeneity not as values or attitudes but as a statistical
property of error distributions. This choice excludes two things. The first is generational
theory — categorical claims about generational traits of the form "Generation Z values X."
Effects specific to generational categories find little support in peer-reviewed research, and
this paper does not ask generational labels for explanatory power. What Definition 5 refers to
is the history of experience — exposure to different technological environments and different
failure cases — and the trace it leaves on the correlation structure of errors. The second is the
presumption of normative implications. Decorrelation is in itself neither good nor bad; it
acquires value in the context of oversight only by way of WP8's formalization — oversight
value = independence × detection probability. By placing the definition on a measurable
quantity, error correlation, Proposition 5 can have a direct refutation procedure: the
comparison of cross-generational and within-generational correlations of detection errors.
Working Paper | Ageless Management in the AI Era 56
Definition 6 (AI-Mediated Complementarity)
The state in which, through AI substituting for or complementing the Gf-type
components of a task bundle, heterogeneous cognitive assets that conventionally had to
be combined within a single individual (Gf-type execution capacity, Gc-type audit
capacity, and problem perception as a person directly affected) become connectable and
exchangeable across individuals. This connection is not a one-way prosthesis: it is
bidirectional in that the benefit of complementation (substitution for Gf-type
components) and the supply of auditing (provision of Gc-type verification) flow mutually
among participants (bidirectional complementarity).
Definition 6 carries this paper's claim to novelty (Section 1.4). The point is that AI is
positioned not as an agent but as a medium. As confirmed in Section 2.3, the idea of using AI
as a prosthesis for Gf-type components itself belongs to the lineage of assistive-technology
and prosthetics research and is not this paper's invention. A prosthesis is a one-way relation
that fills an individual's deficit. What Definition 6 captures is the structure beyond it: the state
in which the prosthesis loosens the conventional requirement that all components of a job
bundle be performed by the same individual, so that heterogeneous cognitive assets become
connectable across individuals. Of the three assets listed, "problem perception as a person
directly affected" refers to the problem-finding capacity that those with the experience of
illness, disability, or exclusion possess precisely because of that experience; it is the term that
allows the participation of socially marginalized groups to be treated as the connection of
assets rather than as accommodation (Proposition 11). The stipulation of bidirectionality at
the end anchors the introduction's (Section 1.4) core term "bidirectional complementarity" to
this definition, and is distinguished from the "bidirectional capital formation" of Proposition 8
— a claim at a different level, that brain capital increases across all participating strata. Note
also that Definition 6 is a definition of a state; it is distinguished from the claim that the state
produces outcomes (Proposition 6), and the latter is untested.
Definition 7 (Multigenerational ecosystem)
An organizational form that, not limited to the firm's own boundary of employment,
connects the cognitive assets of participants ranging from youth to super-seniors and
including socially marginalized groups, through multiple contractual forms such as
employment, outsourced engagement, advisory roles, PBL-based educational
partnerships, and NPO partnerships.
That Definition 7 lifts the restriction to the employment boundary is demanded by both
practice and theory. In practice, the participation of super-seniors (aged 60 to their 90s) and of
youth still in school more often takes the form of outsourced engagement, advisory roles, or
educational partnership than of full-time employment. In theory, the Experience Corps
Working Paper | Ageless Management in the AI Era 57
evidence confirmed in Section 3 comes from designed, non-employment roles, and if what
operates is the quality of the role, then the contractual form is not essential. This choice,
however, carries a heavy price. Outside the employment boundary lies a vacuum of labor-law
protection (Section 7), and an ecosystem built on Definition 7 can, if designed badly,
degenerate into unpaid labor and exploitation. This danger is built into the theory as
Proposition 9, and it is answered by pairing Definition 8's Brain Safety with it. Definitions 7
and 8 must not be used apart.
Definition 8 (Brain Safety)
A health-and-safety standard that protects the brain capital of ecosystem participants
from depreciation, comprising both directions: (i) protection from cognitive load and
brain fatigue, and (ii) protection from unpaid-labor conversion and exploitation arising
from asymmetries of bargaining power.
The distinctive feature of Definition 8 is that it binds protection from cognitive load (i) and
protection from exploitation (ii) into a single standard. (i) is the individual-level version of the
solvency condition WP8 imposed on supervisors — the cognitive resources that can be
devoted to oversight have an upper bound — and follows from the fact that oversight- and
audit-type roles are inherently roles of high cognitive load. The audit participation of superseniors
is subject to WP8's capacity constraints (including alarm fatigue and span
constraints), and Brain Safety redesigns these constraints as an individual-level health-andsafety
standard (Section 8). (ii) responds to Definition 7's lifting of the employment boundary.
Turning people into unpaid advisors with "purpose" or "social contribution" as a substitute
for compensation, and converting youth PBL into labor beyond its educational purpose, are
both expropriations of brain capital arising from asymmetries of bargaining power, and they
must be prohibited within the same system of standards as (i). A welfare-benefit reading of
Brain Safety, which regards protection as a discretionary benefit, is what this definition
excludes.
5.3 Propositions
The following twelve propositions constitute the entirety of this paper's theory. Each
proposition carries a refutation condition, and the commentary immediately following shows
its derivation and evidentiary status.
Working Paper | Ageless Management in the AI Era 58
Proposition 1 (Marginal value shift)
The diffusion of generative AI lowers the marginal cost of Gf-type tasks and raises the
relative marginal value of Gc-type tasks (contextual interpretation and interpersonal
judgment in the sense of Definition 2, and above all evaluative and audit-type tasks).
This rise is conditional on Gf-type production and Gc-type verification being
complementary in production, and on verification actually detecting errors
(Propositions 3 and 4).
Refutation condition: Systematic evidence that, in labor markets after AI adoption, demand for
or compensation premiums on evaluative, audit, and contextual-judgment tasks do not rise, or
that the relative compensation of Gf-type tasks rises persistently. The observation window is, as
a guideline, occupational wage and task-demand data covering roughly ten years from the fullscale
diffusion of generative AI.
The derivation of Proposition 1 is an application, to cognitive tasks, of the standard economic
logic that a fall in the price of a substitute raises the relative value of complements. The first
half — the fall in the marginal cost of Gf-type tasks — is directly supported by the empirical
work in Section 4.1. In customer support, writing, consulting, and software development
alike, the production of deliverables within AI's competence range was substantially
accelerated and leveled, with the largest gains accruing to the less experienced (Brynjolfsson,
Li & Raymond 2025; Noy & Zhang 2023; Dell'Acqua et al. 2023). This is evidence that the cost of
procuring the performance of Gf-type components from the market is falling rapidly. The
parenthetical in the proposition text refers to Definition 2's exemplification of the Gc type
(contextual interpretation and interpersonal judgment), with the evaluative and audit-type
tasks central to this paper's concern attached as a restriction. It does not introduce an
extension not appearing in Definition 2.
The second half — the rise in the relative marginal value of Gc-type tasks — is a theoretical
consequence of the first half, but the derivation involves an assumption that must be made
explicit: the assumption that Gf-type production and Gc-type verification are complementary
in the production function. The complementarity is assumed, not derived, and since
generative AI is itself absorbing summarization, evaluation, and verification functions
(Section 10.5), the possibility that the two turn into substitutes is not empty. That is why the
second sentence of the proposition text states this assumption as a condition of the
proposition. There is, however, a structural limit to the progress of this substitution. As WP8
(Kadowaki 2026h) laid out, AI self-verification shares its error distribution with the generator
and therefore does not constitute an independent verification channel — configurations
using an LLM as verifier show systematic false negatives for the types of error the generator
itself is prone to, and as long as generation and verification derive from the same training
distribution, their blind spots are correlated. Accordingly, improvements in AI's verification
capacity can erode the marginal value of Gc-type verification, but the erosion remains inside
Working Paper | Ageless Management in the AI Era 59
the correlated blind spots, and the niche of detecting systematic blind spots by verifiers who
do not share the generator's error distribution — decorrelated verification — remains. It is to
this niche that the rise in the marginal value of Gc-type verification asserted by Proposition 1
is ultimately anchored. This residual, however, is not a fixed quantity. The AI-side diversity
required by condition (iii) of Proposition 5 — cross-verification by models of different
architectures and developers — is an operation that lowers the error correlation of AI
verification on the machine side, and its progress works to narrow further the residual niche
of decorrelation supplied by humans. How much human verification value remains is an
open empirical question that depends on the speed of this machine-side decorrelation.
Furthermore, the very evidence this paper cited in Section 4.2 shows that adding verification
does not always generate value: human-AI combinations underperform the best of either
alone on average (the g=−0.23 of Vaccaro, Almaatouq & Malone 2024). If so, what Proposition
1 can assert unconditionally extends only to an increase in the demand for (necessity of)
verification, and the rise in marginal value is a claim conditional on verification actually
detecting errors — the validity of Propositions 3 and 4, that is, the effectiveness of detection.
Even in interpreting the refutation condition, the possibility cannot be excluded that
spending on verification labor increases while detecting no errors, and this conditionality
connects Proposition 1 to Propositions 3 and 4. Note that Proposition 1 is a claim about
relative value and does not imply a rise in the absolute wages of those engaged in Gc-type
tasks. Moreover, "the value of Gc-type tasks rises" and "who can perform them" are separate
questions, the latter being the subject of Propositions 3 and 4.
The "demand for and compensation premiums on evaluative and audit-type tasks" that the
refutation condition of Proposition 1 takes as its observable can be given a concrete
generative pathway from the economics of information. The collapse in production costs
brought by generative AI means the mass circulation of unverified artifacts, widening the
informational asymmetry of quality as seen by buyers — clients, readers, regulators. Akerlof
(1970) formalized, as the analysis of the market for lemons, that in markets where quality is
hard to discern, adverse selection can arise in which inferior goods drive out good ones. A
market flooded with unverified AI-generated artifacts approaches the conditions for this
lemons market. In that situation, certification of having passed verification by experiential
audit (HOTL) can function in the structure of Spence's (1973) signaling. Because the incentive
to bear the cost of verification and obtain certification is skewed toward suppliers of quality
that withstands verification, audit certification can support a separating equilibrium as a
signal of quality, and a trust premium on certified artifacts is realized as the market value of
verification labor — a pathway that gives a generating mechanism to the compensation
premium the refutation condition of Proposition 1 takes as its observable. This paper presents
this as a theoretical pathway, however, and does not assert that the premium will arise.
Whether the credibility of certification itself is maintained (the gaming of certification — a
problem of the same form as Proposition 12), and whether the premium exceeds the cost of
verification, are both open empirical questions, to be adjudicated within the observation
Working Paper | Ageless Management in the AI Era 60
window of the refutation condition. Note that the root of demand in this pathway is anchored
not in the performance limits of current-generation models but in the structure of
verification independence — demand for verification that does not share the generator's
error distribution (Section 10.5) — and does not disappear with the turnover of model
generations.
Proposition 2 (Removal of the cognitive bottleneck)
(a) The Gf-component bottleneck (Definition 3) exists as an independent job-exit
pathway alongside institutional compulsion such as mandatory retirement, health
constraints, and demand-side age discrimination. (b) AI's complementation of Gf-type
components removes this bottleneck and lets the Gc-type abilities a person holds reach
the work. The share of pathway (a) in total exits is not quantified in this paper and is
treated as an open empirical question. The validity of (b) has two boundary conditions.
First, unassisted cognitive exercise and the periodic insertion of unmediated audit must
be maintained — constant dependence on AI without them can, through the skilldepreciation
pathway of Section 4.3, erode over the medium to long term the very
foundations of the Gc-type abilities and metacognition that the complementation is
meant to serve. Second, complementation holds only within the range in which
standard cognitive screening does not fall below the threshold of mild cognitive
impairment — the decline of Gf can accelerate nonlinearly in old age, and this model
cannot complement a state in which attentional resources themselves are exhausted.
This lower bound is a boundary on allocation to experiential-audit roles, not a
constraint on participation in the ecosystem in general.
Refutation condition: For (a): systematic evidence that, in studies decomposing reasons for
exit, difficulty in performing Gf-type components is not observed as an independent reason for
exit. For (b): experimental or quasi-experimental evidence that, even after tools complementing
Gf-type components are provided, the job performance and job continuation of older workers
holding Gc-type abilities do not improve.
Proposition 2 is split into the existence claim (a) and the removal claim (b), each with its own
refutation condition. The reason for the split should be made explicit. This paper squarely
acknowledges that the broad pathways of job exit for older workers are institutional
compulsion — mandatory retirement and ceilings on continued employment (Section 7) —
health constraints (the HWLE evidence of Section 3), and age discrimination on the labordemand
side. Institutional compulsion operates regardless of an individual's Gf level, and
66.0% of Japanese aged 65 and over are retirees (Section 10.4). Accordingly, (a) does not claim
that the bottleneck is the main cause of exit or that a "substantial share" is attributable to it.
What it claims is only that Gf-bottleneck-induced exit exists as an independent pathway
alongside these known pathways, and the quantification of its share of total exits is, as the
proposition text states, an open empirical question. The test design for (a) is the
Working Paper | Ageless Management in the AI Era 61
decomposition of exit reasons. In surveys of leavers and panel data, exit reasons are
decomposed into institutional compulsion, health, discrimination, and difficulty of job
performance, and the question is whether difficulty in performing Gf-type components
(difficulty adapting to new procedures and new technologies, etc.) is observed as an
independent exit reason not reducible to the others. If it is not observed, (a) is rejected. The
reality of the Gf/Gc divergence (Section 2) shows the structural possibility of (a) arising, but
the existence of the divergence is not evidence that the divergence causes exit, and (a) is an
untested empirical claim.
Figure 3 The cognitive bottleneck (Definition 3), its removal by AI, and the formation mechanism of
experiential audit capacity (Definition 4) (Proposition 4). Note: schematic. On the left, difficulty in
performing the Gf-type components of a job bundle forces exit from the whole bundle, and the Gc-type
abilities held never reach the work (Definition 3). On the right, AI substitutes for or complements the Gftype
components, the Gc-type abilities reach the work, and experiential audit capacity (Definition 4) is
connected to audit-type roles. The formation mechanism through the interaction of long-term domain
experience with Gc-type abilities and metacognition is the claim of Proposition 4. Together with the
removal effect (Proposition 2(b)) and the verification power of experiential audit capacity (Proposition
3), the formation mechanism is an untested theoretical claim and an object of testing by Hypothesis H3
and related studies.
Claim (b) (that complementation by AI removes the bottleneck) is the transition depicted in
Figure 3, and at present it lacks direct experimental or quasi-experimental evidence.
Research of the form the refutation condition for (b) demands — providing Gfcomplementing
tools to older workers and measuring changes in job performance and
continuation — is not found within the scope of this paper's search. On the contrary, as seen
in Section 4.5, current data show a headwind — AI usage rates among older groups are
roughly half those of younger groups — which this paper reinterpreted as a consequence of
job design and the allocation of training; but the validity of that reinterpretation is itself part
of the test of (b). Furthermore, removal of the bottleneck is a necessary condition for reach,
(a) Status quo: the Gf-type component gates exit from the whole job
Job (task bundle)
Gf-type component Gc-type component
Impairment → job exit
Held Gc never reaches
the work (lies idle)
(b) Ageless Management: AI complements the Gf-type component
Gf-type component Gc-type component
AI (complement)Deployed as experiential audit capacity
Formation of experiential audit capacity (Def. 4) — Prop. 4: a function of accumulated experience, not age (untested claim)
Long-term domain experience× Gc-type capability (contextual knowledge, interpersonal × Metjaucdoggmneitniot)n
Detection of contextual errors and practical/ethical risks in AI output (tested via Props. 3–5 and H3)
Working Paper | Ageless Management in the AI Era 62
not a sufficient one. Whether the Gc-type abilities that reach the work produce value depends
on Propositions 3 and 4, and the institutional pathways of participation depend on Section 7.
In addition, the "Gc-type abilities a person holds" in (b) requires a temporal qualification.
Following the logic of Section 4.3, unexercised abilities depreciate and knowledge of the
domain context becomes obsolete. The reservoir image — that unutilized Gc-type abilities are
preserved intact in those long past exit — is not permitted by this paper's depreciation logic.
The removal effect of (b), and the resource-conversion claims that presuppose it (Propositions
8 and 11), are therefore conditional on the time elapsed since exit. The eligibility criterion of
Hypothesis H3 (no more than five years away from practice) is the experimental-design
reflection of this condition, and it means at the same time that extrapolation to super-seniors
long after exit requires a separate estimation of a depreciation function with years since
leaving work as a continuous variable (Section 9).
The two boundary conditions the proposition text attaches to (b) are each demanded by the
logic of other parts of this paper. The first boundary condition (maintenance of unassisted
cognitive exercise and the periodic insertion of unmediated audit) is a requirement of
consistency with Section 4.3. The evidence of Section 4.3 showed that constant dependence on
AI can erode performance under unassisted conditions. If, while accepting this skilldepreciation
logic, (b) recommended constant AI complementation unconditionally, this
paper would fall into self-contradiction — dependence on the complementation would, over
the medium to long term, undermine the very foundations of the Gc-type abilities and
metacognition that the complementation is supposed to make reachable. The boundary
condition converts this contradiction into a design requirement: the removal effect of (b)
persists only when unassisted cognitive exercise and unmediated audit of raw output are
periodically inserted into the design of the role, cutting off the depreciation pathway. This
requirement is given operational form in the Brain Safety design of Section 8 (the dual-track
auditing of Section 8.1). The second boundary condition (the lower bound of cognitive
screening) is a requirement to distinguish the object of complementation from its foundation.
What AI complements is the Gf-type components of the task bundle, not the attentional
resources that the exercise of Gc-type abilities itself demands. As confirmed in Section 2.4, the
decline of Gf can accelerate nonlinearly in old age, and in a state where standard cognitive
screening falls below the threshold of mild cognitive impairment and attentional resources
themselves are exhausted, the foundation for the Gc-type exercise to be complemented is lost.
Making this neurological lower bound explicit does not justify the exclusion of any particular
group; it is the theory's honest demarcation, acknowledging that the complementation model
has a physical limit — as the end of the proposition text states, the lower bound is a boundary
on allocation to experiential-audit roles and does not constrain participation in the ecosystem
in general. The implications and misuse risks of this boundary are treated self-critically in
Section 10.4.
Working Paper | Ageless Management in the AI Era 63
Proposition 3 (Non-compressibility of experience)
Contemporaneous assistance by AI compresses differences in production (deliverable
quality, working speed, procedural knowledge) but does not compress differences in
verification (differences in the experiential audit capacity of Definition 4).
Refutation condition: Experimental evidence that, under AI-assisted conditions, the difference
in contextual-error detection rates between holders of long-term domain experience and nonholders
disappears or reverses. Or longitudinal evidence showing that the acquisition of
verification capacity accelerates in AI-use environments to the point of substantially
substituting for the effect of accumulated experience.
The boundary of Proposition 3 is placed on the axis "differences in production vs. differences
in verification." The choice of this axis rests on scrutiny of what the compression evidence
actually compressed. The substance of the compression shown by Brynjolfsson, Li &
Raymond (2025) was the transfer of tacit behavioral patterns extracted from top performers'
interactions — clarifying questions, listening, adjustment of tone — which, in the vocabulary
of Definition 2, includes interpersonal judgment, that is, Gc-type components. The withinfrontier
compression of Dell'Acqua et al. (2023) is likewise a context-dependent output: the
quality of consulting deliverables. In other words, what AI compresses is not limited to Gftype
components; insofar as they are used in production, differences in Gc-type components
are compressed as well. The boundary of compression versus non-compression must
therefore be drawn not between kinds of ability — Gf versus Gc — but between functions:
production (making the deliverable) and verification (detecting its errors). That is why
Proposition 3 is formulated on this axis rather than as "surface versus context," and the
evidentiary status is asymmetric, as follows. The first half — compression of differences in
production — is supported by multiple independent empirical studies in Section 4.1 (quasiexperiments,
preregistered experiments, and field experiments) and is the most heavily
evidenced part of this paper's propositions. The second half — differences in verification are
not compressed — is a theoretical claim of this paper that at present has no direct evidence of
the same standard. The indirect support is limited to the following: the compression studies
all measured the production of deliverables within AI's competence range and did not
measure verification capacity; and the deterioration of AI users outside the frontier
(Dell'Acqua et al. 2023) suggests the persistence of a human-side function of discerning the
boundary of that competence range.
The qualifier "contemporaneous assistance" in the proposition text, and the second
sentence of the refutation condition (longitudinal evidence), are also demanded by the
implications of the compression evidence. Brynjolfsson's compression is a shortening of the
experience curve — a compression of acquisition time — and the possibility that the same
mechanism operates on the acquisition of verification capacity — a pathway in which AI
turns failure cases and regulatory context into teaching material and accelerates the
Working Paper | Ageless Management in the AI Era 64
formation of junior auditors' capacity — cannot be excluded a priori. If this longitudinal
compression pathway is real, then even if experience differences persist in the
contemporaneous 2×2 experiment (H3), the effect of accumulated experience would be
substituted over time, and Proposition 3 would effectively fail. By writing this pathway —
which a static experiment cannot observe in principle — into the refutation condition, and
placing the corresponding longitudinal measurement in the implementation roadmap of
Section 9, the refutation condition was made to cover the full range of the proposition's claim.
Furthermore, the existence of reports unfavorable to the second half must be mentioned. A
randomized-experiment report said to show that even experts miss AI errors has appeared in
PNAS Nexus (Section 4.4; as its content and figures are unverified, mention is limited to its
existence), and whether expertise improves verification performance is itself still an open
question. Precisely for this reason, the second half of Proposition 3 is tested directly by the
2×2 factorial design of Hypothesis H3 (experience level × presence of AI assistance). H3's
prediction is that in the detection of surface-level errors the experience difference shrinks
with AI assistance (consistent with the first half), while in the detection of contextual and
practical risks and in the quality of proposed corrections, the main effect of experience
persists and does not shrink under AI assistance. If this prediction fails — if, as the refutation
condition states, the experience difference in contextual-error detection rates disappears or
reverses — Proposition 3 is rejected and this paper's theory loses its core (Section 5.5).
Working Paper | Ageless Management in the AI Era 65
Proposition 4 (Formation of experiential audit capacity)
Experiential audit capacity is formed not by chronological age but by the interaction
between domain-specific knowledge and operational schemata in a particular field (the
product of long-term domain experience) and generalized crystallized intelligence
(vocabulary, reading comprehension, general knowledge) together with metacognition.
The two are conceptually distinct — generalized Gc transfers across domains, whereas
domain knowledge is bound to its domain. This formation has three boundary
conditions. (i) When AI output conforms to the auditor's own past successes and
industry received wisdom, the synergy of confirmation bias and automation bias means
that experience can instead impede detection. (ii) In domains where the speed of
technological change is high and the half-life of domain knowledge is short, the audit
effectiveness of accumulated experience declines and can turn negative. (iii) When the
fluency and stylistic polish of AI output are high, the effect of processing fluency raises
the threshold of the auditor's cognitive sense of incongruity, and the detection
performance of experiential audit capacity can decline — the very trigger of detection,
the "sense of incongruity" itself, becomes less likely to fire in the face of fluent output.
Refutation condition: Systematic evidence that chronological age independently predicts
detection capacity even after controlling for years of domain experience, or that years of
experience lose predictive power after such control. For the boundary conditions: evidence that
experienced auditors' detection rates do not fall below non-experienced auditors' even on
confirmation-conforming errors empties (i); evidence that the audit effectiveness of experience
is preserved even in high-change-speed domains empties (ii); and evidence that detection
performance does not decline when output fluency is manipulated empties (iii) (each is a
refutation of a boundary and works in the direction of strengthening the main proposition).
Proposition 4 puts the first theoretical operation (the reattribution from age to experience)
into refutable form. Note that the refutation condition is bidirectional. If chronological age
has independent predictive power even after controlling for years of experience, the
reattribution is wrong and there is something in age itself — this paper's framework loses its
entitlement to call itself "ageless." Conversely, if years of experience lose predictive power
after control, the paper's reattribution of detection performance (Definition 4) to experience
loses support, and the proposition collapses in a different way. Because Definition 4 is
purified into an operational definition (Section 5.2), either result rejects Proposition 4 directly,
with no retreat into the definition. Proposition 4 survives only when a specific variance
decomposition is observed: experience predicts, and age does not.
The proposition text explicitly distinguishes generalized crystallized intelligence from
domain-specific knowledge in order to make the interaction claim withstand the criticism of
multicollinearity. Holders of long-term domain experience tend to have high levels of
generalized Gc as well, and if the two were measured as a single "experience-and-knowledge"
construct, the interaction term would become inseparable from the main effects, and
Working Paper | Ageless Management in the AI Era 66
Proposition 4 would degenerate into an untestable paraphrase. The substance of the
distinction lies in transferability. Generalized Gc (vocabulary, reading comprehension,
general knowledge) transfers across domains, whereas domain knowledge and operational
schemata are bound to their domain. From this distinction a testable divergence of
predictions is obtained. A non-experienced person high only in generalized Gc (a welleducated
outsider) should be inferior to a long-term experienced person in detecting
contextual risks in the domain, and a person with domain knowledge but low generalized Gc
and metacognition should likewise be inferior in detection — because Proposition 4's claim is
an interaction, not an addition. Corresponding to this prediction, Hypothesis H3 measures
generalized Gc and domain knowledge with separate standardized indicators and explicitly
estimates the interaction term (Appendix A).
The theoretical grounds of the boundary conditions are as follows. Boundary condition (i) is
the proposition-level reflection of the synergy of confirmation bias and automation bias
organized in Sections 4.2 and 4.4. Experience supplies the prior distribution in verification.
When AI output conforms to the auditor's own past successes and industry received wisdom,
the experience-derived prior works in the direction of endorsing the output's validity and
aligns with overtrust in AI output (automation bias) — at that point experience turns from a
resource for detection into a cause of missed detection. Boundary condition (ii) is a
consequence of the half-life of knowledge. The audit effectiveness of domain knowledge holds
only insofar as the correspondence between accumulated schemata and current practice,
technology, and regulation is preserved; in domains of rapid technological change, the
schemata systematically mispredict the current risk structure, so the audit effectiveness of
experience declines and can turn negative. Boundary condition (iii) is the reflection, in the
audit context, of the cognitive-psychology findings on processing fluency. When a stimulus is
subjectively easy to process, people judge its content to be more truthful — it has been shown
experimentally that manipulating perceptual fluency alone raises judgments of a sentence's
truth (Reber & Schwarz 1999), and the finding that fluency broadly elevates judgments of
truth, liking, and confidence has been organized in a systematic review (Alter &
Oppenheimer 2009). The output of generative AI sits precisely at this high-fluency pole —
grammatically well-formed, stylistically smooth, low in processing resistance. The trigger of
detection for experiential audit capacity is the cognitive sense of incongruity generated by a
mismatch between accumulated schemata and the output, but high-fluency output raises the
very threshold of this sense of incongruity, so the detection mechanism can be bypassed
without any deficit in the auditor's ability. That is, (iii) is a pathway in which the surface
properties of the object of verification operate, independently of (i), which originates in the
auditor's internal prior distribution. Note that the structure of the refutation condition is
asymmetric. Refutation of the main body (independent predictive power of age, or loss of
predictive power of experience) rejects the proposition, whereas refutation of the boundaries
— experienced auditors not inferior even on confirmation-conforming errors, effectiveness
preserved even in high-change-speed domains, no decline in detection when fluency is
Working Paper | Ageless Management in the AI Era 67
manipulated — empties the boundary conditions and leaves the main proposition standing in
a stronger form. Hypothesis H3 constructs the embedded errors in both confirmationconforming
and deviating types and controls the fluency of the task documents as a factor or
covariate, precisely in order to test the main body and boundary conditions (i) and (iii)
simultaneously in a single experiment (Appendix A).
The evidentiary status is untested. As confirmed in Section 4.4, no study operationalizing
age or tenure and measuring performance in AI oversight and verification was found within
the scope of this paper's search, and this absence is the basis of the paper's research-gap
claim. Because Hypothesis H3 manipulates experience level (long-term domain experience
holders versus juniors) as a factor, the complete test of Proposition 4 — a variance
decomposition varying experience and age independently — requires an extension that
includes older non-experienced participants (career changers, returnees) and young longterm-
experienced participants in the sample, and this is reflected in the participant
requirements of the experimental protocol in Appendix A. Note that Proposition 4 is a claim
about the formation factors of experiential audit capacity, not a claim of super-senior
superiority. There may be a distributional tendency for holders of long-term domain
experience to be relatively numerous among super-seniors, but that is an implication
mediated by a bridging assumption, not the content of the proposition (Section 4.4).
Proposition 5 (Generational decorrelation)
Under the following three conditions, a generationally heterogeneous supervisor group
has less overlap in what it misses in AI output than a same-generation supervisor group,
a wider collective detection set, and detections that reach decision-making. (i) Channel
condition: auditors access the raw output under verification in an unmediated way. This
need not cover the full volume; raw audit of a randomly sampled portion suffices (dualtrack
auditing — Section 8.1). (ii) Organizational condition: power gradients are
flattened so that decorrelated observations reach decision-making without suppression
or anticipatory deference (corresponding to the failure mode of "silence" that Belonging
in BCM 2026e guards against). (iii) Model condition: the AI output under audit is not
monopolized by a single foundation model, and cross-verification by models of different
architectures and developers is used in parallel — human-side generational
heterogeneity cannot override a situation in which the systematic blind spots of a single
model dominate all output. Absent these, the decorrelation supplied by generational
heterogeneity can be lost at the channel, organizational, or model level.
Refutation condition: Empirical evidence that, under a design satisfying the three conditions,
the cross-generational correlation of detection errors is equal to or higher than the withingeneration
correlation, or that the adoption rate of decorrelated observations into decisionmaking
does not differ significantly from the homogeneous-group case.
Working Paper | Ageless Management in the AI Era 68
The derivation of Proposition 5 is most accurately presented as a connection to WP8
(Kadowaki 2026h). WP8, facing squarely the systematic failures of human oversight —
automation bias, alarm fatigue, and the average inferiority of human-AI combinations
(Vaccaro, Almaatouq & Malone 2024) — formalized the value of an oversight channel as
independence × detection probability and generalized the substance of independence to
"error decorrelation." If multiple oversight channels share the same blind spots, adding
channels does not widen the detection set. And WP8 organized, from case analyses, the
observation that homogeneous supervisor groups sharing the same contemporaneous
training and information environment as the AI are precisely the ones prone to falling into
this state of shared blind spots. Proposition 5 is a supply-side response to this decorrelation
condition. If the correlation of error distributions is low among individuals who have
experienced different historical environments, technological generations, and failure cases
(Definition 5), then a generationally heterogeneous supervisor group should be able to supply
decorrelation organizationally.
The three conditions in the proposition text serve to build into the proposition itself the
three levels of pathways by which the supply of decorrelation can be lost. First, the reason for
(i), the channel condition. What generational heterogeneity supplies is the decorrelation of
error distributions inside individual auditors. But if the route by which auditors reach the
object of verification — the information channel — is shared, individual-level decorrelation is
canceled at the channel level. When all auditors view the output through a summary
produced by the same AI, the context dropped by that summary is equally unseen by all
auditors. Pre-screening and prioritization based on the AI's displayed confidence pushes the
places where the AI errs confidently — the typical hallucination — equally outside all
auditors' attention. That is, the heterogeneous blind-spot structures that different historical
environments have given individuals are homogenized at the entrance of the audit process
by passage through a single channel, and WP8's independence condition — oversight value =
independence × detection probability — is broken at the level of the channel. On the other
hand, unmediated audit of the full volume exceeds WP8's solvency condition — the upper
bound on cognitive resources that can be devoted to oversight — and collides head-on with
Brain Safety's demand for load reduction. That the proposition text specifies "this need not
cover the full volume; raw audit of a randomly sampled portion suffices" is the resolution of
this collision, and it is designed in Section 8.1 as dual-track auditing, combining raw audit —
original-output, unsummarized, unscreened — of a randomly sampled portion with AIsummary-
assisted screening of the remainder. Random sampling is essential — partial audit
filtered by AI reintroduces channel sharing through the filtering itself. The measurement of
detection-overlap rates with the audit channel (AI-mediated versus unmediated) as a factor is
detailed in Section 9.3.
(ii) The organizational condition is demanded by the distinction between statistical
decorrelation and organizational adoption. Even if generational heterogeneity actually
lowers the correlation of error distributions — statistical decorrelation obtains — oversight
Working Paper | Ageless Management in the AI Era 69
value is zero unless the detections reach decision-making. What blocks their reach is the
organization's power gradient. Youth, non-employment participants, and super-seniors alike
tend to sit downstream of the organization's power gradient, and observations that conflict
with the judgment of the majority or of superiors — decorrelated observations are precisely
such observations — can vanish short of decision-making, through suppression (explicit
dismissal) or anticipatory deference (voluntary silence). This is the reappearance, in the audit
context, of the failure mode of "silence" that BCM (Kadowaki 2026e) formalized as the
Belonging of the 3Bs. That the refutation condition of Proposition 5 lists, as an independent
refutation quantity, not only the correlation of detection errors but "the adoption rate of
decorrelated observations into decision-making" follows from this distinction — if statistical
decorrelation obtains but the adoption rate does not differ from the homogeneous group, the
oversight value claimed by Proposition 5 has not been realized.
(iii) The model condition is demanded by the limits of the level at which human-side
decorrelation can operate. What generational heterogeneity manipulates is the error
distribution on the human side. But when the AI output under audit all derives from a single
foundation model, the systematic blind spots of that model — errors rooted in the training
distribution and architecture, appearing in correlated fashion across all output — dominate
every audit task as a common factor on the output side. However heterogeneous the blindspot
structures the human auditors bring, if the object of verification is monopolized by a
single error-generating source, the model-level correlation cannot be overridden by them —
an error for which no one is given any detection cue is not detected even if the generation is
changed. Accordingly, the parallel use of cross-verification by models of different
architectures and developers is a precondition for human-side decorrelation to function. As
stated in the commentary on Proposition 1, this AI-side diversity is at the same time an
operation that narrows, from the machine side, the residual niche of human verification
value — the two stand in a relation of both substitution and precondition, and where the
equilibrium lies is an empirical question. Yet wherever this equilibrium eventually settles, the
structure itself — that the detection of systematic blind spots requires decorrelated verifiers
— does not disappear with capability gains, and Proposition 5 persists as a framework for
allocating the sources of that decorrelation — humans with heterogeneous experience and
models of different lineages (Section 10.5).
This extension to the generational axis, however, is not WP8's own claim but a new claim
made by this paper, and it is untested. To blur this point would be nothing other than the
operation of making untested claims look established through a chain of self-citations within
the series, and this paper explicitly forbids it (Section 10). What can be inherited from WP8 is
the framework — that decorrelation is part of the necessary conditions of oversight value —
and no more; whether generational heterogeneity actually lowers error correlation is an
empirical question to be tested by the quantity Definition 5 specifies in measurable form (the
comparison of cross-generational and within-generational correlations of detection errors).
As related indirect evidence, a mock-jury experiment manipulating racial diversity reported
Working Paper | Ageless Management in the AI Era 70
that diverse groups exchanged a wider range of information and that majority members' own
factual errors decreased (Sommers 2006), but this is laboratory evidence on racial diversity,
and replication with age and generation has not been confirmed. No peer-reviewed empirical
study directly showing that age diversity improves error detection or red-team performance
was found within the scope of this paper's survey.
Moreover, WP8's formalization teaches that even if Proposition 5 were true, it alone would
not guarantee oversight value. Independence is a necessary condition, not a sufficient one;
value arises only when detection probability is multiplied in. That is, Proposition 5 (the
supply of decorrelation) and Propositions 3 and 4 (the reality and attribution of detection
capacity) stand in a multiplicative relation, and only when both hold is the design of a
multigenerational supervisor group (Section 6.3) justified. This is why the detection-error
overlap analysis is built into the test design of Hypothesis H3, and why extension to oversight
experiments manipulating generational composition is needed in future research.
Proposition 6 (AI-mediated diversity effect)
Given the known conditions of task complexity and an inclusive climate, AI-mediated
complementarity (Definition 6) is an additional moderator of the relation between age
and experience diversity and organizational outcomes, and in its presence the relation
moves in a more positive direction. Outcomes here include not only the quality of ideas
but the avoidance of excess risk and the reduction of rework, and are evaluated as net
benefit after deducting the cost of decision-making time required for verification.
Refutation condition: Experimental or quasi-experimental evidence that, in comparisons
manipulating the presence of AI mediation while controlling task and climate conditions, no
difference arises in the relation between diversity and outcomes. Furthermore, if an observed
interaction derives solely from the deterioration of homogeneous teams through AI overtrust
and is not accompanied by an improvement in the absolute outcomes of mixed teams, this
proposition is not regarded as supported (decomposition of simple main effects is required —
H2).
The starting point of Proposition 6 is the frontal acceptance of the fact that the empirical
baseline for age diversity is severe. The average relation between age diversity and team
outcomes shown by meta-analyses lies near zero, from r=−.06 (Joshi & Roh 2009) to r=.014
(Wallrich et al. 2024), and Schneid et al. (2016) conclude no significant relation (the sole
exception being turnover). The unconditional claim that "multigenerational means more
value" cannot be supported empirically. At the same time, it has also been consistently
observed that the effect is condition-dependent. Age diversity has positive effects on
productivity only in firms engaged in creative tasks (Backes-Gellner & Veen 2013); the relation
between diversity and outcomes becomes more positive for tasks of high complexity that
depend on creative divergence (Wallrich et al. 2024); and a climate of low age discrimination,
positive valuation of diversity, and supportive leadership are conditions of success (Wegge et
Working Paper | Ageless Management in the AI Era 71
al. 2012). At the macro level, too, the relation between age diversity and productivity is humpshaped,
suggesting the existence of an optimum (Zélity 2023).
Proposition 6 is a theoretical claim that places an additional moderator — AI-mediated
complementarity — on top of this evidence structure in which "conditions decide everything,"
and its formulation takes the form of comparative statics. That is, Proposition 6 does not
make the absolute-level claim that the effect of diversity "becomes positive" under AI
mediation. What it claims is only a change of relation: in a comparison holding fixed the
known necessary conditions of task complexity and an inclusive climate, the presence of AImediated
complementarity moves the relation between diversity and outcomes in a more
positive direction. In this formulation, the proposition contradicts neither the empirical work
that detected positive effects in particular subpopulations with pre-AI data (Backes-Gellner &
Veen 2013) nor the known set of moderators (Wegge et al. 2012 and others). The theoretical
mechanism is as follows. The main part of the cost of diversity arises from communication
costs among heterogeneous members and from each member's need to fill their own weak
components on their own. AI-mediated complementarity (Definition 6) lowers both —
complementation of Gf-type components loosens the interdependence of weaknesses and
lowers the cost of connecting heterogeneous cognitive assets — so the region where the
benefits of diversity (complementarity of perspectives and blind spots) exceed its costs should
expand. However, the existing moderator studies (task complexity, climate) did not measure
AI mediation. Proposition 6 has a form that can explain the mixture of existing evidence after
the fact, but post hoc explanation is not verification. Until a comparison manipulating the
presence of AI mediation while controlling task and climate conditions is carried out — the
mixed-team experiment of Hypothesis H2 is the first step — Proposition 6 remains an
untested theoretical proposition.
The proposition text defines outcomes as net benefit in order to put the benefits and costs
of decorrelated auditing on the same ledger. The verification supplied by multigenerational
composition is not free — processing and coordinating heterogeneous observations delays
decision-making, and verification labor consumes cognitive resources (a consequence of
WP8's solvency condition). If outcomes were measured by idea quality alone, these costs
would go unrecorded while only benefits were observed, and Proposition 6 would be biased
toward irrefutability. Conversely, the principal benefits of auditing — the avoidance of excess
risk and the reduction of rework — are hard to see in immediate evaluation of outputs, and
looking only at costs would make diversity appear permanently inferior. The proposition text
therefore defines outcomes as net benefit, including the avoidance of excess risk and the
reduction of rework and deducting the cost of decision-making time required for verification,
and correspondingly Hypothesis H2 requires that time-to-decision and verification effort be
recorded as cost variables (Section 9). This net-benefit framework connects to the transactioncost
discussion of Section 6.5. In the analysis of H2, moreover, detecting an interaction is not
enough. Since human-AI combinations can on average fall below the best single agent
(Vaccaro, Almaatouq & Malone 2024), an interaction can also arise because the AI-mediated
Working Paper | Ageless Management in the AI Era 72
condition lowers the performance of homogeneous teams. To distinguish this from support
for Proposition 6, a decomposition of simple main effects is needed — the direction of the
diversity effect in each condition cell and the identification of the source of the interaction's
sign — and this requirement is reflected in the analysis plan for H2 in Section 9. The second
sentence of the revised refutation condition writes this requirement — excluding from
support for the proposition a spurious interaction deriving solely from the deterioration of
homogeneous teams through AI overtrust — into the proposition itself.
Proposition 7 (Cognitive engagement pathway)
If a brain-health effect of work exists, it is mediated not by the length of working hours
but by the maintenance of cognitive engagement through occupation in Gc-type roles.
Refutation condition: Evidence that, in mediation analysis, working hours themselves still
predict the maintenance of cognitive function after controlling for cognitive engagement, or
that the mediation pathway is rejected.
Note that Proposition 7 is a conditional. This paper does not presuppose that "work protects
brain health." As confirmed in Section 3, the empirical literature on retirement and cognitive
function is deadlocked between IV estimates showing negative effects (Rohwedder & Willis
2010) and estimates of comparable standing showing no effect or improvement, and the
verdict of systematic reviews is likewise mixed (Meng et al. 2017). Proposition 7 is an
explanatory hypothesis for this deadlock. If the category "work" mixes cognitively rich roles
with depleting ones, it is no surprise that the sign of the average effect fails to settle, and the
variable carrying the effect should be not the presence or duration of work but active
occupation in Gc-type components — evaluation, contextual interpretation, interpersonal
judgment — that is, cognitive engagement.
The evidentiary status is suggestive. That the variables separating the sign of the effect are
occupation type, the voluntariness of retirement, and the cognitive content of the work; that
accelerated decline on the Gc side was observed only for retirement from jobs of high
interpersonal complexity (Meng et al. 2017); and that RCTs of a non-employment program
with designed role quality showed effects on cognitive function and brain structure in limited
populations (Experience Corps: Fried et al. 2004; Carlson et al. 2008) are all consistent with
the mediation structure. But consistency is not verification. No study has identified the
mediation pathway itself, and Proposition 7 is the direct test target of Hypothesis H1 — an
identification strategy using exogenous variation together with mediation analysis of
cognitive engagement. If Proposition 7 is rejected, the brain-health benefit claim (part of
Proposition 8) loses its basis, but the economic benefit claims (Propositions 1–5) stand
independently (Section 5.5).
Working Paper | Ageless Management in the AI Era 73
Proposition 8 (Bidirectional capital formation)
A multigenerational ecosystem that satisfies ex-ante verifiable conditions — the
connection modes of Definition 7 and the satisfaction of Definition 8, plus the existence
of a developmental pathway in which youth progressively experience Gc-type roles in
small, low-risk projects with full delegation and acceptance of responsibility for failure
(micro-ownership) (cognitive apprenticeship — simulated audit without responsibility
does not qualify), operated as a checklist that a third party can adjudicate prior to the
observation of outcomes — increases brain capital = K (stock) × u (utilization rate)
bidirectionally. It appears as restrained depreciation (maintenance) of K among superseniors,
early formation of K among youth and marginalized groups, and a rise in u for
the organization.
Refutation condition: Evidence that, under a design adjudicated as satisfying the conditions
prior to the observation of outcomes, the brain-capital indicators of any participating stratum
systematically deteriorate ex post. Retracting the adjudication of condition satisfaction
retroactively from ex-post outcomes is not admitted as a defense of this proposition.
Proposition 8 extends the brain capital = K × u framework inherited from WP5's Brain Capital
Management (Kadowaki 2026e) from the single organization to the multigenerational
ecosystem. "Bidirectional" means the claim that the benefit appears not as a grant to a
particular stratum but as capital formation across all participating strata. Among superseniors,
the continued exercise of Gc-type roles restrains the depreciation of K (the pathway
of Proposition 7, and the reverse side of Section 4.3 — abilities not exercised depreciate).
Among youth and marginalized groups, connection with holders of experience promotes the
early formation of K. For the organization, the connection of previously idle Gc-type assets
raises u. Through this three-part structure, Proposition 8 stands not as an argument for
supporting seniors nor for developing the young, but as a theory of capital formation.
The dimensionality and structure of K should be made explicit here. K is not a single
number. This paper conceptualizes K as a construct measured in three dimensions: (1) clinical
cognitive function scores (standardized tests such as MoCA), (2) structural and functional
brain indicators, and (3) standardized indicators of domain knowledge. This inherits the
measurement framework of BCM (Kadowaki 2026e): brain capital is not a metaphor but a
construct whose measurement procedures can be specified dimension by dimension. The
integration of the three dimensions, however, is not an additive composite score. The
structure of K is a hierarchical function in which the foundational cognitive function Kbase
measured by (1) and (2) exceeding a threshold θ serves as a gate (precondition), on top of
which the domain knowledge and operational schemata Kdomain of (3) enter multiplicatively.
K = 1[Kbase ≥ θ] × f(Kbase, Kdomain)
Working Paper | Ageless Management in the AI Era 74
The reason for adopting the hierarchical structure lies in the absurdity of an additive sum. If
the dimensions were composited additively, the substitution of offsetting a decline in clinical
cognitive function scores with abundant domain knowledge to hold K constant would be
formally permitted. But in a state where the foundation of attentional resources is lost, no
amount of domain knowledge reaches audit exercise — the audit value of domain knowledge
manifests only on top of foundational cognitive function and is not a substitute for it. The gate
term 1[Kbase ≥ θ] expresses this non-substitutability at the level of functional form. And this
threshold θ is the same gate as the cognitive-screening lower bound specified by the second
boundary condition of Proposition 2(b) — that standard cognitive screening not fall below the
threshold of mild cognitive impairment. That is, the lower bound Proposition 2 placed as a
boundary on allocation to experiential-audit roles and the gate placed by the measurement
structure of K in Proposition 8 are two manifestations of the same theoretical fact — the
distinction between the object of complementation and its foundation — and this coincidence
secures the internal consistency of the theory. This structure refines the content of
Proposition 8's claim. The restrained depreciation of K among super-seniors should be
observed mainly as a restrained rate of decline of Kbase, and the early formation of K among
youth and marginalized groups mainly as the accumulation of Kdomain and the
developmental formation of Kbase; the "systematic deterioration of brain-capital indicators"
in the refutation condition is adjudicated by dimension-specific measurement in line with
this hierarchical structure. The details of measurement are placed in the organization-level
indicators of Section 9.3.
The proposition text adds cognitive apprenticeship to the conditions as a response to a
diachronic risk. If the AI-mediated division of labor were optimized only statically, it could
converge on a configuration that fixedly assigns Gf-type prototyping to youth and Gc-type
auditing to holders of experience. But this division of labor deprives youth of the opportunity
to accumulate experience — the pathway of Gc formation that consists of judging, failing, and
bearing the consequences. If Proposition 4 is correct, experiential audit capacity is the
product of long-term domain experience, so a division of labor lacking a youth Gc-formation
pathway destroys the very future supply of experiential audit capacity — a diachronic tradeoff
between present efficiency and the future supply of experiential audit capacity. Cognitive
apprenticeship — a developmental pathway in which youth progressively experience Gc-type
roles in small, low-risk projects (Section 6.2) — is the condition that resolves this trade-off at
the level of an ex-ante verifiable checklist. That the proposition text explicitly requires microownership
of this pathway — full delegation and acceptance of responsibility for failure —
and excludes simulated audit without responsibility, has a reason grounded in the formation
mechanism of metacognition. The core of the experiential audit capacity claimed by
Proposition 4 is the metacognition of knowing where one's own judgment can err, and it is
calibrated only under the condition that one bears the consequences of one's judgments
oneself. In simulated audit whose consequences are attributed to others, the cost of error
does not return to the auditor, so the mapping between judgment and consequence — the
Working Paper | Ageless Management in the AI Era 75
feedback loop necessary for calibrating metacognition — does not close. However much
simulated audit without responsibility is repeated, what is formed is knowledge of the formal
procedures of auditing, not genuine metacognition. The requirement of micro-ownership is
therefore not an additional desideratum of the developmental pathway but a constitutive
condition for cognitive apprenticeship to function as a Gc-formation pathway. An ecosystem
lacking it degenerates into a one-way design that maintains the K of super-seniors while
sacrificing the future K formation of youth. That the refutation condition rejects the whole
proposition upon the deterioration of "any participating stratum" is meant to refuse to
condone this one-way degeneration under the name of bidirectionality.
The evidentiary status differs by stratum. That intergenerational knowledge transfer
carries motivational benefits for both sender and receiver (Burmeister, Wang & Hirschi 2020),
that reverse mentoring can yield skill development on both sides (Kaše, Saksida & Miheliฤ
2019), and the qualitative description that intergenerational learning is bidirectional (Gerpott,
Lehmann-Willenbrock & Voelpel 2017) are peer-reviewed findings consistent with the
structure of bidirectionality. What these measured, however, was motivation, retention, and
skill development, not brain-capital indicators themselves. Verification on brain-capital
indicators is entrusted to Hypotheses H1 and H2 and the organization-level indicators of
Section 9, and in that sense Proposition 8 is an untested proposition with suggestive evidence.
The refutation condition cites the deterioration of "any participating stratum" because
bidirectionality is precisely the content of the proposition. A design that maintains the K of
super-seniors at the cost of youths' learning, or the reverse, rejects Proposition 8 even if other
strata benefit.
The proposition text requires that the conditions be operated as an "ex-ante verifiable
checklist," and the refutation condition explicitly forbids retroactive denial of the conditions,
in order to seal off a refutation-evasion structure peculiar to conditional propositions. Since
Definition 8 is itself the standard protecting brain capital from depreciation, if the
adjudication of condition satisfaction were made dependent on ex-post outcomes, every case
in which brain-capital indicators deteriorated could be reclassified retroactively as "Brain
Safety was not satisfied," and no failure would ever wound the proposition — an
immunization of the no-true-Scotsman type, under which one can keep saying "it was not
true Ageless Management." The only way to sever this structure is to make the adjudication of
condition satisfaction independent of outcomes. That is, the connection modes of Definition 7
and each requirement of Definition 8 (Table 9 in Section 8 gives their skeleton) are operated
in checklist form such that a third party can adjudicate satisfaction before seeing results, and
the adjudication procedure is written into the proposition itself: if deterioration is observed
ex post under a design adjudicated ex ante as "conditions satisfied," the proposition is
rejected. The second sentence of the refutation condition is the consequence of this
procedure, making explicit that retroactive reclassification from ex-post outcomes is not
admitted as a defense of the proposition.
Working Paper | Ageless Management in the AI Era 76
Proposition 9 (The protection vacuum of non-employment forms)
Current labor and social-security institutions are designed with the employment
relationship as the principal unit of protection and restraint, and the non-employment
participation on which Ageless Management depends (outsourced engagement, advisory
roles, PBL, NPO partnerships) falls into a vacuum of institutional protection. Ageless
Management that leaves this vacuum unaddressed can degenerate into exploitation.
Refutation condition: That in jurisdictions granting non-employment participants protections
equivalent to employment (accident compensation, adequacy of remuneration, correction of
bargaining power), cases of exploitation and unpaid-labor conversion attributable to the
protection vacuum are not systematically observed — in which case this proposition becomes
empty in that jurisdiction.
Proposition 9 is the proposition by which the theory internalizes the price of Definition 7's
lifting of the employment boundary. Protective devices such as working-hour regulation,
minimum wages, accident compensation, and dismissal regulation are designed with the
employment relationship as their unit, and outsourced engagement, advisory roles,
educational partnerships, and NPO partnerships lie outside many of them. Under this
structure, there is an ever-present danger that, in the name of Ageless Management, levels of
load and unpaid work that would be illegal in employment are legitimated as "flexible
participation." The evidentiary status of Proposition 9 is at the level of institutional analysis
rather than experimental verification, and the actual state of each jurisdiction's institutions
— Japan's freelance-protection legislation, labor protections for youth, and the institutional
environment of older-age employment — is described in Section 7 on the basis of verified
primary legal sources.
The refutation condition of Proposition 9 occupies a singular position within this theory,
because the state in which it is satisfied — the existence of jurisdictions granting nonemployment
participants protections equivalent to employment — is a desirable state for this
paper. Proposition 9 is not a claim of permanent truth but a pointer to a structural feature of
current institutions, and it is a proposition that positively hopes to be rendered "empty" by
institutional reform. In this respect Proposition 9 functions as the normative node that calls
for the institutional comparison of Section 7 and the Brain Safety design of Section 8. What
the pursuit of Proposition 8 in disregard of Proposition 9 — capital formation without
protection — degenerates into is stated explicitly at the end of the proposition.
Working Paper | Ageless Management in the AI Era 77
Proposition 10 (Transferability of the design principles of protection)
Labor protection for youth and the protection of super-seniors from exploitation and
cognitive overload are not identical in their grounds — for youth there exist principles
with no counterpart for older strata: consideration for developmental stage and the
priority of education. At the level of design principles, however, the two share a common
structure as responses to asymmetries of bargaining power and to exit costs (priority of
health, ceilings on load, adequacy of compensation), and this common part is
transferable across strata.
Refutation condition: Legal-theoretical or empirical evidence showing, for any of the design
principles held to be common structure, that it fails to function as protection when applied to
one stratum, or that it is structurally incompatible with the protection of the other stratum.
Proposition 10 does not claim identity in the grounds of protection. As the standard legal
doctrine of the ILO conventions (C138/C182) and the minor-protection provisions of the Labor
Standards Act shows (Section 7.3), at the core of youth protection lie principles with no
counterpart in the protection of older strata — consideration for developmental immaturity
and the connection to compulsory education — and the proposition text acknowledges this
asymmetry squarely within the proposition. What is claimed is weaker, but substantive as
design theory: in the residual after removing stratum-specific principles, the protections of
both strata share the character of responses to a common structure — asymmetry of
bargaining power and exit costs — and the design principles derived from it (priority of
health, ceilings on load, adequacy of compensation) are transferable across strata — the
formulation of common structure plus stratum-specific principles. Youth are prone to
accepting unfavorable terms through lack of experience, information, and alternative
opportunities; super-seniors through the exit costs of scarce reemployment opportunities and
role loss. Within the limits of this common structure, Brain Safety (Definition 8) can be
constructed as a hierarchical design that constitutes the common part as a single system of
standards and stacks stratum-specific principles (such as the primacy of education for youth)
on top. This is the theoretical basis of the design argument of Section 8.
The evidentiary status is untested, and the character of the test also differs from other
propositions. The validity of Proposition 10 is adjudicated by legal-theoretical examination
and comparative institutional analysis rather than by experiment. The refutation condition is
placed not on the interpretive question of the sameness or difference of grounds but on an
operable criterion: failure of transfer. If any of the design principles held to be common
structure is shown to fail to function as protection when applied to one stratum, or to be
structurally incompatible with the protection of the other, the transferability claim is rejected
and Brain Safety retreats entirely to stratum-specific design. This paper does not exclude that
possibility. Even if Proposition 10 falls, Definition 8 itself remains maintainable in stratified
form, and there is no propagation to the core of the theory (Section 5.5).
Working Paper | Ageless Management in the AI Era 78
Proposition 11 (Conversion of social problems into resources)
The three social problems of population aging, constrained youth participation, and the
exclusion of marginalized groups are converted into untapped sources of brain capital,
conditional on AI-mediated complementarity (Definition 6) and on Brain Safety
(Definition 8) operated in an ex-ante verifiable form. The conversion is not automatic.
Refutation condition: Systematic evidence that, even under a design adjudicated as satisfying
the conditions prior to the observation of outcomes, the participation of the strata in question
does not contribute to the organization's brain-capital and performance indicators. As with
Proposition 8, retroactive denial of the conditions ex post is not admitted as a defense.
Proposition 11 is the third theoretical operation (the conversion of social problems into
resources) itself, and it is the integrative proposition of the upstream propositions. The
conversion pathway runs through Proposition 2 (reach through removal of the bottleneck),
Propositions 3–5 (the audit value of the abilities that reach), Proposition 6 (diversity effects
under AI mediation), and Proposition 8 (bidirectional capital formation), and is conditioned
by Propositions 9 and 10 (the design of protection). Proposition 11 is therefore less an
independent claim than an integrative consequence that holds only under a conjunction of
conditions — that both AI-mediated complementarity and Brain Safety are satisfied — and it
weakens in tandem if any upstream proposition falls. The final sentence, "The conversion is
not automatic," is at once a summary of this conditionality and an explicit break with the
discourse that unconditionally relabels population aging an "asset" — an optimism that is
merely the flip side of the accommodation-and-compensation paradigm this paper rejected in
Section 1.
The reason the proposition text attaches the qualifier "operated in an ex-ante verifiable
form" to the Brain Safety condition, and the refutation condition forbids retroactive denial of
the conditions, is identical to the sealing of immunization described in the commentary on
Proposition 8. In a conditional proposition, if the adjudication of condition satisfaction
depends on ex-post outcomes, every failure case can be reclassified as "the conditions were
not satisfied," and the proposition becomes effectively irrefutable. As an integrative
proposition, Proposition 11 sits in the position most prone to this structure — every case in
which conversion failed to occur could be blamed on unsatisfied conditions. Proposition 11
therefore adopts the same adjudication procedure as Proposition 8. The satisfaction of
Definitions 6 and 8 is operated as a checklist that a third party can adjudicate prior to the
observation of outcomes (Table 9 in Section 8), and if no contribution is observed under a
design adjudicated ex ante as satisfying the conditions, the proposition is rejected.
Reclassification tracing back from ex-post outcomes — "it was not true condition satisfaction"
— is not admitted as a defense.
The evidentiary status is untested. For each of the three social problems' strata, no study
exists showing that participation under a design satisfying the conditions contributes to the
Working Paper | Ageless Management in the AI Era 79
organization's brain-capital and performance indicators. What this paper can offer is the
evidentiary status of each proposition composing the pathway (Table 7) and an order of
verification: after testing the pathway through H1–H3, integrative evaluation through the
implementation roadmap of Section 9 (pilot → case study → longitudinal). Note that the
benefit of the participation of marginalized groups, including problem perception as a person
directly affected (Definition 6), has the thinnest peer-reviewed backing of the three strata, and
— including connection to prior practice in inclusive design — it is a task for future research.
Proposition 12 (Superiority of dynamic role allocation)
Fixed role allocation based on chronological age is inferior to dynamic role allocation
based on measured characteristics (Definition 1). The grounds are three: (i) intraindividual
variation in cognitive characteristics, (ii) the magnitude of the overlap of
distributions across age groups, and (iii) the malleability and trainability of
characteristics. Measurement-based allocation, however, is structurally exposed to
gaming under Goodhart's law. The operational requirements of Definition 1 therefore
include the aperiodic replacement of measurement tasks; ex-post verification
(backtesting) by unmediated audit of real work logs rather than one-off test scores; and
real-time blind in-situ verification without advance notice (unannounced in-situ
verification) — the latter constituting a double barrier against not only adaptation to the
tests but the gaming of "ways of keeping logs that backtesting cannot detect" itself.
Refutation condition: Comparative evidence that age-fixed allocation achieves outcomes and
welfare equal to or better than dynamic allocation, or evidence that the costs of measuring
characteristics, mismeasurement, and gaming exceed the gains of dynamic allocation.
Proposition 12 is the proposition that justifies the operation of Definition 1, and its three
grounds are all supported by the evidence of Section 2. (i) Cognitive characteristics vary over
time even within the same individual, and each ability traces a different trajectory (the
asynchrony of peak ages: Hartshorne & Germine 2015). (ii) The distributions of age groups
overlap widely, and differences in group means are too coarse to use for predicting
individuals (the effect sizes of Salthouse 2009 are likewise not large enough to erase the
overlap of distributions; Section 2.2). (iii) Abilities are malleable through compensation and
training, supported by the lineage of SOC theory (Section 2.3) and the intervention evidence
of reducing age discrimination through manager training (Wegge et al. 2012). Age-fixed
allocation appears rational only when all three points are ignored.
That the grounds are empirically established, however, is distinct from the superiority of
dynamic allocation being empirically established. No evidence exists comparing the outcomes
and welfare of age-fixed and dynamic allocation, and the superiority claim itself is untested.
Moreover, the second limb of the refutation condition — the possibility that the costs of
measurement, mismeasurement, and gaming exceed the gains — is a serious implementation
issue. Measuring cognitive characteristics involves cost and error; if measurement is tied
Working Paper | Ageless Management in the AI Era 80
directly to treatment, incentives to manipulate scores arise, and the measurement apparatus
itself can turn into a new apparatus of selection and discrimination (Section 10.3). By building
this possibility into the refutation condition, Proposition 12 presents the superiority of
dynamic allocation as an empirical claim conditional on the quality of the measurement
regime. Identifying the design conditions of a measurement regime under which the
superiority holds is a task for Section 9 and future research.
The proposition text writes the operational requirements into the proposition itself because
gaming is not a peripheral implementation problem for dynamic allocation but a structural
vulnerability. Under a regime in which measurement is tied directly to treatment, the
indicator becomes a target of optimization, and a targeted indicator loses its correspondence
with what it measures — when a measure becomes a target, it ceases to be a good measure, in
the formulation of Goodhart's law (Strathern 1997). Whereas age-fixed allocation is nearly
immune to gaming precisely because chronological age cannot be manipulated, allocation
based on measured characteristics is structurally exposed to this law in the form of score
inflation through coached preparation and overfitting to the measurement tasks. The
proposition text specifies three barriers as operational requirements of Definition 1. First, the
aperiodic replacement of measurement tasks — the more fixed and predictable the tasks, the
more investment in overfitting pays off, so the very irregularity of replacement lowers the
expected return on overfitting. Second, ex-post verification by unmediated audit of real work
logs (backtesting) rather than one-off test scores — a test result at a single point can be
prepared for, but the cost of continuously falsifying a running record of the quality of
judgment in real work is high, and unmediated audit (Section 8.1) bypasses window-dressing
at the stage of summarizing and filtering the logs. Third, real-time blind in-situ verification
without advance notice (unannounced in-situ verification). This third barrier is necessary
because the second barrier can itself become a new target of optimization. What backtesting
verifies is the logs that were kept, and as long as the manner of keeping logs — what to
record, what not to record, at what granularity — is under the control of the person being
allocated, a deeper level of gaming becomes possible: adaptation toward "ways of keeping
logs that backtesting cannot detect." Unannounced in-situ verification bypasses the mediating
layer of logs altogether by observing real-time judgment itself, blind and without advance
notice, rather than via records. The second and third barriers thus constitute a double barrier
against two levels of gaming: adaptation to tests (addressed by the first and second) and
adaptation of log formation itself (addressed by the third). With these three requirements, the
superiority claim of Proposition 12 is explicitly conditioned on the existence of a gamingresistant
measurement regime. The residual risks of gaming and measurement misuse that
remain nonetheless are treated self-critically in Section 10.3.
The operation of unannounced in-situ verification carries a principle of separating gaming
deterrence from surveillance pressure on individuals. What the third barrier primarily
calibrates is system-level measurement accuracy — a test of the divergence between
backtesting results and real-time observation — not the condemnation of individuals. The
Working Paper | Ageless Management in the AI Era 81
results of unannounced verification are therefore used, in the first instance, in anonymized
and aggregated form for calibrating the system. Their use for individual treatment (such as
changes in role allocation) is limited to cases that pass through aggregation over long
windows and due process including disclosure to the person and the opportunity to respond;
no demotion or revocation of authority is carried out on the basis of a single unannounced
result. The deterrent effect on gaming is achieved by making known the fact that the system
is calibrated, and does not require diverting individual observations to individual
surveillance. Unannounced verification lacking this separation degenerates into precisely the
surveillance pressure that Section 8's Brain Safety must exclude, and destroys, upstream of
measurement, the flattening of power gradients that supports the reach of decorrelated
observations to decision-making (the organizational condition of Proposition 5). Operational
details are placed in Section 8.1.
The prior literature to which Proposition 12 should be connected is made explicit: the
economics of statistical discrimination (Phelps 1972; Aigner & Cain 1977). The conditions
under which allocation based on a coarse proxy such as age can be rational — when
observing an individual's true productivity is costly and measurement carries error, using
group statistics as a proxy can be second-best optimal under informational constraints — are
precisely the subject this literature formalized. Within this framework, Proposition 12 is
positioned as the claim that "falling measurement costs and rising measurement accuracy
undermine the rationality conditions of proxy use." In an environment where measuring
cognitive characteristics is costly and error-prone, reliance on the proxy variable of age can
persist by the logic of Aigner & Cain, and the second limb of Proposition 12's refutation
condition — evidence that the costs of measuring characteristics, mismeasurement, and
gaming exceed the gains of dynamic allocation — takes over, as its refutation condition,
exactly the rationality conditions of proxy variables that this literature formalized. That is,
Proposition 12 does not deny the theory of statistical discrimination; it asserts the superiority
of dynamic allocation as a function of the theory's conditions, and whether the superiority
holds is an empirical question depending on the state of measurement technology and the
measurement regime.
5.4 Table 7: The Map of Propositions
Table 7 lists the content of the twelve propositions, the evidence they rest on, their
evidentiary status, and the means of testing them. Evidentiary status is graded in three
categories: "established" means directly supported by independent peer-reviewed empirical
research; "suggestive" means indirect or consistent evidence exists but direct testing is
lacking; "untested" means a theoretical claim of this paper lacking direct evidence. Composite
propositions are graded part by part.
Table 7 The map of propositions: evidentiary status and means of testing
Proposition Summary of content Evidence relied on
Evidentiary
status Means of testing
Working Paper | Ageless Management in the AI Era 82
1 Marginal
value shift
AI lowers the marginal
cost of Gf-type tasks and
raises the relative
marginal value of Gc-type
tasks (conditional on
complementarity and the
effectiveness of detection
= Propositions 3 and 4)
Sections 4.1
(compression
evidence) and 4.2
(oversight failures)
Cost-decline
side
established /
value-rise
side
suggestive
(conditional)
Future research
with labor-market
data (observation
window: roughly
10 years after
diffusion)
2(a) Existence
of the
bottleneck
The Gf bottleneck exists
as an independent exit
pathway alongside
mandatory retirement,
health, and
discrimination
(quantification of its
share is an open
empirical question)
Section 2 (Gf/Gc
divergence =
structural
possibility)
Untested Exit-reason
decomposition
studies (future
research)
2(b) Removal of
the bottleneck
AI complementation of
Gf-type components
removes the bottleneck
and lets Gc-type abilities
reach the work
(boundary conditions:
maintenance of
unassisted exercise and
unmediated audit /
cognitive-screening lower
bound; also conditional
on time elapsed since
exit)
Sections 4.3 (skill
depreciation) and
4.5
(reinterpretation of
headwind data)
Untested Experiments and
quasi-experiments
providing Gfcomplementing
tools (future
research)
3 Noncompressibility
of experience
Contemporaneous AI
assistance compresses
differences in production
but not differences in
verification
Section 4.1
(production side);
Sections 4.2–4.4
(indirect evidence
on the verification
side)
Production
side
established /
verification
side untested
H3 (2×2 factorial
design) +
measurement of
the longitudinal
compression
pathway (Section 9)
4 Formation of
experiential
audit capacity
Experiential audit
capacity is formed not by
chronological age but by
the interaction of domain
knowledge and
operational schemata
with generalized Gc and
metacognition (boundary
conditions: confirmation
bias (i), the half-life of
knowledge (ii),
processing fluency (iii))
Sections 4.2 and 4.4
(the bias synergy;
absence of age-axis
empirical work =
research gap);
processing-fluency
findings (Reber &
Schwarz 1999;
Alter &
Oppenheimer 2009)
Untested H3 (separate
measurement of
generalized Gc and
domain
knowledge;
confirmationconforming/
deviating tasks;
fluency control) +
extension varying
experience and age
independently
(Appendix A,
future research)
Working Paper | Ageless Management in the AI Era 83
5 Generational
decorrelation
Under three conditions —
channel (raw audit of
random samples),
organization (flattened
power gradients and
reach of observations to
decision-making), and
model (AI-side crossverification)
—
generationally
heterogeneous
supervisor groups have
less overlap in misses,
wider detection sets, and
detections that reach
decision-making
WP8 framework
(inherited); BCM
2026e (Belonging);
Sommers (2006)
(indirect)
Untested (the
extension to
the
generational
axis is this
paper's new
claim)
H3 overlap analysis
+ oversight
experiments
measuring audit
channel and
adoption rates
(Section 9.3, future
research)
6 AI-mediated
diversity effect
Given task and climate
conditions, AI-mediated
complementarity is an
additional moderator of
the diversity-outcome
relation, moving it in a
more positive direction
(outcomes evaluated as
net benefit after
deducting decision-time
costs)
Section 5.3 (nearzero
meta-analytic
baseline and
condition
dependence)
Baseline and
condition
dependence
established /
AI-mediated
moderation
untested
H2 (mixed-team
experiment,
including
decomposition of
simple main effects
and recording of
cost variables)
7 Cognitive
engagement
pathway
If a brain-health effect of
work exists, cognitive
engagement mediates it
Section 3
(heterogeneity of
effects; Experience
Corps)
Suggestive H1 (mediation
analysis +
identification
through exogenous
variation)
8 Bidirectional
capital
formation
Ecosystems satisfying exante
verifiable conditions
(satisfaction of
Definitions 7 and 8 +
checklist operation of a
cognitive-apprenticeship
developmental pathway
with micro-ownership;
simulated audit without
responsibility does not
qualify) increase brain
capital K × u across all
participating strata (K is a
hierarchical function in
which domain
knowledge enters
multiplicatively above a
gate of foundational
cognitive function)
Section 5.3
(knowledgetransfer
and
reverse-mentoring
evidence); WP5
(inherited)
Motivation
and skill side
suggestive /
brain-capital
indicators
untested
H1 and H2 +
organization-level
indicators (Section
9.3)
Working Paper | Ageless Management in the AI Era 84
9 Protection
vacuum of nonemployment
forms
Non-employment
participation falls into a
vacuum of institutional
protection and, if left
unaddressed, can
degenerate into
exploitation
Section 7
(institutional
comparison based
on primary legal
sources)
Suggestive
(institutional
analysis)
Comparative legal
research and crossjurisdiction
comparison (future
research)
10
Transferability
of the design
principles of
protection
The grounds of
protection are stratumspecific
(the
developmental principle
for youth), but commonstructure
design
principles (priority of
health, ceilings on load,
adequacy of
compensation) are
transferable across strata
Sections 7 and 8
(analysis of
common structure
and stratumspecific
principles)
Untested Legal-theoretical
examination and
institutional
comparison (future
research)
11 Conversion
of social
problems into
resources
The three social
problems convert into
sources of brain capital
conditional on Definition
6 and Definition 8
operated in ex-ante
verifiable form
Integration of
Propositions 2–10
(depends on
upstream
propositions)
Untested H1–H3 +
implementation
roadmap (pilot →
case study →
longitudinal)
12 Superiority
of dynamic role
allocation
Age-fixed allocation is
inferior to measurementbased
dynamic allocation
(operational
requirements: as barriers
against Goodhart's law,
aperiodic replacement of
measurement tasks,
backtesting of real work
logs, and the double
barrier of unannounced
in-situ verification
without advance notice)
Section 2 (grounds
(i)–(iii)); Strathern
(1997)
Grounds
established /
allocation
comparison
untested
Comparative
experiments and
field studies of
allocation schemes
(future research)
Note: The grading of evidentiary status is this paper's own. "Established" means that independent peer-reviewed
empirical research exists for the part in question, not that the proposition as a whole has been established. The
content of H1–H3 is detailed in Section 9, and the experimental protocol of H3 in Appendix A.
5.5 Relations among the Propositions and the Structure of Refutation
The twelve propositions are not parallel; they form a structure with dependencies. Making
that structure explicit amounts to specifying which parts of the theory collapse when which
proposition falls — that is, to writing the refutation condition of the theory as a whole. The
structure can be organized into five layers. The core is Propositions 3 and 4 (the reality and
attribution of experiential audit capacity); their premise connection is Propositions 1 and 2
Working Paper | Ageless Management in the AI Era 85
(value shift and reach); the extension to organizational design is Propositions 5 and 6
(decorrelation and the mediated effect); the extension to welfare and capital formation is
Propositions 7 and 8; the institutional conditions are Propositions 9 and 10; and then the
integration is Proposition 11 and the operation is Proposition 12.
The core of the theory is Propositions 3 and 4. If Proposition 3 falls — if the experience
difference in contextual-error detection disappears under AI assistance — experiential audit
capacity (Definition 4) is definable as a concept but economically worthless. If AI compresses
even differences in verification capacity, the reason to seek the supply side of oversight in
holders of experience disappears, Proposition 5's decorrelation argument loses its supply-side
meaning, and Proposition 11's resource conversion loses its core pathway. Proposition 1 itself
survives, but the raised marginal value is no longer attributed to holders of experience, and
this paper's theory dissolves into a generic "theory of oversight in the age of AI." The fall of
Proposition 4 is graver still. If chronological age independently predicts detection capacity
even after controlling for years of experience, then the reattribution from age to experience
(the first theoretical operation) is wrong, and the very name of the theory, "ageless," fails to
hold. The theory then degenerates into one of the age-based arguments this paper rejected in
Section 1 — celebration of seniors or exclusion of seniors. Conversely, if years of experience
lose predictive power after control, the reattribution that assigns the variance of detection
performance to experience loses support, and even if Proposition 3 held, its bearers could not
be identified. It is because of this structure that the testing of Propositions 3 and 4 (H3) is
placed at the top priority of this paper's verification plan (Section 9).
The premise-connecting Propositions 1 and 2 are the conditions for the core to have value.
If Proposition 1 falls — if the relative marginal value of Gc-type tasks does not rise — then
even if experiential audit capacity is real, it is a capacity that does not become scarce, and
Ageless Management may remain an ethical imperative but does not stand as a claim of
managerial rationality. If Proposition 2(b) falls — if AI's Gf-type complementation does not
improve the job performance and continuation of older workers — Gc-type abilities remain
held without reaching the work, and the capital formation of Proposition 8 lacks supply. The
rejection of Proposition 2, however, does not directly wound the core (Propositions 3 and 4):
even if the pathway of reach is blocked, the audit value of those who have reached is an
independent question. In that case, the theory survives in reduced form, shrinking from a
"theory of expanded participation" to a "theory of role reallocation among those already
participating."
The organizational-design Propositions 5 and 6 are the layer that extends the theory from
the individual to the group level. Even if Proposition 5 falls — even if generational
heterogeneity does not supply error decorrelation — individual-level experiential audit
capacity (Propositions 3 and 4) is unscathed, and supervisors may be selected not by
generational composition but by measuring individual detection capacity. What is lost is the
oversight rationale for "multigenerational" as an organizing principle, and the organizational
design of Section 6 (Section 6.3) requires substantial revision. If Proposition 6 falls, the
Working Paper | Ageless Management in the AI Era 86
performance effect of diversity returns to the meta-analytic baseline (near zero), and the
outcome-side justification of the multigenerational ecosystem is confined to the oversight
value of Proposition 5 and the knowledge transfer and capital formation of Proposition 8. If
Propositions 7 and 8 fall, the brain-health and human-capital benefit claims disappear, but the
economic core stands independently. Independence in the reverse direction can be claimed
only more narrowly. Proposition 7 is a claim about the mediation structure of the brainhealth
effect of work and can be tested independently of any of Propositions 1–6. But
Proposition 8 is not so. The restrained depreciation of K among super-seniors depends on the
continued exercise of Gc-type roles (the pathway of Proposition 7), and reaching those roles
depends on the bottleneck removal of Proposition 2. In a world where Proposition 2 has
fallen, the capital formation of Proposition 8 lacks supply (previous paragraph). Therefore, if
Propositions 1–6 all fall, what remains unscathed is Proposition 7 alone, and the theory is
then lost as a "theory of managerial rationality," shrinking to one hypothesis in health and
welfare research on the quality of work and cognitive engagement. The economic strand (1–
6) and the welfare strand (7 and 8) are not two independent pillars; what can stand
independently is confined to the single support of Proposition 7.
The institutional Propositions 9 and 10 and the operational Proposition 12 govern not the
truth of the theory but its implementability. The refutation of Proposition 9 means, as noted
above, the realization of a desirable institutional state, and does not wound the theory. The
refutation of Proposition 10 pushes Brain Safety back to stratified design, but the need for
protection itself is unmoved. The refutation of Proposition 12 is grave. If the gains of dynamic
allocation are eaten up by measurement costs, mismeasurement, and gaming, Definition 1 is
unimplementable, and this paper's theory can maintain the critique that "age criteria lack
cognitive-scientific grounds" (Section 2) but fails to present an alternative. Finally, since
Proposition 11 is the integration of the upstream propositions, it is hard to refute in isolation;
instead, the rejection of any upstream proposition propagates to it in tandem. In sum, the
single vital point of this theory is Propositions 3 and 4, and Hypothesis H3 is designed to
strike this vital point directly (Section 9, Appendix A). For the proponent of a theory to
identify its most fragile point and direct the first test at it — that is the minimum
responsibility owed by a section that lines up twelve untested propositions.
6. Designing the Multigenerational Ecosystem
The definitions and propositions of the preceding section are, in themselves, abstract
theoretical apparatus. This section translates them into the language of organizational design:
how to configure the initial placement of the roles of the four players constituting the
multigenerational ecosystem of Definition 7 (Section 6.1); how to keep that placement in
motion without letting it degenerate into fixation by age (Section 6.2, operationalizing
Proposition 12); how to design the generational composition of the supervisor group (Section
6.3, operationalizing Proposition 5); how to condition the connections with NPOs and
Working Paper | Ageless Management in the AI Era 87
educational institutions outside the boundary of employment (Section 6.4); and where this
design does — and does not — connect to Future Value Theory (FVT) (Section 6.5).
The character of this section should be delimited in advance. It is a normative design
argument, and much of its basis consists of the untested propositions of Section 5. The parts
that are empirically established (the evidence on compression, the baseline of the diversity
meta-analyses, the evidence on conditional moderators) and the parts that this paper asserts
theoretically (AI-mediated complementarity, the oversight value of generational
decorrelation) are distinguished explicitly, subsection by subsection. Not allowing the design
argument to be read as an "established prescription" is the descriptive discipline of this
section.
6.1 Reconfiguring the Roles of the Four Players
Table 5 organizes, for the four players constituting the multigenerational ecosystem, the
bottlenecks that conventional job design has imposed, the content of complementation by AI,
and the default roles in the ecosystem. The meaning of the word "default" should be fixed
first. The role placement of Table 5 is a design of initial values derived from the distribution
of cognitive assets statistically expected to be relatively more prevalent in each group
(Section 2); it is not a rule that determines an individual's role from age. As confirmed in
Section 2, the distributions of cognitive characteristics across age groups overlap
substantially, and every age group contains a substantial number of individuals whose
characteristic profiles correspond to other rows of the table. The distinction between default
design and fixed roles is a principle running through this entire section, and its operation is
the dynamic role allocation of the next subsection.
Table 5 Reconfiguration of the four players' roles in the multigenerational ecosystem (default
design)
Player
Conventional
bottleneck
Complementation
by AI
Default role (initial
placement) Design cautions
Super-seniors
(aged 60 to
their 90s)
Difficulty in
performing the
Gf-type
components of
the job bundle
(processing
speed, working
memory,
learning of novel
procedures)
forces exit from
the job as a
whole (Definition
3); barriers to
digital operation
Substitution for and
complementation of
Gf-type components
through search,
aggregation,
summarization,
documentation, and
navigation of
operating procedures
Experiential
auditing: detection
of contextual errors
and practical and
ethical risks in AI
output (Definition
4), contextual
evaluation,
matching against
failure cases
Auditing is a
consumed resource
subject to capacity
constraints (WP8;
Section 8). Empirical
demonstration of
detection probability
is the object of
testing under H3.
Fixation into the
auditor role is itself
a new form of
stereotyping (Section
10). Securing
opportunities for
Working Paper | Ageless Management in the AI Era 88
unassisted exercise
(Sections 4.3 and 8)
Youth (under
20)
Lack of practical
experience and
domain
knowledge; long
apprentice
periods in
onboarding
Immediate supply of
past cases and
procedural
knowledge;
execution support for
generation and
analysis (experience
compression; Section
4.1)
Rapid prototyping:
iteration of trial and
error, bringing in a
feel for novel
technological
environments
Primacy of
educational purpose
and precedence of
schooling
(Proposition 10;
Section 7). Protection
of unassisted
learning. AI use
without guardrails
can harm
independent
learning (Bastani et
al. 2025)
Socially
marginalized
groups (via
NPO and
similar
partnerships)
Restricted
employment
opportunities,
accessibility
barriers, lack of
practical training
opportunities
Support for job
performance through
task subdivision and
multimodal work
support by voice,
image, and other
channels (evidence
on effectiveness in
Section 6.4)
Supplying problem
perception as those
directly affected:
surfacing
overlooked
problems and latent
needs
AI does not advance
inclusion
unconditionally
(Omri et al. 2025;
Section 6.4). Bias in
AI hiring tools (El
Morr et al. 2024).
Confinement to the
"insider-perspective"
role trivializes
participation
The midcareer
generation
(20s–50s)
Depletion of
cognitive
resources under
the double
burden of
execution and
management
Relief of load
through automation
of routine
management tasks
and reporting
Orchestration:
integrating
multigenerational
insight and AI
output, crystallizing
them into business
decisions
The integrator is
itself subject to
automation bias and
the constraints of
oversight span
(WP8). Risk of
depreciation of one's
own skills through
AI dependence in
execution (Section
4.3)
Note: This table reconstructs the role scheme of the original draft on the basis of the definitions of Section 5 and the
empirical qualifications of Section 4. Roles are defaults (initial placements), not fixed. The role placement in each row
is a design initial value based on distributional tendencies of age groups; an individual's placement is determined by
measured characteristics, experience, condition, and volition (Section 6.2). The "design cautions" column was added to
make explicit that the propositions each role relies on are untested and that role fixation is a risk.
What Table 5 inherits from the original draft's scheme is the four-player configuration;
what it changes is three points. First, the role of the super-seniors was changed from
expressions of status such as "meta-auditor and arbiter of value" to the functional expression
of experiential auditing. The ground of audit value is neither age nor status but the
interaction of accumulated domain experience with Gc-type abilities and metacognition
(Definition 4; Proposition 4), and its detection probability remains unestablished (H3).
Working Paper | Ageless Management in the AI Era 89
Second, a "design cautions" column was attached to each role. As seen in Section 4, the
benefits of AI complementation are not unconditional, oversight is failure-prone, and
dependence on assistance can depreciate skills. Role design deserves the name of design only
when it prices in these headwinds. Third, for the AI complementation of socially
marginalized groups, the original draft's assertion that "participation becomes possible" was
weakened into a conditional formulation. The grounds are presented in Section 6.4.
On the default role of the youth, one change in the environment must be added.
Conventionally, the organizational entry of the young presupposed a pathway that began
with routine entry-level tasks and accumulated domain knowledge through apprenticeship.
But precisely those routine entry-level tasks stand at the front line of substitution by AI, and it
has been pointed out that youth may be vulnerable to the effects of AI adoption alongside, or
even more than, older groups (Ayalon 2026 — a non-systematic review, cited as a prior
example of problem framing). This asymmetry is an exogenous given for the design argument
of this section. If the conventional entry pathway of apprenticeship thins out, the pathway
connecting youths' cognitive assets (adaptation to novel technological environments, less
constrained ideation) to the ecosystem will cease to exist unless deliberately designed. The
youth row of Table 5 is, in this sense, a proposal for the reconstruction of an entry pathway,
not an endorsement of the status quo. Its institutional mode (PBL-type; primacy of
educational purpose) is treated in Section 6.4 and Section 7.
The theoretical claim of this paper is that the interaction of the four players becomes
connectable only through AI-mediated complementarity (Definition 6). Conventionally,
contextual verification capacity, execution capacity, and problem perception were required to
be combined within the same individual. That is why the "mid-career worker with both
experience and execution capacity" was made the core of the organization, while holders of
experience only, or of perception only, were marginalized. When AI substitutes for the Gftype
components of task bundles, this combination requirement loosens, and a division of
labor that connects heterogeneous cognitive assets across individuals becomes technically
possible. As Definition 6 makes explicit, this connection is not a one-way prosthesis: it is
bidirectional in that the benefits of complementation (substitution of Gf-type components)
and the supply of auditing (provision of Gc-type verification) flow mutually among
participants (bidirectional complementarity). That said, "becoming possible" and "producing
results" are separate propositions, and the latter hangs on the test of Proposition 6 (Section
6.5).
The context-partition protocol for hybrid domains. The orchestration of the mid-career
generation explicitly includes a design task: the assignment of audit targets. Boundary
condition (ii) of Proposition 4 warned that in domains where technological change is fast and
the half-life of domain knowledge is short, the audit effectiveness of accumulated experience
declines and can turn negative. But this boundary condition cannot be operated through the
crude rule "do not use experiential auditing on fast-changing projects," because real projects
belong neither to purely fast domains nor to purely slow ones: they are hybrid domains.
Working Paper | Ageless Management in the AI Era 90
Within a single business proposal or deliverable, slowly changing components — legal and
governance matters, the configuration of relationships and interests, industry practice — and
fast-changing components — the specifications of frontier technologies, tool environments,
the capability range of models — are inseparably mixed. The unit of assignment is therefore
the component, not the project. The orchestrator decomposes the output under audit into
components with different speeds of change, assigns experiential auditing to the components
with long half-lives, assigns the components with short half-lives to auditing by younger and
mid-career members and to diversity on the AI side — cross-validation by different
foundation models, as the model condition of Proposition 5 requires — and evaluates each
component separately. This separated evaluation prevents erroneous auditing of experience
on the fast components (overconfidence in a previous generation's specifications) from
contaminating the detection value on the slow components, and vice versa. This protocol is
thus an operational response to boundary condition (ii) of Proposition 4: rather than leaving
the boundary condition as a limit of the theory, it converts it into a design variable, the
assignment of audit targets. Decomposition into components is itself, however, work that
requires judgment, and erroneous decomposition can become a new pathway of misses. The
effort decomposition requires and the risk of misdecomposition are entered as verification
costs in the net-benefit ledger of Section 6.5.
6.2 Dynamic Role Allocation — Operationalizing Proposition 12
Proposition 12 asserts that fixed role allocation based on chronological age is inferior to
dynamic role allocation based on measured characteristics (Definition 1). Its grounds were
three: (i) the intraindividual variability of cognitive characteristics, (ii) the magnitude of
distributional overlap across age groups, and (iii) the malleability and trainability of
characteristics. This section develops the proposition into operating rules for allocation.
The core of the operation is to allocate roles not by age but by four variables. First,
measured cognitive characteristics: describe the job bundle by the relative weights of its Gftype
and Gc-type components (Definition 2) and match it against the individual's
characteristic profile. Second, accumulated experience: the allocation criterion for audit-type
roles is the length and quality of domain experience, not chronological age (Section 4.4;
Proposition 4). Third, condition: cognitive characteristics show intraindividual variation
within the day and the week, and the same individual has states suited to auditing and states
unsuited to it. Allocation is not a qualification fixed by a single measurement but an
assignment updated according to state. Fourth, the person's own volition: unless roles are
something the person can re-choose, measurement-based allocation turns into an apparatus
of selection (Section 10).
If this operation functions properly, the correspondence between the rows of Table 5 and
individuals becomes fluid. A person who moved to a different industry at 60 is on the
inexperienced side in that domain and enters through the roles of prototyping and learning.
A practitioner in their 30s who has trained in a single domain since their teens can hold high
Working Paper | Ageless Management in the AI Era 91
experiential audit capacity (the thought experiment of Section 4.4). Participants with large
fluctuations in condition have their audit tasks reassigned to lighter-load times and formats.
Only when Table 5 is operated in this way is the default design distinguished from fixed roles.
A word on the unit of allocation. The unit of dynamic role allocation is not the "job" but the
task or session carved out of the job bundle. That the same person handles audit tasks in their
own domain in the morning and, on another day, takes the prototyping side in an adjacent
domain where their experience is shallow, is if anything a natural consequence, given that
Definition 2 defines Gf-type and Gc-type as relative weights on a continuum. Nor is the
direction of transition one-way. In place of the seniority-based single-track path of "execution
→ management → audit," the ecosystem permits multidirectional transitions, including
returns from auditing to execution, concurrent holding of execution and auditing, and
increases and decreases in the volume of participation. In particular, a design in which those
placed in audit roles are severed completely from execution undermines the very basis of
experiential audit capacity, by the skill-depreciation logic seen in Section 4.3. Transition
possibility is a requirement of fairness and, at the same time, a requirement of capacity
maintenance (Section 8).
Multi-role concurrency (identity buffer). Since the unit of allocation is the task or
session, one further operating principle can be derived: changes of allocation are operated
not as "demotion" from a single role but as reweighting of a portfolio premised on the
simultaneous holding of multiple roles. That is, every participant always holds concurrently
more than one of the roles of auditing, learning, and execution, and allocation changes driven
by measurement, condition, and volition are implemented as adjustments of the weights
within that role portfolio — lowering the weight of auditing while raising the weights of
learning and execution, and so on. This principle is needed as a response to an
organizational-psychology risk. Under a regime in which a participant's occupational identity
is unified into a single role — "I am the auditor" — a measurement-based change of allocation
is experienced by the person as the stripping of that single identity: being "removed from the
auditor role." A change experienced as role deprivation readily invites loss of self-efficacy, the
stigma of perceived decline from those around, and quiet disengagement in which the person
formally continues to participate while withdrawing inner engagement. This is nothing other
than damage to the pathway Proposition 7 identified — if a brain-health effect of work exists,
what mediates it is cognitive engagement. If dynamic role allocation is operated in a way that
destroys participants' engagement, it raises the precision of allocation while undermining the
very thing allocation is meant to protect — cognitive engagement and, beyond it, brain
capital. Multi-role concurrency is, in this sense, not a device of efficiency but a design for
protecting engagement, and it is included among the operating requirements of dynamic role
allocation as a buffer (identity buffer) that absorbs the psychological cost of allocation
changes.
The bidirectional feedback protocol. The operation of dynamic role allocation also
includes the design of how information flows between roles. This paper places the destination
Working Paper | Ageless Management in the AI Era 92
of experiential-audit feedback, as a rule, on the AI side rather than on persons. That is, the
contextual errors and risks detected by auditing are reflected, as the basic form, not as direct
instruction or correction of younger individuals but as indirect feedback into the revision of
the AI's prompts, evaluation criteria, and verification procedures. There are three reasons.
First, the object of auditing is AI output, not persons (Definition 4). An operation that diverts
detection results into the evaluation of individuals degrades auditing into over-surveillance
and destroys the psychological safety of those audited. Second, if direct instruction from
seniors to juniors were made the basic form, the division of labor in Table 5 would harden
into a one-way relation of authority between "the generation that audits" and "the generation
that is audited." A reverse-seniority relation in which junior execution constantly passes
through senior approval is a ready source of intergenerational friction. Third, reflection into
the system side prevents the recurrence of detected errors through design improvement
rather than individual attention, which is also consistent with the suppression of alarm
fatigue (Section 8).
This protocol is not complete in one direction. Its counterpart is reverse mentoring from
juniors to seniors — the operation of AI tools, trends in new models and technological
environments, sharing of the information environment of younger generations. That reverse
mentoring can yield skill development for both younger mentors and older learners has peerreviewed
evidence (Kaše et al. 2019; Section 6.5). The pair of indirect auditing and reverse
mentoring implements, at the level of information flows, the bidirectional complementarity
(Definition 6) in which the benefits of complementation and the supply of auditing flow
mutually, and is a preventive design against the audit regime sliding into a structure of
intergenerational surveillance and conflict.
The cognitive apprenticeship loop — micro-ownership. The cognitive apprenticeship
that Proposition 8 requires as an ex-ante verifiable condition — a development pathway in
which youth experience Gc-type roles in stages, in small, low-risk projects, with full
delegation and acceptance of responsibility for failure (micro-ownership) — is implemented
within the operation of this section as a development track. The core of the implementation is
the combination of the smallness of the risk and the realness of the responsibility. That is,
small, low-risk projects whose erroneous consequences are limited for the organization —
improvements to internal processes, pilots within a restricted customer scope, small-budget
initiatives — are carved out as units; within that scope, full authority over judgment,
including auditing as Human on the Loop (HOTL), is delegated to the junior, and the
consequences of failure — rework, explanation to stakeholders, recovery — are borne by the
person. Verification exercises with known embedded errors (Section 8.1) are useful as a
calibration device at induction, but they do not themselves constitute the development
pathway. As Proposition 8 expressly excludes, simulated auditing without responsibility does
not cultivate genuine metacognition — where the consequences of judgment do not come
back to oneself, the metacognitive calibration loop that maps judgments to consequences
does not close, and what forms is limited to knowledge of the formal procedures of auditing
Working Paper | Ageless Management in the AI Era 93
(Section 5, commentary on Proposition 8). The development track is completed by
progressively raising the scale and risk level of the delegated projects, together with feedback
on detection performance and the consequences of failure. This implementation is needed
because a division of labor that fixes Table 5's initial placement (juniors = prototyping,
seniors = auditing) for single-period efficiency carries a diachronic risk. If Proposition 4 is
correct, experiential audit capacity is the product of long-term domain experience, and a
fixed division of labor deprives youth of the pathway of judging, failing, and bearing the
consequences through which Gc forms, destroying the future supply of experiential audit
capacity itself — a diachronic trade-off between present efficiency and future supply (Section
5; Proposition 8). The cognitive apprenticeship loop is an operational response to this
diachronic risk, and it makes explicit that dynamic role allocation is not only single-period
matching of person and place but also a pathway of capacity formation.
Responsibility layering. The cognitive apprenticeship loop requires an explicit
demarcation concerning where responsibility lies. Even in a micro-ownership domain fully
delegated to a junior, senior experiential auditing — involvement as HOTL — can exist; but if
the mode of that involvement is left unbounded, the delegation is hollowed out. If senior
detection is operated as a veto over the junior's decisions, final decision authority reverts in
effect to the senior, and the core of micro-ownership — the experience of judgment that
carries responsibility — is lost. Conversely, if senior involvement is construed as shared
responsibility, a moral hazard arises on the junior's side — "the senior saw it, so it is not my
responsibility" — and the mapping of judgment to consequence — the metacognitive
calibration loop — again fails to close. Further, over-inhibition, in which a junior constantly
conscious of the senior's gaze optimizes for failure avoidance, damages the very purpose of
this track, trial and error. This paper therefore layers responsibility. Senior HOTL in the
junior's micro-ownership domain is confined not to a veto that halts decisions but to Socratic
coaching that prompts the junior's own reconsideration through questions — "What happens
if this premise fails?" "Who will need an explanation?" — while final decision authority and
the attributed responsibility for its consequences remain with the junior. Whether to adopt
the senior's observations is also the junior's decision. This demarcation prevents inhibition
because there is no veto, and prevents moral hazard because responsibility is not shared.
This demarcation of responsibility among participants is, moreover, not in conflict with the
external attribution of responsibility treated in Section 7 — the contractual cap on liability
that precludes shifting responsibility onto auditors and attributes final risk to the business
entity. The former is a demarcation of development and evaluation inside the ecosystem; the
latter is the allocation of legal risk in relations with third parties — a problem of a different
layer. What micro-ownership's "acceptance of responsibility for failure" means is the internal
acceptance of consequences — rework, explanation to stakeholders, recovery — not
unlimited external liability for damages.
The greatest design risk is that dynamic role allocation reproduces age stereotypes. When
Table 5's initial values — "seniors audit, juniors prototype" — slide in operation into the rule
Working Paper | Ageless Management in the AI Era 94
"audit because senior, prototype because young," Ageless Management becomes ageless in
name only. This paper specifies four design principles against this slide. First, do not close the
entrance by age: define the placement requirements for audit roles by measured experience
and characteristics, and use age neither as a requirement nor as reference information.
Second, do not impose accountability for deviation from the default: an operation that
demands justification only from deviators (juniors who audit, seniors who prototype) is de
facto fixation. Third, guarantee the transparency of measurement and a channel of appeal:
measurement-based allocation carries risks of measurement error, gaming, and
discriminatory diversion (the refutation condition of Proposition 12; Section 10). Fourth,
prepare the climate conditions. One of the few robust findings of age-diversity research is
that effects are conditioned on climate. Wegge et al. (2012), from data on more than 745 teams
and 8,848 persons across three industries, identified as conditions for age diversity to
function, in addition to high task complexity, low age stereotyping and age discrimination
and positive appraisal of diversity. Li et al.'s (2021) survey of 3,888 establishments likewise
reports that, absent age-inclusive management, the effects of age diversity are nonsignificant
or negative. Dynamic role allocation is an operation that functions only on top of these
climate conditions.
A direction on the practice of measurement should also be indicated. The measurement
dynamic role allocation requires is not a one-shot qualifying examination for selection but
continuous, low-burden monitoring in support of placement. The candidates are a
combination of structured description of domain experience (domains, years, types of failure
cases handled), standardized measurement of Gc-type abilities and metacognition, work
samples (performance on actual audit and prototyping tasks), and the person's stated
preferences; their indicator construction at the organizational level is treated in Section 9.
What must be stressed here is the restriction on the use of measurement. The same
measurement, used for placement support, becomes the foundation of dynamic role
allocation; used for selection and exclusion, it becomes an apparatus of selection more
precise than the age criterion that Definition 1 rejected. This danger is confronted head-on in
Section 10 as an internal critique of this paper's theory.
Dynamic role allocation, moreover, requires measurement costs. Measuring characteristics,
re-measuring, and maintaining role descriptions cost money; measurement that is too coarse
produces misallocation, and measurement that is too precise produces surveillance. This is
why the refutation condition of Proposition 12 explicitly includes the case in which "the costs
of measurement, mismeasurement, and gaming exceed the gains of dynamic allocation." The
superiority of dynamic role allocation is a theoretical claim, open to comparative testing
against age-fixed allocation (Section 9).
Working Paper | Ageless Management in the AI Era 95
6.3 Organizational Design for Generational Decorrelation — Operationalizing
Proposition 5
As seen in Section 4, WP8 (Kadowaki 2026h) formalized the value of oversight as
independence × detection probability and identified as a necessary condition that the
supervisor's errors be independent of the AI's errors (error decorrelation). WP8 further
organized, from case analyses, the observation that a homogeneous supervisor group trained
contemporaneously with the AI shares blind spots and has difficulty satisfying the
independence condition. Proposition 5 positioned generational heterogeneity as a supply-side
response to this decorrelation condition: the claim that individuals who have experienced
different historical environments, technological generations, and failure cases have a smaller
overlap of judgmental blind spots than within a single generation (Definition 5). Proposition 5
does not, however, make this claim unconditionally. Generational heterogeneity manifests as
an expansion of the group's detection set, and that detection reaches decision-making, only
under three conditions: (i) that auditors have unmediated access to the original output under
verification (the channel condition); (ii) that power gradients are flattened, so that
decorrelated observations reach decision-making without suppression or self-censorship (the
organizational condition); and (iii) that the AI output under audit not be monopolized by a
single foundation model (the model condition) (Section 5). What these conditions demand of
organizational design is developed in turn in the latter part of this section.
Translated into organizational design, this proposition becomes a problem of supervisorgroup
composition. That is, when designing an audit regime for AI output, the design variable
is not only the capacity of the individual auditors but the correlation of errors among
auditors. Isomorphic to the way a financial portfolio is designed not only on the expected
returns of the individual assets but on the correlations among assets, the supervisor group
should be designed not only for individual detection capacity but so that the overlap of misses
is small — this is the design idea of this section. Generational composition is one axis of this
portfolio. However many auditors are lined up who were trained in the same technological
environment and share the same failure cases, the shared blind spots remain. A formation
that combines different technological generations — for example, a generation that
experienced the domain's failures in pre-AI manual work and a generation trained together
with AI — in theory widens the detection set.
Working Paper | Ageless Management in the AI Era 96
Figure 4 The structure of the multigenerational ecosystem and generational decorrelation. The
cognitive assets of the four players are connected through AI-mediated complementarity (Definition 6),
and roles are allocated dynamically on the basis of measured characteristics rather than age
(Proposition 12). The generational heterogeneity of the supervisor group is designed as the
organizational source of error decorrelation (Definition 5). Note: this figure is a schematic of the
theoretical structure; the effectiveness of each connection hangs on the tests of Propositions 5 and 6 and
Hypotheses H2 and H3.
Three constraints on this design idea must nevertheless be made explicit. First, WP8's
capacity constraint binds here as well. Oversight is a consumed resource; adding supervisors
creates problems of coordination costs and the distribution of alarms; there is an upper
bound on the number of objects one supervisor can effectively oversee (span); and the
solvency condition operates on the cognitive resources, time, and money an organization can
devote to oversight. Diversification for decorrelation is not free, and the size of the oversight
portfolio is finite. The monotone prescription "the more diverse auditors, the better" does not
hold under the solvency condition and alarm fatigue. Second, decorrelation is a necessary
condition, not a sufficient one. As WP8 made explicit, the value of oversight arises only when
independence is multiplied by detection probability. Whether a generationally heterogeneous
auditor group in fact has less overlap of misses, and whether each auditor can detect errors
at all, are both untested, open to the refutation condition of Proposition 5 (that, under a
design satisfying the three conditions, the intergenerational correlation of detection errors be
equal to or greater than the within-generation correlation, or that the adoption rate of
decorrelated observations into decision-making not differ significantly from the
homogeneous-group case) and to the tests of Hypotheses H2 and H3 (Section 9). Third, to turn
generational decorrelation into a rule about the age composition of supervisors — quotas
such as "so many members aged 60 or over on the audit committee" — is an operation that
AI-mediated complementarity
(Definition 6)
Super-seniors (aged 60–90s)
Experiential audit, context evaluation
Receives Gf support, supplies Gc
Youth (under 20)
Rapid prototyping
AI compresses the experience barrier
Marginalized participants
Lived-experience problem perception
Cognitive accessibility
Mid-career generation (20s–50s)
Orchestration
Integration and decision-making
Dynamic role allocation (Prop. 12): roles assigned by measured traits, not age
Generational decorrelation (Def. 5): experience of different eras and technology generations → overseer pools with less-overlapping blind spots (supply-side answer to WP8)
Working Paper | Ageless Management in the AI Era 97
proxies the heterogeneity of experience, which is the ground of decorrelation, by the
heterogeneity of age, contrary to the conceptual separation of Section 4.4. The design variable
is strictly the heterogeneity of experience, training environments, and failure cases; age
composition is only its coarse observable proxy.
Independence of the audit channel. In addition to the three constraints above, the
conditional clause of Proposition 5 — that auditors access the original output under
verification without mediation — imposes an independent requirement on portfolio design.
What generational decorrelation supplies is heterogeneity of the error distributions among
auditors; but that heterogeneity manifests as an expansion of the detection set only when
each auditor reads the original output itself with their own experience and their own frame
of judgment. When all auditors share, as their information channel, summarization and prescreening
by the same AI, the heterogeneity is homogenized at the entrance of the audit
process. The context that the AI summary dropped is equally unseen by all auditors.
Prioritization based on confidence pushes the places where the AI errs with confidence — the
typical hallucination — equally outside the attention of all auditors. Then, however
generationally heterogeneous the auditor group, misses re-correlate at the level of the
channel, and the first factor of oversight value — independence × detection probability —
collapses. The financial analogy from the opening of this section applies directly: however
heterogeneous the individual assets, if the price information of all assets passes through a
single source, the diversification effect of the portfolio is lost. The audit regime of the
ecosystem must therefore always include, within the audit portfolio, an unmediated channel
— a slot in which at least some auditors access the original output directly, with neither
summarization nor pre-screening. It need not be the full volume: as Proposition 5 makes
explicit, raw audit of a randomly sampled portion suffices. This requirement, however, stands
in direct tension with the designs Section 8 considers for containing cognitive load — AI
summarization, dialogic auditing, prioritization. The management of this trade-off between
load reduction and independence, and its resolution as dual-track auditing, is treated in
Section 8.1.
Flattening the power gradient. The organizational condition (ii) of Proposition 5 derives
from the distinction between statistical decorrelation and organizational adoption. Even if
generational heterogeneity actually lowers the correlation of error distributions, oversight
value is not realized unless the detection reaches decision-making (Section 5). What blocks
arrival is the organization's power gradient. The auditors of the multigenerational ecosystem
— non-employment super-seniors, youth participating through educational partnerships,
participants via NPOs — are all likely to sit downstream of the power gradient. And
observations that conflict with the judgment of the majority or of superiors — decorrelated
observations are precisely such observations — can vanish before decision-making, through
explicit dismissal (suppression) or voluntary withholding (self-censorship). This is the
reproduction, in the audit context, of the "silence" that BCM (Kadowaki 2026e) formalized as
the failure mode of Belonging. The design requirement comes down to not making the arrival
Working Paper | Ageless Management in the AI Era 98
of observations depend on individual courage. The candidate devices are the
institutionalization of the recording of detections and of escalation channels, a duty to
respond to audit opinions (recording the reasons when dismissing), and an operation that
removes the speaker's age, contract form, and status as variables from the evaluation of an
observation; these are incorporated as the "voice channel" in Table 9 of Section 8. Isomorphic
to the way the formation of the audit portfolio is a device protecting the independence of the
channel, these are devices protecting the independence of adoption.
The model condition. Condition (iii) of Proposition 5 is a design variable on the side of the
audited object, not of the auditors. When the AI outputs under audit all derive from a single
foundation model, that model's systematic blind spots — errors rooted in the training
distribution and the architecture and appearing correlated across all outputs — become a
common factor dominating all audit tasks, and human generational heterogeneity cannot
override them (Section 5). The organizational design requirement is to use in parallel, at least
for outputs with grave consequences, cross-validation by models of different architectures
and developers. The method for incorporating this model factor into the measurement of
detection-overlap rates is treated in Section 9.
The bulwark against cognitive hold-up. Flattening the power gradient has a reverse side.
A regime in which observations are not suppressed simultaneously requires a bulwark
against the incentive to oversupply observations. Experiential auditing has no established
market price (Section 8.2), and under the information asymmetry in which the
commissioning side can hardly verify directly the validity of a flagged risk, auditors — above
all non-employment seniors whose contract continuation depends on the impression of the
"usefulness" of their observations — face a structural incentive to perpetuate their own status
and contracts by continuing to flag fictitious or inflated risks. This is a hold-up inherent in the
agency relation of auditing, and it requires no malice — excess caution is, in one's own
introspection, indistinguishable from conscientiousness. Three bulwarks can be specified.
First, exclude solo auditing by a single senior: a solitary auditor becomes, in effect, the sole
evaluator of their own observations, and the information asymmetry is maximized. Second,
make independent double-checking by multiple heterogeneous seniors the basic form: the
multiple placement of heterogeneous auditors that Proposition 5 required for the
decorrelation of misses performs here a second function as the decorrelation of over-flagging
— the likelihood that independent, heterogeneous auditors flag the same fictitious risk is low
(an application of Proposition 5). Third, periodic blind evaluation of audit validity: an
operation that blindly mixes in known errors and unproblematic outputs and periodically
measures detection rates and over-flagging rates is the operational version of the quantity
that Hypothesis H3 measures as its dependent variable (iv), the over-rejection rate (Section 9),
and it re-anchors the evaluation of auditors on the validity, not the quantity, of observations.
An audit regime lacking these bulwarks can, in the name of decorrelation, multiply
verification costs without limit. This cost ledger connects to the net-benefit framework of
Section 6.5.
Working Paper | Ageless Management in the AI Era 99
It should also be made explicit that generation is not the only axis of the portfolio. The
heterogeneities that can produce error decorrelation include domain heterogeneity (failure
cases from different industries), functional heterogeneity (legal, technical, front-line),
training-environment heterogeneity (pre-AI manual training versus AI-native training), and
heterogeneity of lived experience (Section 6.4). Generational heterogeneity can be understood
as bundling, along the time axis, the heterogeneity of training environments and of failure
cases among these. The design of the oversight portfolio is therefore not an optimization
problem of generational composition but a multi-axis composition problem that minimizes
the overlap of blind spots, in which generation merely provides one easily observable axis.
Also, though a macro-level finding, Zélity's (2023) finding of an optimal level (a hump shape)
in the relation between age diversity and productivity is suggestive for oversight portfolios as
well. Since the benefits of heterogeneity are eaten by coordination costs and capacity
constraints, the size of the portfolio and the degree of diversification have an optimum, and
the design goal is "optimization," not "maximization." The location of this optimum cannot be
derived from theory and is left to organization-specific measurement (Section 9).
The existing empirical evidence on the value of diversity in error detection should be
confirmed honestly. Direct peer-reviewed evidence that age diversity improves error
detection or audit performance is not found within the scope of this paper's search. What
exists is indirect evidence. Sommers (2006), in a mock-jury experiment, reported that racially
diverse groups exchanged a wider range of information than homogeneous groups and that
factual errors by the majority members themselves decreased — a finding showing that
diversity can reduce errors not only through the pathway of "bringing in different
information" but also through "making the majority's information processing more careful";
but this is an experiment on racial diversity, and replication with age is not confirmed. Also,
Börsch-Supan & Weiss (2016), from unique data connecting production-process errors on
assembly lines with worker attributes, reported that no evidence was obtained supporting
the conventional view that productivity declines with age; and by the summary of the
National Academies (2022), older workers make slightly more minor errors but grave errors
are rare, and experience offset the productivity decline. But this is an effect of experience (or
of age), not an effect of diversity, and the two must not be conflated. Proposition 5 is
consistent with these pieces of indirect evidence, but it is an untested theoretical proposition
not supported by them.
6.4 Modes of Connection with NPOs and Educational Institutions
The multigenerational ecosystem of Definition 7 is not limited to the firm's own boundary of
employment. Youth participation typically passes through PBL-type educational partnerships
and contest-type schemes; the participation of socially marginalized groups through NPOs
and intermediary support organizations; and the participation of super-seniors through
outsourcing and advisory contracts. The institutional conditions of these modes of connection
(labor law, social security, freelance protection) are treated in Section 7. This section confirms
Working Paper | Ageless Management in the AI Era 100
the empirical evidence on the effectiveness of connection. To state the conclusion first: the
evidence converges on a single point — connection is possible, but not unconditional.
First, connection with educational institutions. PBL-type partnership is a mode that
connects youths' ideation and prototyping capacity to the ecosystem within the frame of
educational purpose, and it is also institutionally required as a form of participation not
based on an employment contract (Section 7; Proposition 10). But the finding of Bastani et al.
(2025), seen in Section 4, gives a direct warning here: use of generative AI without guardrails
greatly raised assisted performance while harming independent learning outcomes. A PBL
design with primacy of educational purpose is therefore not only incompatible with a design
whose main purpose is the firm's acquisition of deliverables; it must also include the design
of AI use (guardrails; secured unassisted practice) as a requirement on the educational side.
Next, connection with NPOs and intermediary support organizations. As organizational
forms bearing the work integration of people who face barriers to employment, peerreviewed
research has accumulated, centered on work-integration social enterprises (WISEs),
and there exist syntheses of the knowledge on work integration of people with brain injury,
mental illness, and intellectual disability (Kirsh et al. 2009) as well as reviews of evaluation
frameworks. This field, however, is dominated by case studies and qualitative research, and
controlled effectiveness studies are few. From the standpoint of ecosystem design, what can
be expected of NPO partnership extends to the existence of participation pathways and the
accumulation of operational know-how — not to a quantitative guarantee of outcomes.
On assistive technology (AT), held to be the technical basis of connection, a systematic
review exists. Marinaci et al. (2023) systematically reviewed 41 publications from 2017
onward and concluded that assistive technologies show effectiveness for overcoming
accessibility barriers, improving job performance, independence, and expanding career
opportunities. The review itself, however, states as limitations its geographic skew and the
lack of research on the emotional and sociocultural dimensions of AT use in the workplace.
As a further important qualification, the evidence that AT and AI-based cognitive-accessibility
technologies (voice UIs, summarization, read-aloud) support job performance differs in level
from the evidence that they improve labor-market outcomes — job acquisition, retention, and
wages. Peer-reviewed evidence directly measuring the latter is not established within the
scope of this paper's search. The cautionary view — "too much promise, yet too little
substance" (Smith & Smith 2021) — remains valid.
And the currently most precise empirical study testing the relation between AI and the
employment of people with disabilities on an international panel does not permit optimism.
Omri et al. (2025), using data from 27 high-technology advanced countries over 2006–2022,
estimated with a moderated mediation analysis the effect of AI adoption (proxied by the
number of industrial robots installed) on unemployment among people with disabilities. The
direct effect was an increase in unemployment, and the authors' initial hypothesis (that AI
reduces unemployment) was rejected. The only significant unemployment-reducing pathway
Working Paper | Ageless Management in the AI Era 101
was the indirect effect through higher education (coefficient −0.0322, 95%CI [−0.0550,
−0.0151]); the indirect effect via basic and secondary education was nonsignificant. Moreover,
a counterintuitive moderation was estimated whereby the higher the quality of governance,
the more the unemployment-increasing effect of AI is amplified; the simple expectation that
"good institutions cancel out AI's harms" is not supported either. Even allowing for the study's
limitations (the proxy variable is pre-generative-AI robot adoption; causal identification is
weak), the implication to draw is clear. AI does not advance inclusion unconditionally.
Inclusion is a possibility that opens only when educational investment, accessible design, and
the conditional design of modes of connection are all in place — which is the very reason
Proposition 11 states expressly that "the conversion is not automatic."
These pieces of evidence bring into view the design elements demanded in common of the
modes of connection with NPOs and educational institutions. First, the translation and
correction function of the intermediary organization. NPOs and educational institutions bear,
between participants and firms, functions that individuals cannot bear themselves: securing
the cognitive accessibility of tasks, negotiating conditions, supervising educational purpose.
Designing the mode of connection is, in substance, designing this intermediary function.
Second, the explicit statement of conditions. Since inclusion through AI depends on the
conditions of education, design, and institutions, the agreement of connection must state
explicitly the design of AI use (guardrails; unassisted opportunities), compensation and
attribution, and the ceiling load of participation (Table 9 in Section 8). Third, the calibration
of outcome expectations. Measuring the initial outcomes of connection by employment
outcomes exceeds what the current state of the evidence confirms (the level of support for job
performance) and readily invites withdrawal through disappointment. Starting measurement
at the level of job performance, continuation of participation, and brain-capital indicators is
what is consistent with the evidence (Section 9).
IP defense (appropriability). For connection beyond the boundary of employment, the
price from the firm's standpoint must also be stated. As noted in Section 1, the theoretical
backbone of this paper lies in the lineage of the resource-based view (Barney 1991) and
dynamic capabilities (Teece, Pisano & Shuen 1997), and experiential audit capacity can be a
source of competitive advantage because it is rare and hard to imitate — the immobility of
the resource. Yet the ecosystem of Definition 7 works precisely in the direction of weakening
this immobility. Under the open modes of connection — outsourcing, advisory, PBL, NPO
partnership — the tacit knowledge transferred to participants in the course of auditing — the
business's criteria of judgment, interpretations of failure cases, organization-specific context
— can leak beyond the boundary of employment through participants' departure or
simultaneous participation in other firms (knowledge spillover). An organization that
depends on outside participants for the source of its competitive advantage structurally faces
the question of whether it can appropriate the returns from that source (appropriability).
Contractual instruments such as NDAs are necessary, but the leakage of tacit knowledge
cannot be fully captured by contract. This paper therefore specifies, in addition to contract,
Working Paper | Ageless Management in the AI Era 102
two architectural controls as complementary governance. First, the modularization and
encapsulation of interpretive context: partition the contextual information necessary for
auditing into engagement-level modules, and disclose to each participant only the scope
necessary for their assigned component — the context-partition protocol of Section 6.2 can
share the same partition with this appropriability control. Second, federated management of
access rights to the RAG layer: compartmentalize the read and write permissions to the
accumulation layer of interpretive context (the RAG layer) treated in Section 8.1 by
participant and by engagement, so that no single participant can reach the full picture of the
organization's interpretive context. These controls are an operation that re-seats
organization-specific contextual integration — the orchestration that connects the
participants' knowledge, rather than the knowledge of the individual participants — in the
seat of inimitability, a design for reconciling open participation with the defense of
appropriability. Excessive compartmentalization, however, thins the very supply of context
necessary for auditing and lowers detection probability. The adjustment of this trade-off
between the defense of appropriation and the effectiveness of auditing is likewise entered in
the net-benefit ledger of Section 6.5.
Two related points should be added. First, bias in AI hiring tools. The systematic scoping
review of El Morr et al. (2024), while indicating that AI as assistive technology can enhance
the lived experience of people with disabilities, pointed out that all 34 of the reviewed AI
model-building studies neither measured nor addressed bias, and that AI hiring tools
perpetuate discrimination against people with disabilities. A selection AI placed at the
entrance of the ecosystem can close off the mode of connection itself. Second, the relation
between inclusion and firm performance. A corporate survey exists correlating leading firms
in disability inclusion with high performance (Accenture 2018), but it is grey literature
without peer review, a correlation on a self-selected sample, and it cannot exclude reverse
causality (high-performing firms can afford to invest in inclusion). The rollout of
neurodiversity employment programs (SAP, Microsoft, and others) is a fact (Krzeminska et al.
2019), but peer-reviewed independent effectiveness studies are limited. This paper does not
use these as grounds for the proposition that "inclusion pays."
Evidence-grade note: Accenture (2018) is a corporate survey (grey literature), and its figures are not cited
in the text. The WISE literature is centered on case and qualitative research, and individual effect sizes
are not cited in this paper. Omri et al. (2025) is a peer-reviewed empirical study, but it is a panel
mediation analysis with weak causal identification, and its data predate the diffusion of generative AI
(through 2022). The statements of this section on the inclusion effects of generative AI stand inside these
qualifications.
6.5 The Connection to FVT — As a Conditional Pathway
The original draft argued that organizations with diversity of thought are more likely to
produce discontinuous innovation, and that this leads to firm value in the sense of Future
Working Paper | Ageless Management in the AI Era 103
Value Theory (FVT; Kadowaki 2026a). This section revises that claim into a conditional form,
in the light of the empirical baseline.
First, the baseline does not support optimism. The conclusions of the meta-analyses on the
relation between age diversity and team outcomes agree near zero. Joshi & Roh (2009), in a
meta-analysis of 39 studies and 8,757 teams, estimated the direct effect of age diversity on
performance at r = −.06 (k = 21, 95%CI [−.09, −.04]). Schneid et al. (2016), in a meta-analysis of
74 studies, reported that the overall relation between age diversity and team outcomes was
nonsignificant and that the only significant relation was increased turnover (r = .11),
themselves positioning this as a refutation of arguments emphasizing age diversity. In the
most recent and largest registered-report meta-analysis, Wallrich et al. (2024) (615 reports,
2,638 effect sizes), the overall effect of demographic diversity is r = .014, effectively zero, and
age diversity alone is nonsignificant. The unconditional claim that "being multigenerational
raises value" is incompatible with the totality of the existing evidence. This paper does not
make that claim.
On the other hand, the same body of evidence also identifies the conditions under which
effects appear. Backes-Gellner & Veen (2013) showed, from large-scale German linked data,
that age diversity has a positive effect on firm productivity when, and only when, the firm is
engaged in creative rather than routine tasks. In the moderator analysis of Wallrich et al.
(2024) as well, the diversity–performance relation is more positive for tasks that are high in
complexity and whose outcomes depend on creative divergence. The climate conditions of
Wegge et al. (2012) were stated in Section 6.2. At the macro level, moreover, Zélity (2023)
shows that the relation between age diversity and GDP per capita is hump-shaped, that is,
that an optimal level exists (a country-level aggregate relation, not directly transferable to
intra-firm formation). Taken together, the claim that can be written is this: age diversity can
contribute to outcomes when the task is complex and creative, when there is a climate of low
age discrimination in which diversity is positively appraised, and when age-inclusive
management accompanies it. And diversity is not monotonically good; it has an optimum.
This paper's theoretical wager is to add AI-mediated complementarity to this list of
conditions. Proposition 6 claims that, taking as given the known conditions of task complexity
and inclusive climate, AI-mediated complementarity (Definition 6) is an additional moderator
of the relation between age and experience diversity and organizational outcomes, and that
in its presence the relation moves in a more positive direction. What Proposition 6 claims is
not the absolute level of the relation (that it will necessarily be positive) but a comparative
static: the difference in the relation with and without AI mediation. The near-zero baseline of
the existing meta-analyses was measured almost entirely in environments lacking AI
mediation, and on this paper's reading it is compatible with Proposition 6. Against the
standard explanation that the benefits of diversity (the complementarity of perspectives and
knowledge) are eaten by coordination costs (communication, friction), if AI lowers the costs
of the Gf-type components and of translation and brokerage, the balance of benefits and costs
can move — this is the theoretical ground of Proposition 6. It is, however, untested, and it will
Working Paper | Ageless Management in the AI Era 104
be adjudicated only by comparisons that manipulate the presence of AI mediation (the
refutation condition of Proposition 6; Hypothesis H2). As for the diversity theorem of Hong &
Page (2004), often invoked in this context, its qualifications must be stated. The theorem is a
result of a computational model under specific conditions, and its mathematical generality
has been criticized (Thompson 2014). Moreover, the "diversity" of the theorem is cognitive
and functional diversity, and the empirical evidence that age diversity proxies for it is weak.
The pathway age → heterogeneity of experience and technological environment →
(conditionally) cognitive complementarity is a hypothesis in this paper as well.
On how outcomes are to be measured, the framework of net benefit and transaction costs is
made explicit here. Multigenerational auditing and verification are not free of charge. The
processing and coordination of heterogeneous observations delay decision-making;
independent double-checking (Section 6.3) increases verification effort; and the portfolio
formation for decorrelation itself carries coordination costs. These are the transaction costs
of multigenerational formation. On the other hand, the principal benefits of auditing — the
avoidance of collapse through excessive risk, the reduction of rework — appear only weakly
in the immediate evaluation of outputs. If this asymmetric manifestation of benefits and costs
is left unaddressed, the evaluation of multigenerational formation can be manipulated
arbitrarily through the choice of measure. It is to foreclose this manipulation that Proposition
6 defines outcomes as "net benefit — not idea quality alone, but including the avoidance of
excessive risk and the reduction of rework, and net of the cost of the decision time required
for verification." That is, the claim of the superiority of multigenerational formation is made
only on the ledger of risk-adjusted, verification-cost-deducted net benefit, not on the single
indicator of idea quality. That Hypothesis H2 requires recording time to decision and
verification effort as cost variables is the measurement-side counterpart of this definition
(Section 9).
The ledger of transaction costs also demands the consideration of pathways that lower
costs. If the infrastructure this section and Section 8 specify — the maintenance of
characteristic measurement and role descriptions, audit assignment and context partition,
the sampling management of dual-track auditing, the compartmentalization of the RAG layer
and the preservation of its diversity, the standardization of contracts and compensation —
were all built in-house by each individual firm, the fixed cost would not be small.
Organizations of a scale at which net benefit exceeds the fixed cost are limited, and the
applicability of this model shrinks to that range — this limit of external validity is confronted
head-on in Section 10.5. On the other hand, much of this infrastructure — measurement
tasks, audit-assignment algorithms, standard contract templates, sampling designs — is low in
firm-specificity and lends itself to standardization. If, therefore, the measurement and
governance infrastructure were provided to multiple organizations as common
infrastructure (standardized SaaS-type services), the fixed costs of development and
maintenance would be shared among the user organizations, and the marginal cost per firm
could fall. If this pathway materializes, the lower bound of the organizational scale at which
Working Paper | Ageless Management in the AI Era 105
this model is applicable comes down. This, however, is the presentation of a possibility, not an
assertion. Whether such a service market actually forms, and how far standardization is
compatible with each organization's contextual specificity, are both empirical questions not
derivable from this paper's theory, and they should be read paired with the statement of
limitations in Section 10.5.
Empirical findings that should be written separately from team outcomes concern
intergenerational knowledge transfer. That knowledge transfer in age-mixed dyads raises the
motivation and organizational commitment not only of the receiver but of the sender
(Burmeister et al. 2020), that reverse mentoring can yield skill development for both younger
mentors and older learners (Kaše et al. 2019), and that intergenerational learning is
bidirectional (Gerpott et al. 2017) all have peer-reviewed evidence. These are pathways — set
apart from performance effects — to retention, skill formation, and the formation of the
brain-capital stock (K), and they constitute adjacent evidence for Proposition 8 (bidirectional
capital formation). Not conflating "the performance effects of diversity" and "the effects on
knowledge transfer and retention" into a single virtue is the descriptive discipline of this
area.
Finally, the scope of the connection to FVT is demarcated. FVT (Kadowaki 2026a) is a
framework that inverts the origin of firm value from past accumulation to the creation of
future value, and the organizational conditions that raise the probability of the occurrence of
discontinuous innovation are its central concern. Two pathways can be considered by which
the design argument of this section connects to FVT. The first pathway is that
multigenerational formations satisfying the conditions raise the production of ideas that
combine novelty and feasibility; this is tested directly as Hypothesis H2. The second pathway
is that social impacts — the lightening of the social-security burden through the continued
participation of super-seniors, the lifetime-income effects of early youth participation — are
reflected in firm value via ESG evaluation and the cost of capital. A proposal exists to expand
investment in late-life brain capital as an object of ESG investment (Dawson et al. 2022), but it
is a proposal paper, and to this paper's knowledge there is no empirical study of this pathway.
The second pathway is stated in this paper only as a pathway hypothesis, and no value claim
premised on its holding is made. The original draft's expression that the multigenerational
ecosystem is a "robust infrastructure" of Future Value is replaced, in this paper's framework,
by "a conditionally testable design hypothesis."
The design argument of this section can be summarized as follows. (i) The role placement
of Table 5 is a default design, not fixed roles; its operation is dynamic role allocation, which
includes the pair of indirect audit feedback and reverse mentoring, the cognitive
apprenticeship loop with micro-ownership (Proposition 8) and its responsibility layering
(HOTL confined to coaching, with attributed responsibility on the junior), multi-role
concurrency (identity buffer — the engagement protection of Proposition 7), and the context
partition for hybrid domains (the operationalization of boundary condition (ii) of Proposition
4). (ii) The supervisor group is formed as a portfolio with the correlation of errors as a design
Working Paper | Ageless Management in the AI Era 106
variable, but that formation is subject to WP8's capacity constraints, carries no guarantee of
detection probability, and requires the three conditions of Proposition 5 — an unmediated
channel (a randomly sampled raw-audit slot), the flattening of the power gradient, and
diversity on the AI side of the audited objects — together with the bulwarks against cognitive
hold-up (independent double-checking; periodic blind evaluation of audit validity). (iii)
Connection beyond the boundary of employment is possible but not unconditional, requiring,
in addition to the conditional design of education, accessibility, intermediary functions, and
compensation, the defense of appropriability (the modularization of interpretive context and
the compartmentalization of RAG-layer access rights). (iv) The outcome effects of
multigenerational formation are conditional; AI-mediated complementarity is this paper's
theoretical addition to that list of conditions, and the claim of its superiority is made on net
benefit after deducting the transaction costs of verification — with the pathway of lowering
that cost through common infrastructure left open as a possibility (Section 10.5). And all of
these designs are sustainable only on the premise that participants' brain capital is protected.
That the institutional conditions of participation are treated in the next section (Section 7),
and the design criteria of protection in the one after (Section 8), follows from this order.
7. Institutional Hurdles: An International Comparison
The theory and design arguments of the preceding sections have been, so to speak, matters
internal to the organization. The implementability of Ageless Management (Definition 1),
however, is strongly conditioned by the legal institutions outside the organization. Even if an
organization seeks to remove calendar age from the allocation of roles, if the pension system
imposes a de facto marginal tax rate on work beyond a certain age, and if the choice of
contractual form determines the presence or absence of protection, dynamic role allocation
(Proposition 12) runs into an institutional wall. For three settings — employment-based work
at older ages (Section 7.1), non-employment participation at older ages (Section 7.2), and
youth participation (Section 7.3) — this section compares the institutions of Japan, the United
States, Germany, the EU, Singapore, South Korea, and the ILO (Section 7.4, Table 6), and
confirms the institutional foundations of Proposition 9 (the protection vacuum of nonemployment
forms) and Proposition 10 (the symmetry of Brain Safety) (Section 7.5).
Two caveats at the outset. First, this section is a description and comparison of institutions,
not legal advice. The application of any particular institution is determined by the facts of the
case and the latest statutes and circulars, and practical decisions require consultation with
professionals. Second, the statute names, article numbers, effective dates, and monetary
amounts in this section are limited to matters checked against primary sources — the e-Gov
statute database, the Ministry of Health, Labour and Welfare, the Japan Pension Service, the
U.S. Social Security Administration (SSA), the German Pension Insurance (DRV), EUR-Lex, and
the ILO — or reliable professional commentary (evidence grade: primary legal and
administrative sources). Details that could not be so verified are not stated.
Working Paper | Ageless Management in the AI Era 107
7.1 Employment-Based Work at Older Ages — Mandatory Retirement,
Employment Security, and Pension Work Disincentives
Japan — the Two-Tier Structure of the Act on Stabilization of Employment of Elderly
Persons and the Emergence of Non-Employment Options
The basic statute governing employment-based work at older ages in Japan is the Act on
Stabilization of Employment of Elderly Persons, etc. (Act No. 68 of 1971; commonly, the Act on
Stabilization of Employment of Elderly Persons). Article 8 of the Act provides that where a
mandatory retirement age (teinen) is set, it may not be below 60 (with an exception for work
in which elderly persons have difficulty engaging). Article 9 obliges employers that set a
mandatory retirement age below 65 to take one of the following employment-security
measures for elderly persons: (i) raising the retirement age, (ii) introducing a continuedemployment
system, or (iii) abolishing the mandatory retirement age. This is a legal
obligation.
By contrast, Article 10-2, newly established by the amendment under Act No. 14 of 2020,
prescribed the securing of work opportunities from age 65 to 70 as an obligation to make
efforts (effective April 1, 2021). The difference in the nature of the obligations — the measures
up to 65 (Article 9) being a legal obligation, the measures up to 70 (Article 10-2) an effort
obligation — is important, and the age-70 measures are not a mandate of retirement at 70.
The Ministry of Health, Labour and Welfare, too, states explicitly that the amendment does
not obligate raising the mandatory retirement age to 70.
From this paper's standpoint, what most deserves attention is the composition of the menu
of measures for securing work up to age 70. There are five options: in addition to (i) raising
the retirement age to 70, (ii) abolishing the mandatory retirement system, and (iii)
introducing a continued-employment system up to 70 (reemployment or extended
employment), they include (iv) a scheme for continuously concluding outsourcing contracts
up to 70, and (v) a scheme under which the person can continuously engage up to 70 in the
employer's social-contribution projects (or the social-contribution projects of entities the
employer commissions or invests in). Options (iv) and (v) are called the entrepreneurshipsupport
measures (Article 10-2), and options for securing work outside employment were
thereby given explicit statutory standing. That is, Japan's older-age employment legislation
has, for the 65–70 phase, officially recognized participation forms outside the employment
boundary as formal policy instruments. This points in the same direction as the participation
structure spanning employment, outsourced engagement, and social-contribution projects
envisaged by Definition 7 (multigenerational ecosystem). Introducing the entrepreneurshipsupport
measures, however, requires a procedure of preparing a plan and obtaining the
consent of a majority union or equivalent, and, as discussed below (Section 7.2), the measures
carry a protection vacuum precisely because they are non-employment forms.
Working Paper | Ageless Management in the AI Era 108
Japan — the In-Work Old-Age Pension Offset and the April 2026 Revision
The other institutional variable of employment-based work at older ages is Japan's in-work
old-age pension offset (zaishoku rลrei nenkin) — the in-work suspension of the old-age
employees' pension under the Employees' Pension Insurance Act. The mechanism is as
follows. When the sum of the basic monthly amount of the old-age employees' pension and
the total-remuneration-equivalent monthly amount exceeds the suspension-adjustment
amount (the threshold), one half of the excess is suspended (suspended amount = (basic
monthly amount + total-remuneration-equivalent monthly amount − threshold) ÷ 2). This
structure, in which increased earnings from work bring a reduction of the pension, has long
been identified as an institutional factor generating work adjustment among older persons —
the perception that "working means losing."
This institution changes substantially in April 2026. Under the Act Partially Amending the
National Pension Act, etc. for Strengthening the Functions of the Pension System in Light of
Social and Economic Changes (Act No. 74 of 2025; enacted June 2025), the threshold of the inwork
old-age pension offset is raised, effective April 1, 2026. The monetary relationships
require precision. The pre-amendment threshold was at the 500,000-yen level (in actual
terms, 510,000 yen for FY2025). The amending Act raised this to 620,000 yen, but this 620,000
yen is the statutory value based on wage levels at the time of the Act's enactment in 2025.
Because the threshold is automatically revised each fiscal year in line with wage movements
(wage indexation), the threshold actually applied in FY2026, the fiscal year of entry into force,
became 650,000 yen. That is, "from 510,000 yen to 620,000 yen" is the statutory increase (at
2025 prices), while "650,000 yen" is the actual amount applied in FY2026 after wage
indexation; the two are not in contradiction. The Japan Pension Service's dedicated page
likewise states the figures as 510,000 yen per month before the amendment (FY2025) and
650,000 yen per month after (FY2026).
One further point requires precision. Persons aged 70 or over who work at establishments
covered by employees' pension insurance have lost insured status under employees' pension
insurance upon reaching 70 and therefore bear no premiums. The suspension under the inwork
old-age pension offset, however, continues to apply to employed persons aged 70 or
over under the same formula (for employed persons aged 70 or over, it is calculated using an
amount corresponding to the standard monthly remuneration). The understanding that "past
70 one falls outside the in-work old-age pension offset" is mistaken, and the end of premium
liability must not be confused with the continued application of the suspension. For the
employment-based participation of the super-seniors (aged 60 to their 90s) envisaged by
Ageless Management, this suspension is an institutional variable that continues beyond age
70.
United States — Abolition of the Earnings Test at and after FRA, and the ADEA
The United States is a representative example of a contrasting institutional choice. The
retirement earnings test on Social Security old-age benefits was a mechanism that withheld
Working Paper | Ageless Management in the AI Era 109
benefits when earnings from work exceeded exempt amounts, but the Senior Citizens'
Freedom to Work Act of 2000 (Public Law 106-182, signed April 7, 2000) abolished the
earnings test at and after full retirement age (FRA). From FRA onward, old-age benefits are
not reduced no matter how much one earns from work.
One must not, however, simplify this to "the United States abolished the earnings test." For
beneficiaries before FRA the earnings test remains in force. In years before the year of FRA
attainment, $1 of benefits is withheld for every $2 of earnings above the lower exempt
amount, and in the year of FRA attainment (through the month before the attainment
month), $1 is withheld for every $3 of earnings above the upper exempt amount. The SSA's
official 2026 exempt amounts are $24,480 per year (lower) and $65,160 per year (upper).
Moreover, because withheld benefits are effectively recovered through benefit recomputation
after FRA attainment, writing "forfeited" is also inaccurate. The institution's implication lies
in the point that it is a design in which the work-disincentive effect disappears at the age
boundary of FRA.
On the employment-discrimination side, the Age Discrimination in Employment Act of 1967
(ADEA) prohibits age discrimination in employment (hiring, discharge, compensation,
promotion, and the like) against workers aged 40 and over. The upper age limit on protected
status that originally existed was removed by the 1986 amendment, whereby mandatory
retirement itself became unlawful for most occupations (with limited exceptions such as
pilots). The United States is the representative country that adopts "no mandatory retirement"
as a matter of legal institution, and it can be called the jurisdiction that has carried the
direction of Definition 1 — removing calendar age as a criterion for exit from roles — furthest
at the level of employment law.
Germany — Complete Abolition of the Earnings Limit on Early Pensions
Germany is the example that has most thoroughly removed work disincentives on the
pension side. The earnings limit (Hinzuverdienstgrenze) applicable while drawing an early
old-age pension was abolished on January 1, 2023 by the Eighth Act Amending Book IV of the
Social Code (8. SGB IV-ÄndG). The official FAQ of the German Pension Insurance (DRV) states
explicitly that this limit has been abolished entirely (ganz entfallen). Since then, even those
drawing an early pension before the statutory pension age can work without limit while
receiving the full pension, regardless of the amount earned. It is a permanent measure and
applies to all recipients irrespective of when the pension began. Until the abolition, through
2022, an annual limit of 46,060 euros per year (a level raised under COVID special rules)
applied. This abolition, however, concerns early old-age pensions: the earnings limits on
disability pensions (Erwerbsminderungsrente) have not been abolished but have shifted to
dynamic limits linked to wage trends. One must not generalize to "Germany has abolished all
pension earnings limits."
At the EU level, the Employment Equality Directive (Council Directive 2000/78/EC, adopted
November 27, 2000) prohibits direct and indirect discrimination in employment and
Working Paper | Ageless Management in the AI Era 110
occupation on grounds of religion or belief, disability, age, and sexual orientation. Article 6 of
the Directive, however, provides that differences of treatment on grounds of age can be
justified where there is a legitimate aim of employment policy, the labor market, or
vocational training and the means are appropriate and necessary. This is a special exception
granted to age alone among the four grounds of discrimination, and it is the basis provision
for the case law of the Court of Justice of the EU on the permissibility of member states'
mandatory retirement systems. Age is given special treatment, even inside antidiscrimination
law, as a "justifiable distinction" — this asymmetry itself indicates how
institutionally deep-seated the variable of calendar age is.
Singapore and South Korea — Extending Working Lives through Mandates
Asia's leading jurisdictions on older-age work take the route of extending working lives by
strengthening mandates. Singapore's Retirement and Re-employment Act adopts a two-tier
structure of a statutory retirement age (below which forced retirement is prohibited) and,
beyond it, an age up to which re-employment must be offered. From July 1, 2022 the
retirement age has been 63 and the re-employment ceiling 68, and from July 1, 2026 they are
raised to a retirement age of 64 and a re-employment age of 69 (the government has stated a
policy of raising them to 65 and 70 by 2030). In imposing an obligation to offer re-employment
rather than an obligation to retain employment, the structure is close to Japan's continuedemployment
system.
South Korea, through the 2013 amendment of the Act on Prohibition of Age Discrimination
in Employment and Elderly Employment Promotion (the retirement-age extension act,
enacted June 2013), made it mandatory to set the retirement age at 60 or above. It was phased
in from 2016 for workplaces with 300 or more employees and public institutions, and from
2017 for those with fewer than 300. The Act also prescribes the prohibition of employment
discrimination on grounds of age. In recent years, a further extension of the statutory
retirement age (to 65) has been under discussion between labor and management and in the
National Assembly.
These employment-based institutions form a three-stage spectrum in the design of in-work
pensions. Japan maintains partial suspension while raising the threshold (2026), the United
States abolished the test at and after FRA (2000), and Germany abolished it entirely, including
for early claimants (2023). The direction is in every case a shrinking of work disincentives,
but the end points differ. This point is taken up again in Section 7.5.
7.2 Non-Employment Participation at Older Ages — the Protection Vacuum and
Its Partial Filling
The participation forms on which Ageless Management depends are not limited to
employment. Definition 7 (multigenerational ecosystem) explicitly includes outsourced
engagement, advisory roles, PBL-based educational partnerships, and NPO partnerships, and
the design arguments of Section 6 conceived many of the audit and advisory roles of super-
Working Paper | Ageless Management in the AI Era 111
seniors in non-employment form. As seen in Section 7.1, Japan's Act on Stabilization of
Employment of Elderly Persons has itself formalized a non-employment option in the
entrepreneurship-support measures. Here, however, lies a structural problem. Most of the
protections of labor law and social security law are designed with the employment
relationship (the labor contract) as their unit. Workers in outsourced, advisory, or gig-type
arrangements, because they are not based on labor contracts, are in principle not reached by
the protections of labor law, beginning with the automatic application of the Labor Standards
Act and workers' accident compensation insurance. The entrepreneurship-support measures
extended work opportunities into non-employment, but protection does not automatically
follow. Proposition 9 (the protection vacuum of non-employment forms) refers to this
structure: current institutions are designed with the employment relationship as the primary
unit of protection and restraint, non-employment participation falls into a vacuum of
institutional protection, and Ageless Management that leaves the vacuum unattended can
degenerate into exploitation.
Overlaid on this vacuum, beyond the negative problem of absent protection, is the positive
problem of an asymmetry of institutional incentives. As confirmed in Section 7.1, the
suspension under the in-work old-age pension offset does not end at 70. Even after a person,
upon reaching 70, loses insured status under employees' pension insurance and ceases to
bear premiums, so long as the person works in employment at an establishment covered by
employees' pension insurance, the suspension continues to apply under the same formula, as
an "employed person aged 70 or over," based on the basic monthly amount and the totalremuneration-
equivalent monthly amount (calculated using an amount corresponding to the
standard monthly remuneration). If, on the other hand, the same person shifts to nonemployment
work under an outsourcing contract — the entrepreneurship-support measures
of Article 10-2 of the Act on Stabilization of Employment of Elderly Persons — the person does
not qualify as an employee under employees' pension insurance and is therefore not subject
to the suspension of the in-work old-age pension offset. That is, under current institutions
there exists a structure of institutional arbitrage in which, even where the same individual
supplies the same kind of services, whether the pension is suspended or not divides
according to whether the contractual form is employment or outsourcing. This asymmetry
financially steers organizations and individuals who would supply Gc-type roles at
remuneration levels above the threshold in employment form toward moving from the
contractual form with thicker protection to the contractual form with thinner protection.
With the formalization of opportunity (the entrepreneurship-support measures) and the
pension arbitrage pointing in the same direction, the shift to non-employment is not an
institutionally neutral choice but an induced one. What makes the protection vacuum of
Proposition 9 grave is that the destination of this inducement is precisely the place where
protection is thinnest.
The reverse side of this arbitrage is the risk of employee misclassification — so-called
disguised subcontracting. Even if the contract is titled an outsourcing agreement, whether a
Working Paper | Ageless Management in the AI Era 112
person qualifies as a worker is judged by substance. The basic framework of administrative
interpretation, the Report of the Labor Standards Act Study Group, "On the Criteria for
Determining 'Worker' Status under the Labor Standards Act" (December 19, 1985), sets out a
framework of holistic judgment centered on the subordination test — the presence or
absence of freedom to accept or refuse work requests and instructions to engage in tasks, the
presence or absence of direction and supervision in the performance of work, the presence
or absence of constraints on place and hours of work, and the substitutability of the labor
supplied — and the remuneration's character as compensation for labor, with the presence or
absence of entrepreneurial character, the degree of exclusivity, and the like as reinforcing
elements; the judgment of the relationship of use and subordination in the internship
circular examined in Section 7.3 (Circular Kihatsu No. 636) belongs to the same lineage.
Accordingly, if a super-senior on an outsourcing contract is obligated to engage in constant
audit work, has working hours and place managed, is not allowed to accept or refuse
individual requests, and has the performance of work placed under the organization's
direction and command, worker status can be found as a matter of substance regardless of
the name of the contractual form. The consequences do not stop at employment
responsibilities — application of the Labor Standards Act, the Minimum Wage Act, and
workers' accident compensation insurance — and the retroactive incurrence of applicable
social insurance premiums: through qualification as an "employed person aged 70 or over,"
the very pension treatment the preceding arbitrage presupposed can be overturned.
Organizations that design super-seniors' audit and advisory roles in non-employment form
therefore need to build into the design a legal safeguard that aligns the substance of the
contract with the requirements of the non-employment form — this paper calls it the protocol
for the exclusion of direction and control — rather than choosing the contractual form by
looking only at the pension advantage. Concretely: (i) define the unit of engagement not by
hours worked but by tasks and deliverables (such as audit reports on a specified set of AI
outputs); (ii) impose no constraints on the time or place of work; (iii) substantively guarantee
the freedom to accept or refuse each individual engagement; and (iv) issue no concrete
direction or command as to the method of performing the work, giving feedback as
evaluation of deliverables. This is not a technique of evasion but its opposite. If the substance
is constant labor under direction and command, the person should be contracted as an
employee and given the protections of employment; if the non-employment form is chosen,
its discretion and freedom of refusal must be guaranteed in substance, not in name. That the
indirection of audit in Section 6.2 — the design that places the addressee of feedback on the
AI side rather than the person — is consistent with this protocol is no coincidence, and the
ceilings on load and the adequacy of compensation demanded by Brain Safety in Section 8,
only together with this protocol, close the path (Proposition 9) by which the arbitrage-driven
inducement turns into exploitation.
The exclusion of direction and control, however, itself generates a trade-off with
responsibility for quality. An ordering organization that has renounced direction and
Working Paper | Ageless Management in the AI Era 113
command cannot secure audit quality through the management of work performance, and
the discipline of quality migrates to the design of contractual liability. Here neither pole holds.
If unlimited liability for damages arising from an independently contracted super-senior's
erroneous audit — an oversight — is imposed, the individual takes on damage risks orders of
magnitude beyond the engagement fee, and a rational contractor will not accept the
engagement. Older contractors, who have little temporal room to recover losses through
work opportunities or asset formation after damage occurs, have rational grounds to behave
risk-aversely, and this chilling operates all the more strongly. Conversely, complete
exculpation for erroneous audits destroys the incentive to maintain the level of care, and
moral hazard arises. This paper's solution is the combination of two principles. The first is a
contractual cap on liability: the contractor's liability for damages is confined within an agreed
ceiling benchmarked to the engagement fee or the like, keeping it at an assumable level of
risk. The second is the principle that ultimate risk rests with the operating entity. As the legal
analysis of WP8 (Kadowaki 2026h) shows, supervisory responsibility in HOTL rests with the
operating entity, and shifting responsibility onto external auditors does not function as an
externalization of the duty of oversight — audit is an input into the operating entity's
decision-making, not a transfer of decision-making responsibility. Accordingly, of the ultimate
damage arising from errors in AI outputs that slipped past audit, the portion exceeding the
cap on liability is borne by the operating entity. If damages are not the instrument of
discipline, what then controls the moral hazard of erroneous audits? The division of roles is
clear. The allocation of catastrophic risk is the province of contract design including the cap
on liability, while the control of everyday levels of care is the province of the operational
governance of Section 8 — monitoring of each auditor's over-flagging rate (false-alarm rate),
temporary suspension and recalibration of audit authority when thresholds are exceeded,
and the reputation mechanism based on records of audit performance (Section 8.2). This
division of labor, which does not use damage claims as an instrument of everyday control,
prevents the chilling of engagement and, at the same time, places the effectiveness of control
on the side of ex-ante measurable operational indicators rather than on ex-post damages that
are difficult to prove.
In Japan, legislation partially filling this vacuum has recently begun to move. First, the Act
on Ensuring Proper Transactions Involving Specified Entrusted Business Operators (Act No.
25 of 2023; commonly, Japan's Freelance Act) came into force on November 1, 2024. The Act
prescribes, as obligations of ordering businesses when outsourcing work to specified
entrusted business operators (freelancers who employ no workers): (i) clear indication, in
writing or equivalent, of transaction terms (the content of the deliverable, the amount of
remuneration, and so on); (ii) setting a remuneration payment date (within 60 days of receipt
of the deliverable) and payment by that date; (iii) prohibition, in continuing outsourcing
arrangements, of refusal of receipt, reduction of remuneration, returns, unreasonably low
pricing, and the like; (iv) accurate display of recruitment information; (v) accommodation for
balancing work with childcare, family care, and the like; (vi) establishment of systems for
Working Paper | Ageless Management in the AI Era 114
harassment countermeasures; and (vii) advance notice at least 30 days before mid-term
termination and the like of continuing outsourcing arrangements. Jurisdiction lies with the
Fair Trade Commission, the Small and Medium Enterprise Agency, and the Ministry of Health,
Labour and Welfare. It is the first statute to extend, cross-cuttingly, subcontracting-act-style
transaction fairness and work-environment improvement (a partial extension of labor-lawtype
protection) to non-employment workers outside employment labor law, and it can be
called the starting point of the institutional infrastructure for non-employment work.
Second, the scope of special enrollment in workers' accident compensation insurance
(Articles 33 et seq. of the Industrial Accident Compensation Insurance Act) was expanded on
November 1, 2024, the same day the Freelance Act came into force. Freelancers engaged, as
specified entrusted business operators, in business performed under outsourcing from
enterprises and the like became newly eligible for special enrollment (work of the same kind
entrusted by consumers is also covered), and, going beyond the existing individually
designated sectors (construction, IT freelancers, and so on), essentially all freelancers became
able to enroll voluntarily in workers' accident compensation insurance. Premiums are
entirely self-funded; the Class II special enrollment premium rate is the basic daily benefit
amount × 365 × 3/1000 (0.3%), and the basic daily benefit amount is chosen from 16 tiers
between 3,500 yen and 25,000 yen.
These two pieces of legislation show that the vacuum of Proposition 9 has been recognized
by legislators as well, and that its filling has begun. But the filling is partial. First, what the
Freelance Act provides is fairness of transactions, not a floor on remuneration corresponding
to the minimum wage. Clear indication of the remuneration amount is mandated, but the
adequacy of its level is left to the market. Second, special enrollment in accident insurance is
voluntary and fully self-funded in premiums, asymmetric with the automatic application and
employer contribution of the employment form. The decision whether to enroll, and its cost,
are shifted onto the side with weaker bargaining power. Third, no correction of collective
bargaining power (a mechanism corresponding to trade-union law) has been put in place for
non-employment forms. The labor–management consent procedure required by the
entrepreneurship-support measures of the Act on Stabilization of Employment of Elderly
Persons is a control at the entrance to conversion to outsourcing; it does not answer the
continuing disparity in bargaining power after conversion. Accordingly, the refutation
condition of Proposition 9 — the existence of jurisdictions granting non-employment
participants protection equivalent to employment (accident compensation, remuneration
adequacy, and correction of bargaining power) — remains unmet, at least in Japan. The
vacuum has shrunk, but it has not disappeared. The anti-exploitation side of Brain Safety
(Definition 8) in Section 8 seeks to fill this residual vacuum with the organization's design
norms.
Working Paper | Ageless Management in the AI Era 115
7.3 Youth Participation — Child-Labor Regulation and the Design Space for
Education-Based Participation
For youth, who stand at the other end of the multigenerational ecosystem, the character of
the institutions is inverted. Whereas the institutional problem on the older side is the
removal of work disincentives, the institutions on the youth side take protection from labor
as their first principle. This protection is not an obstacle to be circumvented but a premise of
design. Can forms be designed that connect the cognitive assets of youth to the ecosystem
without impairing the protective principle of the priority of education and development —
this is the institutional question on the youth side.
At the base of the international standards lie two ILO conventions. The Minimum Age
Convention, 1973 (No. 138) prescribes a three-tier structure for the minimum age for work.
First, the minimum age may not be below the age of completion of compulsory schooling and
in no case below 15 (Article 2(3)). Second, national laws may permit light work (work not
harmful to health, development, or school attendance) by persons aged 13 to 15 (Article 7).
Third, the minimum age for hazardous work likely to jeopardize health, safety, or morals is
18 (Article 3; an exception in Article 3(3) permits it from 16 subject to protection and
training). For member states whose economic and educational institutions are insufficiently
developed, there is a developing-country exception reading these as 14 in principle and 12–14
for light work (Article 2(4)). The one-sentence summary "the ILO minimum age is 15" is
inaccurate because it drops this three-tier structure (15 in principle, 13 for light work, 18 for
hazardous work). Japan ratified Convention No. 138 in 2000, and the minimum age is
consistent, in a form connected to completion of compulsory education, with Article 56 of the
Labor Standards Act discussed below.
The Worst Forms of Child Labour Convention, 1999 (No. 182) obliges immediate measures
for the prohibition and elimination, for children under 18, of the "worst forms of child
labour": slavery, forced labor, and human trafficking; child soldiers; sexual exploitation; illicit
activities such as drug trafficking; and dangerous and harmful work. On August 4, 2020, the
Kingdom of Tonga ratified as the 187th member state, making it the first convention in ILO
history to be ratified by all member states (universal ratification). Adoption (1999) and the
attainment of unanimous ratification (2020) are distinct events. Japan ratified in 2001, and
the United States, which has not ratified No. 138, has ratified No. 182. The fact that the
protection of children is the domain that has exceptionally reached universal consensus
among labor standards whose ratification status is otherwise divided indicates the
international foundation of the design principle Proposition 10 places on the youth side — the
priority of education and health. At the EU level, the Young Workers Directive (Council
Directive 94/33/EC, adopted June 22, 1994) prescribes the prohibition in principle of work by
children (under 15 or in compulsory education), an exception structure of permits for
cultural and artistic activities, work practice and light work from 14, and light work for
Working Paper | Ageless Management in the AI Era 116
limited weekly hours from 13, together with regulation of night work and hazardous work
for those under 18, giving the same three-tier idea form in EU law.
In Japanese domestic law, Chapter 6 "Minors" of the Labor Standards Act (Act No. 49 of
1947) (Articles 56–63) corresponds to this. Article 56 prohibits the use of a child as a worker
until the end of the first March 31 after the child reaches 15 (there is an exception whereby, in
non-industrial undertakings, children aged 13 or over may be employed outside school hours,
with the permission of the administrative agency, in light work not harmful to health and
welfare, and in film production and theatrical undertakings the same permission makes this
possible even under 13). Article 57 mandates keeping age certificates and the like for persons
under 18, Article 58 prohibits the conclusion of labor contracts by parents or guardians on
behalf of a minor, and Article 59 provides that minors may claim wages independently.
Article 60 prohibits, in principle, overtime and holiday work for minors, and Article 61
prohibits, in principle, night work (from 10 p.m. to 5 a.m.) by persons under 18. Article 62
prescribes restrictions on employment in dangerous and harmful work, and Article 63 the
prohibition of work in mine pits. The range of article numbers requires care: dangerous and
harmful work is Article 62, mine work Article 63, and the whole of minor protection is
"Articles 56–63."
How, then, does youth participation outside employment — internships, PBL (project-based
learning) educational partnerships, contest-based participation — intersect with this
regulation? The key is the determination of worker status. Article 9 of the Labor Standards
Act defines a worker as "one who is employed at a business ... and to whom wages are paid,"
and for internships an administrative circular (Circular Kihatsu No. 636 of September 18,
1997) supplies the framework of judgment. That is, where the benefit or effect of the work in
question accrues to the establishment, as when the student engages directly in production
activity, and a relationship of use and subordination is found between the establishment and
the student, the student qualifies as a worker. Conversely, where the training is observational
or experiential and no relationship of use and subordination is found, as where the student is
not considered to receive business-related direction and command from the employer, the
student does not qualify as a worker. The judgment is made by substance, not by name
(internship, PBL, practicum), and considers (i) the presence or absence of direction and
command, (ii) the accrual of outcomes and benefits to the establishment, and (iii) the
presence or absence of attendance management and sanctions, among other factors. If judged
a worker, the Labor Standards Act, the Minimum Wage Act (the obligation to pay wages at or
above the prefectural minimum wage), and workers' accident compensation insurance apply.
The three-ministry agreement of the Ministry of Education, Culture, Sports, Science and
Technology, the Ministry of Health, Labour and Welfare, and the Ministry of Economy, Trade
and Industry (Basic Approach to the Promotion of Internships, revised 2022) is a policy
document organizing the relationship with recruiting activities; the criteria for worker status
themselves are those of the circular above.
Working Paper | Ageless Management in the AI Era 117
This two-sidedness directly yields the design principles for youth participation. The first
side — one cannot say "student interns may go unpaid." If a student is made to work under
direction and command in a manner whose benefits accrue to the business, that is labor, and
it receives the application of the minimum wage and accident insurance. The procurement of
unpaid labor borrowing the name of education is impermissible both legally and under the
normative argument of Proposition 10. The second side — for observational, experiential, or
learning-type programs designed for educational purposes that avoid incorporation into
operations and direction and command, worker status is denied as a rule, and here lies the
legal space to design education-purpose PBL institutionally as "learning" rather than "labor."
The youth PBL participation conceived in Section 6 — youth-perspective feedback on AIgenerated
outputs, prototyping exercises, contest-style problem submissions — can be
designed inside this space so long as educational purpose is placed first, business use of the
outputs second, and no structure of direction and command is adopted.
This design space, however, has three limits. First, the boundary is continuous, and there is
a constant danger that participation begun as education-based slides into de facto labor as
dependence on its outputs grows. Because worker status is a judgment of substance, the name
given at design time is no rampart. Second, where worker status is denied, young participants
are placed outside the protections of labor law (wages, accident compensation). This is a
protection vacuum of the same shape as the non-employment form on the older side
(Proposition 9), and it needs to be filled by the educational institution's management and
insurance and by the host's duty of care for safety. Third, contest-based participation (prizebased
open calls for problem solutions and the like) adopts a structure in which, out of the
unpaid work products of many participants, compensation goes only to a few winners, so that
designs poor in educational benefit approach the de facto procurement of unpaid labor.
Proposition 10's claim — that the protective principle on the youth side (the priority of
education and health) and the prevention of exploitation on the older side are responses to
the same structure, the asymmetry of bargaining power and exit costs — takes concrete form
in these three limits. Section 8's Brain Safety translates this symmetry into design norms.
7.4 Summary Table of the Institutional Comparison
Table 6 organizes the institutions treated in this section across jurisdictions. The details of
each institution (article numbers, amounts, sources) are left to Table A2 in Appendix B; here,
the implications for the implementation of Ageless Management are mapped to each.
Table 6 International comparison of institutions surrounding older-age and youth work, and their
implications for Ageless Management
Jurisdiction
Institution (effective/
amended year) Key points
Implications for
Ageless Management
Japan Act on Stabilization of
Employment of Elderly
Prohibition of mandatory
retirement below 60 (Article 8).
Legal obligation to secure
Participation forms
outside the employment
boundary formalized
Working Paper | Ageless Management in the AI Era 118
Persons (age-70 measures
in 2021)
employment up to 65 (Article 9).
Effort obligation to secure work
up to 70 (Article 10-2). Options
include outsourcing contracts and
social-contribution projects
(entrepreneurship-support
measures = non-employment
forms)
for ages 65–70. But this
remains an effort
obligation, and nonemployment
forms
carry no accompanying
protection (Section 7.2)
Japan Increase of the threshold
of the in-work old-age
pension offset (April
2026)
Threshold raised from 510,000 yen
(FY2025) to 650,000 yen (FY2026;
statutory value 620,000 yen plus
wage indexation). The mechanism
itself of suspending one half of the
excess is maintained. The
suspension also applies to
employed persons aged 70 or over
The range of "working
means losing" shrinks
but does not vanish.
Supplying highremuneration
Gc-type
roles in employment
form must take account
of the residual workadjustment
incentive
Japan Freelance Act (November
2024); expansion of
special enrollment in
accident insurance (same)
Obligations toward nonemployment
workers: clear
indication of transaction terms,
payment within 60 days, 30 days'
advance notice of mid-term
termination, and so on. Voluntary
special enrollment in workers'
accident insurance expanded to
essentially all freelancers
(premiums self-funded)
Partial filling of
Proposition 9's vacuum.
But there is no
remuneration floor or
bargaining-power
correction, and accident
compensation is
voluntary and selffunded
— the vacuum
remains
United
States
Senior Citizens' Freedom
to Work Act (2000)
Earnings test abolished at and
after FRA. Before FRA, $1 per $2
($1 per $3) of earnings above the
exempt amounts (2026: $24,480/
year; $65,160 in the year of FRA
attainment) withheld (effectively
recovered after FRA)
A precedent removing
work disincentives from
FRA onward. The design
bounded by an age
threshold itself remains
United
States
ADEA (1967; amended
1986)
Prohibition of age discrimination
against those 40 and over.
Removal of the upper age limit
made mandatory retirement
unlawful in most occupations
(limited exceptions)
Legal foreclosure of
forced exit by calendar
age — the example that
carries Definition 1's
direction furthest in
employment law
Germany Abolition of the
Hinzuverdienstgrenze
(January 2023)
Earnings limit while drawing an
early old-age pension completely
abolished (permanent measure).
For disability pensions, not
abolition but dynamic limits
An example of full
abolition of pensionside
work disincentives.
Demonstrates that
complete removal of
"working means losing"
is legislatively possible
EU Employment Equality
Directive 2000/78/EC
(2000)
Prohibition of employment
discrimination including age. But
Article 6 gives age alone a
Age receives special
treatment even inside
anti-discrimination law
Working Paper | Ageless Management in the AI Era 119
justification exception (the basis
provision for permitting
mandatory retirement)
— evidence of the
institutional deeprootedness
of the
variable of calendar age
Singapore Retirement and Reemployment
Act (to 64/69
in July 2026)*
Two-tier structure of statutory
retirement age plus mandatory reemployment-
offer age. Policy of
65/70 by 2030
The gradualist type of
extending work through
mandates. The reemployment-
offer
obligation is structurally
close to Japan's
continued employment
South Korea Mandatory retirement
age of 60 (phased in
2016/2017)
Mandated setting the retirement
age at 60 or above. The same Act
also prescribes the prohibition of
age discrimination. Extension to
65 under discussion
An example of
mandating retirement
ages in progress. The
very existence of a
statutory retirement age
is also a preservation of
the calendar-age
criterion
ILO Convention No. 138
(1973); Convention No.
182 (1999)
Three-tier minimum-age structure
(15 in principle, 13 for light work,
18 for hazardous work;
developing-country exception
14/12). No. 182 ratified by all
member states in 2020
Youth participation
must be designed inside
the protective principle
(priority of education
and health) — the
international
foundation of
Proposition 10
Japan Labor Standards Act,
Chapter 6 on minors
(Articles 56–63); Circular
Kihatsu No. 636 (1997)
Minimum age, prohibition of night
work (10 p.m.–5 a.m.), restrictions
on dangerous and harmful work,
and so on. Worker status of
interns and the like judged by the
substance of the relationship of
use and subordination
Both faces: the legal
space to design
education-purpose PBL
as "learning," and the
risk of sliding into
unpaid labor (Section
7.3)
Note: The formal statute names, act numbers, article numbers, amounts, and sources for each institution are given in
Appendix B (Table A2). This table includes only institutions, verified against primary sources or professional
commentary as of August 2026, that are in force or whose effective date is fixed. *Singapore's 2026 amendment rests
on multiple cross-checks of law-firm and HR-media commentary; final confirmation against primary sources of the
responsible ministry (MOM) remains outstanding (see the evidence-grade note at the end of this section). The
descriptions of the institutions are summaries; the details of application are governed by the respective statutes and
circulars.
7.5 Implications of Propositions 9 and 10 — Ageless Management as a Function
of Institutional Trends
Three implications are drawn from this section's comparison. First, the institutions that have
restrained employment-based work at older ages are moving toward contraction across
jurisdictions. On the work disincentives of in-work pensions, the United States moved in 2000
to abolition at and after FRA, Germany in 2023 to complete abolition including early
Working Paper | Ageless Management in the AI Era 120
claimants, and Japan in 2026 to a substantial increase of the threshold. On the side of
employment opportunity as well, the outer bound of working ages is being extended in each
country: Japan's age-70 effort obligation (2021), Singapore's increases (2026), and South
Korea's mandatory retirement age (2016–2017). Institutions that create "working means
losing" and forced exit by calendar age are, at least as a legislative trend, in retreat. For
Ageless Management, this is a tailwind among the givens.
Second, however, this trend is skewed toward the employment form. As Proposition 9
points out, protection for the non-employment participation on which Ageless Management
depends is incomplete. Japan's Freelance Act and the expansion of special enrollment in
accident insurance (both November 2024) show the beginning of the filling of the vacuum,
but lack all of a floor guarantee on remuneration, correction of bargaining power, and
automatic application of accident compensation. While the entrepreneurship-support
measures (2021) formalized non-employment work opportunities, the very time lag —
protection catching up only partially, three and a half years later — shows that the expansion
of opportunity and the development of protection move at different speeds. Organizations
that build multigenerational ecosystems on non-employment contracts must therefore treat
the law's required level as a floor of protection, not a ceiling, and bear the responsibility of
filling the residual vacuum with their own design — Section 8's Brain Safety. Ageless
Management that neglects this can, as Proposition 9 warns, degenerate into exploitation: on
the older side into unpaid advisorship, and on the youth side into unpaid labor borrowing the
name of education.
Third, the institutions on the youth side and the older side, while apparently opposite —
protection from labor on one side, inclusion into labor on the other — can be read as
responses to the same structure. What the three-tier structure of ILO No. 138 and the minor
provisions of the Labor Standards Act protect is the condition that opportunities for
education and development not be eroded by labor under an asymmetry of bargaining
power. What exploitation prevention on the older side protects is the condition that the
compensation and cognitive load of those in positions with high exit costs be kept adequate.
Proposition 10 (the symmetry of Brain Safety) integrates the two as symmetric design
principles (priority of education and health, ceilings on load, adequacy of compensation)
responding to the same structure of bargaining-power asymmetry and exit costs. This
section's institutional comparison has shown that one side of this symmetry (youth) has
reached nearly universal consensus in international law, while the other side (older, nonemployment)
is still in formation in each country.
In sum, the implementability of Ageless Management is a function of institutional trends.
The more institutions shrink "working means losing" and expand protection for nonemployment
forms, the lower the implementation costs of Definition 1's dynamic role
allocation and Definition 7's multigenerational ecosystem. Conversely, the longer the
protection vacuum is left unattended, the more implementation deepens its dependence on
the organization's autonomous norms, and implementation without norms raises the risk of
Working Paper | Ageless Management in the AI Era 121
exploitation. That this paper positions Brain Safety in Section 8 not as a merely desirable
consideration but as a condition for the validity of Ageless Management (the conditional
clauses of Propositions 8 and 11) is a consequence of this institutional reality.
Evidence-grade note: The institutional descriptions in this section are limited to matters checked against
primary statutes and administrative sources (e-Gov, the Ministry of Health, Labour and Welfare, the
Japan Pension Service, the SSA, the U.S. Congress, the EEOC, the DRV, EUR-Lex, and the ILO) and against
public research institutes (JILPT) and professional commentary. Singapore's 2026 amendment rests on
multiple cross-checks of law-firm and HR-media commentary; final confirmation against primary
sources of the responsible ministry (MOM) remains outstanding. This section is a description of
institutions, not legal advice.
8. Brain Safety: Occupational Health for the Brain
Definition 8 defined Brain Safety as occupational health and safety standards that protect the
brain capital of ecosystem participants from depreciation. It is bidirectional: (i) protection
from cognitive load and brain fatigue, and (ii) protection from conversion into unpaid labor
and from exploitation arising from asymmetries of bargaining power. As the preceding
section showed, the non-employment forms of participation on which Ageless Management
depends fall into the protection gap of current institutions (Proposition 9). Until institutions
catch up — and after institutions are in place as well — the designers of an ecosystem must
have health and safety standards of their own. This section presents the skeleton of those
standards. Section 8.1 addresses the cognitive-load side and Section 8.2 the anti-exploitation
side; Section 8.3 shows the theoretical connection to the framework of this series' BCM
(Kadowaki 2026e), and Section 8.4 consolidates the results into Table 9 (the skeleton of the
guidelines).
8.1 The Cognitive-Load Side — The Individual Version of the Solvency Condition
and Dual-Track Auditing
Conventional occupational health and safety has been designed around physical load and
hazardous work. The audit-type roles that Ageless Management assigns as the default (Section
6) reduce physical load while shifting the center of gravity of the load to the cognitive side.
Verifying AI output continuously consumes working memory and attentional resources in the
form of sustained attention, document reading, and responses to alarms. In light of the
structure of cognitive aging confirmed in Section 2, this load structure constitutes, for superseniors
whose age-related changes lie on the side of processing speed and attention, a health
and safety problem of a different kind from physical load.
This paper formalizes this problem as the individual version of the solvency condition of
WP8 (Kadowaki 2026h). WP8's solvency condition states that there is an upper bound on the
cognitive resources, time, and cost an organization can devote to oversight, and that oversight
can be sustained only within that bound. The same logic operates inside the individual. An
Working Paper | Ageless Management in the AI Era 122
individual's attentional resources and cognitive effort are also subject to a budget constraint,
and auditing can be sustained only within that budget. What happens when auditing is piled
up beyond the budget has also already been described by WP8: alarm fatigue. The more
outputs there are to check, the more sluggish the auditor's responses become, and the
detection probability falls. That is, neglected cognitive load is at once a health and safety
problem that depreciates the auditor's own brain capital and a governance problem that
undermines the detection probability that is the source of oversight value. An overloaded
audit regime damages the person and fails to protect what it is meant to protect. It fails twice.
From this formalization, the three design principles listed in the original draft can be rederived.
The first is control of working time. Dividing an audit session into short units
(roughly 20–30 minutes as a guide) with rest in between is a budget-execution discipline for
keeping the rate of consumption of the attention budget within its ceiling. The commissioning
unit for audit tasks is likewise designed per session, not as continuous engagement. The
second is a multimodal work environment. Auditing that depends on screen gaze and reading
concentrates its consumption on visual attention. The combined use of voice interfaces (Voice
UI) and read-aloud is a candidate design for dispersing this concentration of consumption.
The third is interactive auditing. A format that progressively narrows the object of
verification through AI summaries and question-and-answer, rather than the bulk reading of
long documents, limits the amount of information that must be held at once. What must be
stated explicitly, however, is that all three principles are design hypotheses derived from the
attention-budget formalization, and their effects on audit quality and brain health have not
been tested. That testing belongs to the measurement framework of Section 9.
Managing the budget constraint should not be left to the individual auditor's self-restraint.
The parties who commission audits usually do not know how much of an auditor's attention
budget each request consumes, and if multiple commissioners pile requests onto the same
auditor, the total exceeds the budget even when each individual request is reasonable. The
cognitive-load management of Brain Safety must therefore be designed as a pair: session
design (execution discipline on the individual side) and total-volume management of audit
demand (commissioning discipline on the organizational side). Concretely, the demand-side
management tools include visualizing each auditor's acceptance ceiling, prioritizing requests
(not having every output audited equally), and AI-side design that reduces the volume of
output requiring audit in the first place (such as flagging low-confidence outputs). To
acknowledge the scarcity of oversight is the same thing as to use oversight with care. The
effects of this prioritization and pre-screening on the independence of the audit channel,
however, are addressed explicitly at the end of this section.
The rarity effect that WP8 identified can also be reread at the individual level. The rarer
errors are in an environment, the harder it is to sustain a supervisor's vigilance. As AI
performance rises, errors become rarer, auditing becomes the task of confirming "almost
always fine," and attention adapts to that environment and declines. The experimental
finding seen in Section 4.2 — that human vigilance slackens most under imperfect but high-
Working Paper | Ageless Management in the AI Era 123
performing AI — is consistent with this mechanism. The design implication is not to entrust
the maintenance of vigilance to the auditor's will. Candidate devices include exercises that
deliberately mix known errors into the audit stream to confirm detection (an operational
version of the "embedded errors" adopted by the experimental task of Hypothesis H3),
rotation of audit targets to prevent habituation, and the sharing of detection cases to update
the fact that "errors are real." These, too, are design hypotheses, and testing their
effectiveness belongs to Section 9.
At the same time, load ceilings must not be made uniform by age. Age-based protection
such as "halving audit time across the board because someone is old," however well
intentioned, is an operation that reintroduces, in the name of protection, the chronologicalage
variable that Definition 1 removed — the same slide as in dynamic role allocation
(Section 6.2): protective stereotyping. As confirmed in Section 2, age-related changes in
attention and processing speed show large individual differences and overlapping
distributions, and load tolerance is no exception. The load management of Brain Safety
should be designed not as uniform age-specific standards but as individualized ceilings based
on measured states (fatigue, performance degradation, and the person's own reports). The
individualization of protection is the consistent application of the principle of Ageless
Management to occupational health.
How, then, is the state of load to be measured? There are three candidate layers. First, selfreport
(periodic short scales of fatigue and perceived load). Second, behavioral indicators
(degradation signals derived from work logs, such as changes in audit response time, declines
in detection rate, and increased misses in the latter half of a session). Third, participation
patterns (the frequency of session interruptions and declinations). Behavioral indicators may
be more sensitive than self-report, but continuous collection is one step away from turning
auditors into objects of surveillance, and the design requirements are the separation of the
purpose of measurement (protection) from its use (evaluation and selection), feedback to the
person, and consent to the scope of collection. The danger that measurement is diverted from
protection to selection is taken up again as self-critique in Section 10.
Operational restriction of unannounced verification — anonymized system
calibration. The danger of surveillance appears most sharply in the unannounced, real-time,
blind in-situ verification (unannounced in-situ verification) that Proposition 12 specifies as an
operational requirement of dynamic role allocation. This section fixes, from the Brain Safety
side, the operational principle stated in Section 5 as a restrictive clause. Unannounced in-situ
verification is operated, in the first instance, in anonymized and aggregated form for systemlevel
calibration — testing the divergence between backtest results and real-time observation.
Its use for individual treatment (such as changes in role allocation) is permitted only through
aggregation over long windows and due process including disclosure to the person and an
opportunity to respond; demotion or revocation of authority on the basis of a single
unannounced result is not performed. This restriction is not merely the general fairness
requirement of restraining arbitrariness; it is an intrinsic requirement of Brain Safety. Under
Working Paper | Ageless Management in the AI Era 124
a regime in which one cannot know when one is being observed and a single observation can
directly determine one's treatment, participants are forced to behave under the assumption
of constant surveillance — the panopticon effect. What the sense of constant surveillance
supplies is not discipline but a permanent state of vigilance, which continues to occupy
cognitive bandwidth as an unerasable claim on attentional resources. That is, individually
surveillant unannounced verification erodes, in the name of governance, the very thing this
section seeks to protect as the extension of Base — cognitive bandwidth — and contradicts the
purpose of Brain Safety head-on. The deterrence of gaming is achieved by publicizing the fact
that the system is calibrated, and does not require diverting individual observations to
individual surveillance (Section 5, commentary on Proposition 12). This two-layer separation
— anonymized system calibration, and use for treatment only over long windows with due
process — is included in Table 9 as the only operational mode that reconciles Proposition 12's
anti-gaming safeguard with Brain Safety.
Consent-based load adjustment. How, then, should protection be activated when a
measured state crosses the protective threshold? This paper does not adopt unilateral forced
cutoff by the firm — immediate exclusion from audit work or unilateral termination of the
contract on the grounds of fatigue or degradation indicators — as the mode in which Brain
Safety is activated. There is a conflict here that must be made explicit. Unilateral cutoff is an
operation that overrides, in the name of protection, the person's will to work and right of selfdetermination,
and it can slide into a variant of ageism that merely replaces the
chronological-age variable removed by Definition 1 with fatigue indicators — even when
"because you are old" becomes "because the data say so," it remains exclusion that bypasses
the person's will. On the other hand, leaving degradation signals unattended damages both
the person's brain capital and the detection probability of the audit (earlier in this section).
This paper's design for this conflict is risk alerts and consent-based dynamic contract
renewal. That is, objective data — degradation signals derived from work logs and measured
values of fatigue — are used first as the presentation of a risk alert to the person, and load
adjustments — reduction of session counts, changes of format, pause and resumption — are
implemented as renewals of contract terms under the person's consent (informed consent),
accompanied by an explanation of the meaning of the measurements and of the options, and,
where necessary, with the involvement of an intermediary organization (Sections 8.2 and
6.4). The involvement of the intermediary is a bulwark against "consent" turning into de facto
coercion under asymmetric bargaining power. Even this mode, however, leaves a residual
tension. If the person does not consent to adjustment while degradation progresses, the
collision between responsibility for audit quality and the person's self-determination does not
disappear. And this situation is not exceptional but the normal case the design must
anticipate. Underestimation of one's own functional decline — anosognosia (reduced illness
awareness) and inflated self-assessment — can itself be a degradation of metacognitive
calibration, so consent-based adjustment carries a structural blind spot: consent is hardest to
obtain precisely in the situations that most require adjustment. Choosing neglect here is a
Working Paper | Ageless Management in the AI Era 125
collapse of governance (damage to both audit quality and the person's brain capital);
choosing forced cutoff is a relapse into the violation of self-determination. Against this
deadlock, this paper specifies a third path as an escalation procedure: an objective second
opinion by a third-party body independent of the interests involved — an occupational
physician, an external intermediary organization (Sections 6.4 and 8.2), or the like. The third
party evaluates the validity of the measurement data and the reasonableness of the proposed
adjustment independently of both the person and the commissioning organization, and its
findings are used both for re-presentation to the person and for revision of the adjustment
proposal. Only if agreement still fails to form after this procedure are organizational
measures justified, in an order that puts reassignment according to the risk level of the
audited domain (transfer to low-risk domains) first and termination of participation as the
last resort. That no general solution to the collision exists remains true. But specifying the
escalation path and the order of measures in advance is a bulwark against both arbitrary
cutoff and irresponsible neglect.
The cognitive-load side has one further aspect. As seen in Section 4.3, dependence on
assistance can depreciate skills that are not exercised (Budzyล et al. 2025). If the reduction of
audit load is pursued to the point of full reliance on AI pre-screening and pre-summarization,
the auditor's own detection skill depreciates and the source of oversight value is lost.
Restraining cognitive load and maintaining skill therefore involve a trade-off, and load design
must incorporate deliberate opportunities for exercise with AI assistance removed — what
BCM calls protected unassisted practice (Section 8.3). Designing the ceiling on load and the
floor on exercise simultaneously is the requirement of the cognitive-load side of Brain Safety.
There is another, independent channel of damage to be placed alongside skill depreciation:
damage to the independence of the audit channel. The load-reduction designs listed in this
section — AI summaries, interactive auditing, confidence-based prioritization — lower
cognitive load while at the same time making the information channel of the act of auditing
dependent on AI. As Proposition 5 makes explicit, generational decorrelation widens the
group's detection set under the condition that auditors access the original output in
unmediated form; if all auditors share summaries and pre-screening by the same AI,
generational heterogeneity is homogenized at the entrance of the audit, and misses recorrelate
at the channel level (Section 6.3). That is, pushing load reduction to the limit not
only depreciates the individual auditor's detection skill (previous paragraph) but can also
damage the group's independence — the other factor of oversight value. Load reduction and
independence stand in a trade-off, and this trade-off must be managed explicitly within the
design.
Dual-track auditing. Neither pole of this trade-off is viable as a design. Full-volume
unmediated auditing exceeds the ceiling on the attentional resources that can be devoted to
oversight — the solvency condition of the organization and of the individual — and collides
head-on with the load-management requirements of the first half of this section. Full-volume
AI-summarized auditing lowers the load but breaks the channel condition of Proposition 5
Working Paper | Ageless Management in the AI Era 126
and nullifies generational decorrelation at the entrance of the audit. This paper's design
therefore takes a dual structure. As the first track, for a portion sampled at random from all
outputs, a raw-audit slot is preserved in which the raw output is audited without summaries
or pre-screening (unmediated, unsummarized, unscreened). As the second track, the
remaining outputs receive screening by AI summaries, interactive auditing, and
prioritization, economizing on the attention budget. That the sampling is random is essential:
raw auditing of the "seemingly important parts" selected by an AI reintroduces channel
sharing through the selection itself, and equally excludes from the first track precisely the
places where the AI errs with confidence (Sections 6.3 and 5).
Sampling design — the breakdown of fixed rates and risk-weighted dynamic sampling.
How the sampling rate of the first track is designed determines the success or failure of the
dual track. First, fixed-rate random sampling necessarily breaks down as processing volume
grows. The capacity arithmetic is simple. The attention required for raw auditing grows in
proportion to (sampling rate × total output volume), whereas the attention that can be
supplied is capped at (number of auditors × per-person attention budget). As long as AI
throughput keeps growing, any fixed sampling rate will eventually exceed this supply ceiling
— a fixed rate collides head-on with the solvency condition. Conversely, if the sampling rate is
lowered continually to stay within the supply ceiling, coverage declines in inverse proportion
to total volume. The limit on oversight coverage that WP8 identified is not dissolved by the
dual track; it is merely reallocated in the form of sampling.
For that reason, the allocation of the scarce raw-audit slots should not be indiscriminate.
This paper's design is risk-weighted dynamic sampling. Outputs are stratified by impact,
based on the severity of the consequences of error — irreversibility (the possibility of ex-post
correction), scope of propagation (the number and kinds of parties affected), and monetary
scale — with the sampling rate raised for high-impact strata and relatively lowered for lowimpact
strata. Sampling within each stratum remains random. Two disciplines attach to this
weighting. First, impact weights must be defined from the attributes on the consequence side
of the output, independently of the AI's self-assessment. In particular, the AI's confidence
must not be used as grounds for lowering the sampling rate as "evidence of low risk." Errors
accompanied by high confidence are precisely the typical manifestation of hallucination, and
confidence-linked pre-screening systematically removes from the audit the places where the
AI errs with confidence — exactly the configuration against which the findings on automation
bias and alarm fatigue in Section 4.2 and WP8 warn. Using low confidence as a trigger for
additional checking (earlier in this section) is permissible; using high confidence as grounds
for exemption from audit is not — the discipline is asymmetric. Second, a purely random
floor that obeys no impact weighting whatsoever (a minimum sampling rate applied across
the entire output domain) is maintained. If all outputs were selected according to the impact
function, the selection function itself would become a new channel shared by all auditors,
and errors in the domains left outside the impact assessment would reach no one's eyes. The
role of the floor is twofold. Its first role is the insurance against fallibility just mentioned — a
Working Paper | Ageless Management in the AI Era 127
safeguard against errors and blind spots in the stratification design itself. Its second role is an
exploration slot for unknown unknowns. The impact weighting by irreversibility, scope of
propagation, and monetary scale is constructed in reliance on knowledge of what brings
about grave consequences — that is, on the lessons of past accidents and failures. Novel
modes of breakdown outside that weighting — errors classified as low-impact precisely
because they resemble none of the past accident types, or errors that do not appear in the
variables of the impact function at all — therefore, by definition, never ride the risk-weighted
track. Only a purely random floor that relies on no existing risk knowledge can, in principle,
catch errors of this kind — exploratory sampling beyond the distribution (exploration) as
against sampling optimized to the known distribution (exploitation). In this twofold role, the
floor is the last safeguard of the channel condition of Proposition 5 — the preservation of
decorrelation.
Forced queue. Sampling design must also include control of the later stage of sampling —
the phase in which the sampled outputs are actually read. Even if sampling is perfectly
random, if auditors can themselves choose what to read from among their assigned outputs, a
second selection by preference arises inside the track. An auditor under an attention-budget
constraint digests short, plain, readable outputs first and defers long and complex ones — and
the deferred outputs are then often left unaudited when the session ends. This auditor-side
selection bias (cherry-picking) is a natural budget-saving behavior requiring no ill will, yet its
consequences are systematic: long and complex outputs are, in many cases, high-risk outputs
with intricate judgments, and selection by preference erodes the coverage that random
sampling was supposed to guarantee precisely from the high-risk side. The countermeasure is
algorithmic delivery control. Assignments within the raw-audit track are delivered as a
forced queue in random order that the auditor cannot choose, and the auditor processes
them in the order received. In addition, the volume of each delivery is calibrated to what can
be digested within the session time ceiling (the 20–30 minutes earlier in this section) — for if
the volume exceeds the budget, truncation of the tail revives selection even when the order is
forced. Audit difficulty arising from low readability is handled not by deferral but by splitting
the session and assigning an additional session — since this is the raw-audit slot, lightening
by summarization would violate the principle of non-mediation. The forced queue is the
enforcement mechanism that insulates the design decision of random sampling from the
individual auditor's preferences.
Even under this refinement, the residue of Type II error — errors slipping through — does
not disappear. It can be disclosed quantitatively. If the raw-audit sampling rate of a given
stratum is p and the miss rate of AI-summarized screening is β, the probability that an error
in that stratum is captured by neither track is approximately (1−p)×β, which is positive as
long as p is below 1 and β is positive. Risk weighting lowers the slip-through probability of
high-impact errors, but does not bring it to zero, and slip-through in the domains classified as
low-impact is, indeed, tolerated by design. Accordingly, even after the introduction of risk
weighting, it remains the case that an operation claiming "we are safe because a fixed rate of
Working Paper | Ageless Management in the AI Era 128
raw auditing is secured" supplies only reassurance detached from the reality of coverage —
no fixed guarantee of safety exists. Ex-ante disclosure of the sampling rates (including
stratum weights and the floor) with periodic revision in response to changes in processing
volume and risk, disclosure of estimated slip-through probabilities, and measurement of
detection performance by channel (raw audit vs. AI summary) (Section 9) are the minimum
requirements for keeping this design variable open to verification.
Dual-track auditing is the independence version of the protected unassisted practice of BCM
(Kadowaki 2026e). Just as unassisted practice is the floor on exercise that prevents the
depreciation of an individual's skill, the raw-audit slot is, in the same form, the floor on
independence that prevents the erosion of the group's decorrelation. Load reduction, skill
maintenance, and independence maintenance cannot all be maximized at once, and the
cognitive-load design of Brain Safety is operated as a constrained design problem that
simultaneously specifies the ceiling on load, the floor on exercise, and the floor on nonmediation
(the raw-audit slot including the random-sampling floor).
Layer separation of audit feedback. The errors and corrections detected by audits are fed
back indirectly into the improvement of AI prompts, evaluation criteria, and verification
procedures (Section 6.2). To the governance of this feedback path, this paper adds one further
requirement: layer separation — human corrections and their interpretive context are not
fed in full volume into the retraining of the AI's base model. The ground for concern lies in
the demonstration of model collapse: when a generative model is trained continually on
recursively generated data, the tails of the original data distribution are lost and the model
loses diversity (Shumailov et al. 2024). Since the object of audit feedback is AI-generated
output and its corrections, an operation that recycles this into retraining without selection
has the structure of recursively biasing the training data toward self-generated artifacts and
their vicinity. What the demonstration directly addressed was recursive training on modelgenerated
data, and extrapolation to feedback containing human corrections carries a margin
of uncertainty; but if degradation of generative diversity does occur, it works in the direction
of deepening from within the systematic blind spots of a single model against which model
condition (iii) of Proposition 5 warns. The design requirement is therefore a separation of
layers: the interpretive context of the audit — context metadata such as the reasons for
corrections, contextual information, and conditions of application — is managed at the layer
of prompts, retrieval references, and evaluation criteria, and input into the training of the
base model is limited to selected and recorded portions. Layer separation alone, however, is
not enough. The retrieval-reference (RAG) layer that manages interpretive context is itself an
accumulating database; if it saturates with the past judgments of a small number of auditors,
then even with the base model protected, the context the retrieval returns converges on
particular auditors and judgment types, and every subsequent audit and generation refers to
it — a loss of diversity of the same form as model collapse is reproduced at the RAG layer. This
paper therefore places, in addition to layer separation, two operational requirements on the
RAG layer: (i) temporal decay of logs — attenuating the reference weight of old judgments
Working Paper | Ageless Management in the AI Era 129
over time to prevent the entrenchment of past judgments — and (ii) source-diversity
weighting — imposing diversity constraints on the composition of references so that retrieval
results are not biased toward particular auditors, generations, or judgment types. These
requirements are included in Table 9.
8.2 The Anti-Exploitation Side — The Symmetry of Proposition 10
The second aspect of Brain Safety is economic protection. The participation forms of Ageless
Management have their center of gravity in non-employment types — outsourcing contracts,
advisory roles, PBL, and NPO partnerships. As Proposition 9 pointed out, these lie outside the
employment relationship that is the principal unit of labor law's protection and restraint, and
an ecosystem left unattended can degenerate into exploitation. There are three typical modes
of degeneration. The first is conversion into unpaid advisors. The practice of procuring the
experiential audits of super-seniors without pay or for nominal consideration, under
honorific titles such as "advisor" or "counselor," is nonpayment for the scarce resource of
experiential audit capacity (Sections 4–5). So-called exploitation through purpose (yarigai
exploitation) — using the logic that "social participation is itself the reward" or "it gives you
purpose in life" as a substitute for compensation — makes the intrinsic motivation of
participation a cover for asymmetric bargaining power. The second is the conversion of youth
PBL participation into unpaid labor in the name of education. The moment the commercial
acquisition of deliverables becomes the primary purpose and education becomes
subordinate, it is no longer PBL but disguised labor. The third is the consumption of the
"insider perspective" of socially marginalized groups as opinion-gathering unaccompanied by
compensation or attribution.
Proposition 10 claims that these apparently different problems share a single structure.
Labor protection for youth and the anti-exploitation and cognitive-load protection of superseniors
are both responses to the same structure — asymmetric bargaining power and exit
costs — and demand symmetric design principles. Youth are prone to accepting unfavorable
terms because of asymmetries of experience, information, and legal status; super-seniors
because of the scarcity of re-employment opportunities and the desire for social recognition;
marginalized groups because of the scarcity of participation opportunities themselves. What
the symmetry claim means in practice is that protection can be designed not as a bundle of
stratum-specific exceptions but as the application of a single design principle — priority of
education and health, ceilings on load, and the adequacy of compensation — to all
participating strata. Concretely: (i) require written contracts and explicit compensation for all
forms of participation, and set compensation for the provision of experiential audit and
insider perception with reference to market levels; (ii) require the primacy of educational
purpose for youth participation, institutionalizing the priority of schooling, restrictions on the
commercial use of deliverables, and the involvement of educational institutions (Section 7);
(iii) apply load ceilings (Section 8.1) to all strata; (iv) contractually provide non-employment
participants with accident compensation and consultation channels comparable to those of
Working Paper | Ageless Management in the AI Era 130
employees. Note that, as the refutation condition of Proposition 10 specifies, if the protection
requirements of the two strata are shown to differ structurally, this symmetric design turns
out to be an oversimplification. Symmetry is a hypothesis that simplifies design, not a
doctrine.
The primacy of educational purpose in youth participation needs operational criteria for
judgment. This paper proposes three. First, documentation of learning objectives. For each
PBL project, the learning objectives of the participating youth are defined in advance, and
their attainment is assessed by the educational institution. Second, subordination of
deliverable use. The firm's acquisition of deliverables is positioned as a byproduct of the
learning process, and the commercial value of deliverables does not govern project design.
The moment the initiative of design shifts to "what should we have them make so that it
sells," the primacy of educational purpose is lost. Third, educational design of AI use. As the
evidence in Section 4.3 shows, AI use without guardrails can raise assisted performance while
impairing learning. "Youth participation that delivers results fast with AI," convenient for the
firm, can be the most dangerous design from the standpoint of education. A connection that
does not explicitly manage this tension does not satisfy the requirement of the primacy of
educational purpose.
For the participation of socially marginalized groups as well, the application of the
symmetric principle should be made concrete. Insider problem perception (Table 5) is a
cognitive asset for the ecosystem and should not be procured as unpaid "cooperation with
interviews." There are three design requirements. First, compensation. If the provision of
insider perception is used to improve products or businesses, it is professional service and
the object of commensurate compensation. Second, attribution. Contributions to ideas and
improvements originating in insiders' perception should be recorded, and attribution to the
person made explicit. Third, consideration for load. Repeatedly narrating the experience of
one's own hardship itself carries cognitive and emotional load. Placing the choice of
frequency and format of participation on the person's side is the insider version of the load
management of Section 8.1. "Insider participation" lacking these is mere extraction wearing
the appearance of inclusion.
The adequacy of compensation faces one practical difficulty: for the task of experiential
audit, no established labor market or reference price yet exists. The absence of a market price
lends itself easily to the pretext for markdown that "there is no going rate, so an honorarium
will do." For the time being, the design can only hold, as self-imposed standards, the
application by analogy of the compensation levels of adjacent professional services (audit,
advisory, expert committee membership, and the like), ex-ante disclosure of per-task
compensation, and periodic review of compensation levels. Note that if Proposition 1 (the
marginal value shift) is correct, referenceable market prices should form as demand for
evaluation- and audit-type tasks increases, and whether that formation occurs is itself one of
the observable implications of Proposition 1.
Working Paper | Ageless Management in the AI Era 131
Governance of the over-flagging rate — suspension and recalibration of audit
authority. The design of the anti-exploitation side needs, as its pair, a bulwark against the
failure mode running in the opposite direction. As the compensation and status of audit-type
roles become established, an incentive can arise to pile up the volume of flags as proof of
contribution. If the response criterion drifts toward "suspect everything," false alarms inflate
verification costs and alarm fatigue, and organizational decision-making heads toward
gridlock through over-auditing. Moreover, if collusion arises in which multiple senior
auditors endorse one another's flags — logrolling — over-auditing becomes fixed as a group
equilibrium beyond individual tendency. The division of labor that does not assign this
control to ex-post damages claims was set out in Section 7.2: liability doctrine handles the
allocation of catastrophic risk, and everyday discipline is handled by operational governance.
That is, the standard operation incorporates a procedure of continuously monitoring each
auditor's false-alarm rate — the estimated bias of the response criterion c in signal detection
theory (Section 9; Appendix A) — temporarily suspending the audit authority of an auditor
who exceeds a pre-set and published threshold, and reinstating that auditor upon passing a
recalibration task in which known errors and sound outputs are mixed in under blinding.
Basing the judgment on performance in blinded tasks rather than on mutual evaluation
among auditors is a bulwark against collusive mutual endorsement contaminating the
calibration judgment. Likewise, ex-ante disclosure of thresholds and procedures, and basing
the judgment solely on calibration indicators rather than on the content of the flags, are
bulwarks against suspension being diverted into the suppression of inconvenient flags — the
erosion of the voice channel that organizational condition (ii) of Proposition 5 protects. In
addition, a reputation mechanism based on the recorded history of detections, false alarms,
and recalibrations supplies everyday discipline against faulty auditing without recourse to
damages claims that are difficult to prove (Section 7.2).
A word is also in order on who implements protection. Individual participants, above all
non-employment individuals, cannot be expected to negotiate load ceilings or the adequacy of
compensation with firms — precisely because of asymmetric bargaining power. Here
intermediary organizations such as NPOs, educational institutions, and employment agencies
(Section 6.4) can perform the bargaining-power-correcting functions of collective negotiation,
the development of standard contracts, and the provision of grievance channels. Correction
by private ordering, however, is no substitute for institutional protection. The protection gap
that Proposition 9 identifies should ultimately be filled by institutional design (such as the
extension of freelance-protection legislation seen in Section 7), and the guidelines of this
section are positioned as the organization-side self-imposed standards to hold until
institutions catch up, and to lay on top of them.
Working Paper | Ageless Management in the AI Era 132
8.3 Connection to BCM's Three Bs — Brain Safety Is the Lifelong Extension of
Base
The BCM of this series (Kadowaki 2026e) formalized a firm's brain capital as stock (K) ×
utilization rate (u), and organized the sequential conditions for individual cognitive ability to
become firm capability into three constraints — Belonging, Base, and Build: that the ability
exists and is maintained (failure mode: decay), that it is available at the moment of work
(failure mode: depletion), and that it does not remain inside the head but is transmitted to the
organization (failure mode: silence). The core discipline of BCM lay in the distinction between
states and measures — the point that what links to value is not expenditure on measures but
the measured state of the workforce.
Brain Safety has a precise location within this framework: Brain Safety is the lifelong
extension of Base. BCM's Base formalized, for the workforce within employment, the claim
that resolving outstanding economic claims on attention (economic hardship and liquidity
constraints) releases cognitive bandwidth. This paper extends this Base in two directions. The
first is the extension of scope. BCM's unit was the employed workforce, but the participants of
a multigenerational ecosystem extend beyond the boundary of employment. The economic
foundation of non-employment participants (adequate compensation and the provision of
accident coverage, Section 8.2) is nothing other than Base extended beyond employment. The
second is the extension of the sources of claims. Claims on attention do not come from
economic hardship alone. The cognitive load, alarms, and brain fatigue emitted by the audittype
role itself are a standing claim on attentional resources, and the load design of Section
8.1 is an intervention that keeps this claim within budget. In Ageless Management, which
presupposes lifelong participation, the protection of Base extended in these two directions is
a precondition for restraining the depreciation of participants' brain-capital stock (K) and
maintaining the utilization rate (u) (Proposition 8).
The phrase "lifelong extension" also carries a temporal implication. BCM's Base addressed
the current-period cognitive bandwidth of the workforce during employment. Participation
in Ageless Management is not a single period of employment but a decades-long process of
repeated increases and decreases in participation volume, interruptions, and re-entries. On
this time axis, the protection of Base is not limited to load management during participation
but includes the maintenance of brain capital in transition periods — the shrinking of roles,
temporary exit, and preparation for re-entry. The evidence since Mental Retirement, seen in
Section 3, suggests that the sharp drop in cognitive engagement accompanying exit is the
principal risk of this transition period, and the gradual adjustment of participation volume
that an ecosystem can offer (gradual decrease and increase of session counts) is a design
alternative to the "all-or-nothing" exit of the employment system. Whether this gradual
adjustment actually restrains the depreciation of K is, however, precisely the object of testing
of Hypothesis H1.
Working Paper | Ageless Management in the AI Era 133
The extension does not stop at Base. In the dimension of Belonging, experiential audit
capacity becomes oversight value only when a detected error reaches others as voice and is
heard. In WP8's terms, detection leads to the correction of error only through reporting and
acceptance. In an organization where an auditor's dissent is discounted on grounds of status,
age, or contract form, assembling the decorrelation portfolio (Section 6.3) realizes no value.
The failure mode of silence is most likely to arise precisely among participants outside the
employment boundary who are in weak positions. In the dimension of Build, as stated at the
end of Section 8.1, the design of unassisted exercise opportunities prevents the depreciation
of experiential audit capacity. Applying the protected unassisted practice that BCM proposed
not only to youth in the course of learning but to the Gc-type roles of super-seniors is an
extension newly claimed by this paper, and its effectiveness is untested (Section 9). BCM's
distinction between states and measures likewise applies to Brain Safety as it stands: Brain
Safety is not to be deemed achieved by the introduction of a list of measures, but is to be
evaluated by the measured states of participants' load, fatigue, and engagement.
Finally, the relation between Brain Safety and the brain-health hypothesis (H1) should be
made clear. As seen in Section 3, the evidence on the brain-health effects of work is mixed,
and Proposition 7 qualified the claim: if an effect exists, it is mediated not by the length of
working hours but by the maintenance of cognitive engagement. Brain Safety is the designside
counterpart of this qualification. That is, this paper does not claim that "participation is
good for the brain." Overloaded auditing, chronic alarm fatigue, and participation under
exploitative terms supply not cognitive engagement but chronic stress and attrition, and can
instead depreciate brain capital. What Hypothesis H1 takes as its object of testing is
engagement in roles that are Gc-exercising and AI-complemented, and its formalization
implicitly includes the conditions of this section — load managed within budget and the
absence of exploitation. The expansion of participation lacking Brain Safety is a path that
worsens participants' brain capital under the banner of Ageless Management, and
Proposition 8 promises nothing about the consequences of such a design. What the refutation
condition of Proposition 8 asks is whether, under a design judged to satisfy the conditions
prior to the observation of outcomes, a deterioration of brain-capital indicators is observed
ex post (Section 5).
8.4 Table 9: The Skeleton of the Brain Safety Guidelines
The foregoing discussion is consolidated into the guideline skeleton of Table 9. Each row
shows the domain, the design principle, examples of concrete measures, and the
corresponding propositions of this paper.
Table 9 The skeleton of the Brain Safety guidelines
Domain Principle Concrete measures (examples)
Corresponding
propositions
Do not exceed the
budget constraint on the
Time ceiling per audit session
(roughly 20–30 minutes as a guide)
Propositions 7 and
8 (WP8's solvency
Working Paper | Ageless Management in the AI Era 134
Cognitive-load
management
individual's attentional
resources (the
individual version of the
solvency condition)
with rest; per-session
commissioning of tasks;
management of the volume of items
to check and of alarms
condition and
alarm fatigue)
Consent-based
load adjustment
(informed
consent)
Protection is activated
not by unilateral forced
cutoff but by consentbased
dynamic contract
renewal
Presentation of risk alerts based on
objective data; renewal of contract
terms involving the person and an
intermediary organization; an
objective second-opinion procedure
by a third-party body (occupational
physician, external intermediary
organization) when agreement fails
to form; an order of measures that
puts reassignment to low-risk
domains first and termination of
participation as the last resort
Propositions 7, 9,
and 10 (managing
the conflict with
self-determination)
Workenvironment
design
Distribute the modes of
load and make work
cognitively accessible
Combined use of Voice UI and readaloud;
interactive auditing through
AI summaries and question-andanswer;
reduction of the cognitive
load of displays and documents
Propositions 2 and
11
Skill
maintenance
Design the ceiling on
load and the floor on
exercise simultaneously
Regular incorporation of protected
unassisted practice; avoidance of
full reliance on AI pre-screening;
assessment of ability under both
assisted and unassisted conditions
Propositions 3 and
8 (extension of
BCM's Build)
Sampling design
for the raw-audit
slot (dual-track
auditing)
Risk-weighted dynamic
sampling — stratify by
impact while
maintaining a purely
random floor (minimum
sampling rate)
(independence of the
audit channel; the floor
is insurance against
fallibility and an
exploration slot for
unknown unknowns)
Design of stratum-specific sampling
rates by impact (irreversibility,
scope of propagation, monetary
scale) with random sampling within
strata; the discipline of not using
the AI's confidence as grounds for
lowering sampling rates;
application of the random-sampling
floor to the entire output domain
(exploratory sampling for novel
modes of breakdown outside the
impact weighting); ex-ante
disclosure of sampling rates and the
floor and disclosure of the slipthrough
probability ((1−p)×β);
measurement of detection
performance by channel (Section 9)
Proposition 5
(WP8's
decorrelation
condition and
solvency condition;
the independence
version of protected
unassisted practice)
Forced queue Insulate assignments
within the raw-audit
track from auditors'
preferences (prevention
of auditor-side selection
bias)
Delivery in a random forced order
that cannot be chosen; calibration
of volume to what can be digested
within the session time ceiling (20–
30 minutes); monitoring of deferral
and tail truncation; handling of
Proposition 5
(enforcement
mechanism of
random sampling)
Working Paper | Ageless Management in the AI Era 135
hard-to-read outputs by session
splitting
Independent
double-checking
(exclusion of
solo auditing)
Do not make the audit
dependent on the
judgment of a single
auditor (a bulwark
against cognitive holdup)
Independent audits by multiple
heterogeneous seniors with crosschecking
of flags; exclusion of solo
auditing by a single senior; plural
evaluation of the validity of flags
(Section 6.3)
Proposition 5
(application of
Section 6.3)
Periodic blinded
evaluation of
audit validity
Base the evaluation of
auditors on the validity
of flags, not their
volume
Blinded mixing-in of known errors
and unproblematic outputs;
periodic measurement of the
detection rate and the over-flagging
rate (operational version of H3
dependent variable (iv)); feedback
of results to the person
Propositions 3, 4,
and 12 (backtesting
on real work logs)
Over-flagging
governance
(suspension and
recalibration of
audit authority)
Control the moral
hazard of over-flagging
by operational
governance, not by
damages claims
(division of labor with
the cap on liability of
Section 7.2)
Continuous monitoring of each
auditor's false-alarm rate
(estimated bias of the response
criterion c); temporary suspension
of audit authority upon exceeding
pre-set and published thresholds;
recalibration by blinded tasks with
reinstatement judgment; a
reputation mechanism based on the
history of detections, false alarms,
and recalibrations; non-reliance on
mutual evaluation as a bulwark
against collusion (logrolling)
Propositions 5 and
12 (monitoring of
the SDT response
criterion c —
Section 9; Appendix
A)
Operational
restriction of
unannounced
verification
(anonymized
system
calibration)
Use unannounced insitu
verification, in the
first instance, for
anonymized and
aggregated system-level
calibration, and do not
divert it to individual
surveillance (exclusion
of the panopticon effect
— consistency with the
protection of cognitive
bandwidth)
System calibration through
anonymized and aggregated
results; use for individual
treatment limited to long-window
aggregation with due process
(disclosure to the person and an
opportunity to respond);
prohibition of demotion or
revocation of authority based on a
single result; gaming deterrence
through publicizing the fact that
calibration is performed (Section 5,
commentary on Proposition 12)
Propositions 7 and
12 (separation of
gaming deterrence
from surveillance
pressure)
Layer separation
of feedback and
diversity
preservation at
the RAG layer
Separate by layer the
management of the
audit's interpretive
context from the
training of the base
model, and preserve
judgment diversity at
the RAG layer
Management of interpretive
context (context metadata) at the
prompt and evaluation-criteria
layers; avoidance of full-volume
input into base-model retraining;
selection and recording of what is
fed into training; temporal decay of
logs at the RAG layer and sourcediversity
weighting
Proposition 5
(model condition
(iii); Shumailov et
al. 2024)
Working Paper | Ageless Management in the AI Era 136
(preservation of
generative diversity)
Compensation
and contracts
Adequacy of
compensation; do not
substitute honorific
titles or a sense of
purpose for
compensation
Written contracts with explicit
compensation; market-referenced
remuneration for experiential
audit; contractual provision of
accident compensation and
consultation channels for nonemployment
participants
Propositions 9 and
10
Youth
participation
Primacy of educational
purpose
Priority of schooling; restrictions on
the commercial use of deliverables;
involvement of educational
institutions; guardrail design for AI
use
Proposition 10
Voice channels
(flattening the
power gradient)
Audit dissent reaches
decision-making
regardless of status, age,
or contract form
(organizational
condition (ii) of
Proposition 5)
Recording of dissent and detected
items with escalation paths; a duty
to respond to audit opinions
(recording the reasons for
rejection); operations that exclude
the speaker's attributes and
contract form from the evaluation
of flags (Section 6.3)
Proposition 5
(extension of BCM's
Belonging)
Measurement Measure states, not the
introduction of
measures
Continuous measurement of load,
fatigue, and cognitive engagement
with threshold-based operation;
brain-capital indicators by
participation stratum (Section 9)
Propositions 7 and
8 (BCM's distinction
between states and
measures)
Note: This table is a skeleton, not a regulatory standard or a finished code. The effectiveness of each concrete measure
(its effects on audit quality, brain health, and the prevention of exploitation) is untested; the cognitive-load side
connects to Hypotheses H1 and H3, and the economic-protection side to the institutional comparison of Section 7 and
the testing frameworks of Propositions 9 and 10. Numerical values such as the time ceiling are initial design values
inherited from the practical guides of the original draft, not empirically grounded thresholds. Further, that each item
of this table is operated in a form whose satisfaction can be judged prior to the observation of outcomes is a
requirement for the testability of Propositions 8 and 11 (the refutation conditions of Section 5; the measurement
framework of Section 9).
The order of implementation also inherits BCM's discipline. BCM derived the conclusions
that the three constraints operate multiplicatively, that the weakest link governs the whole,
and that the foundation comes first and development comes after. The ecosystem version has
the same form. Introducing sophisticated load measurement while the adequacy of
compensation and contracts (the extended Base) is lacking only means that participants are
measured under exploitative terms; assembling a decorrelation portfolio while voice
channels (the extended Belonging) are lacking means that detection ends in silence. The
domains of Table 9 are not a parallel checklist: they have an order in which the foundational
domains of compensation-and-contracts and cognitive-load management come first, and the
refinement of skill maintenance and measurement rests on top of them.
Working Paper | Ageless Management in the AI Era 137
Brain Safety is not an ancillary welfare benefit of Ageless Management. Following the logic
of Sections 4 and 6, it is the very supply condition of oversight value. Without budget
management of cognitive load, detection probability collapses (Section 8.1); without the
adequacy of compensation, the ecosystem degenerates into exploitation and the sustainability
of participation collapses (Section 8.2); without the design of exercise opportunities,
experiential audit capacity itself depreciates (Section 8.3). It is in this sense that Proposition 8
asserts bidirectional capital formation only for multigenerational ecosystems that satisfy exante
verifiable conditions — the connection modes of Definition 7 and the satisfaction of
Definition 8. The framework for measuring and verifying the satisfaction of these conditions
is the subject of the next section.
Regarding the operation of this conditional claim, a normative statement and an epistemic
statement should be clearly separated. At the normative level, this paper claims the following:
an organization that cannot meet the standards of Brain Safety has no standing to claim the
benefits of Ageless Management. This is a statement about the legitimacy of an organization
attaching this paper's name to its own practice. The epistemic level — the judgment of
condition satisfaction in testing Propositions 8 and 11 — is a different matter. Condition
satisfaction is judged by an ex-ante checklist preceding the observation of outcomes, that is,
by the satisfaction of the items of Table 9 operated in a form that a third party can judge
before seeing the results. If, under a design judged ex ante to satisfy the conditions, a
systematic ex-post deterioration of the brain-capital indicators of any participating stratum is
observed, Propositions 8 and 11 are rejected. Retroactive denial of the conditions from expost
outcomes — claiming that "the conditions were not satisfied after all" — is not admitted
as a defense of the propositions (see the refutation conditions of Propositions 8 and 11). The
normative statement disciplines organizations' claims; the epistemic statement disciplines
this paper's theory. An operation that conflates the two and reclassifies every failure case as a
deficiency of Brain Safety is a maneuver that immunizes the theory against refutation, and is
not a correct application of this paper's framework. That this paper asserts its propositions
only conditionally is not rhetoric: the conditions — the design standards of this section — are
the substance of implementation, and at the same time the entrance to verification.
9. Measurement and Verification
The definitions and propositions of Section 5 were developed into design theory, institutional
analysis, and safety and health in Sections 6 through 8. Most of this paper's propositions,
however, remain untested, and a theory stands as a scientific proposal only when it is
equipped with a map of verification. This section draws that map. Section 9.1 organizes three
hypotheses (H1–H3) together with their verification designs, identification threats, and costs
(Table 8), and Section 9.2 presents the design of H3, the highest-priority and lowest-cost
experiment (Figure 5). Section 9.3 extends the K × u framework of Brain Capital Management
(Kadowaki 2026e) to organization-level indicators for the multigenerational ecosystem, and
Working Paper | Ageless Management in the AI Era 138
Section 9.4 sets out a roadmap for implementation. Two principles run through this section.
First, begin with the tests that are lowest in cost and that strike directly at the theory's core
mechanisms. Second, measurement is an instrument for role allocation and protection, not
an instrument of selection and exclusion. The risk of the latter misuse is flagged at various
points in this section and confronted head-on in Section 10.
9.1 The Map of Verification — Three Hypotheses
This paper's twelve propositions were each presented with a refutation condition attached
(Section 5). But enumerating refutation conditions and having an executable verification plan
are two different things. This section bundles the core of the propositions into three testable
hypotheses and specifies a verification design for each. H1 targets the individual-health side
of Proposition 7 (the cognitive-engagement pathway) and Proposition 8 (bidirectional capital
formation); H2 targets the organizational-outcome side of Proposition 5 (generational
decorrelation) and Proposition 6 (the AI-mediated diversity effect); and H3 targets the
theory's core mechanism, Proposition 3 (the non-compressibility of experience) and
Proposition 4 (the formation of experiential audit capacity).
Hypothesis H1 (Brain-Health Hypothesis)
Super-seniors engaged in Gc-exercising, AI-complemented roles show a lower rate of
cognitive decline (MoCA, etc.) than same-age peers who have exited work, mediated by
the maintenance of cognitive engagement. Identification strategy: use the exogenous
variation of pension reforms and mandatory-retirement rules as instrumental variables,
or a quasi-experiment exploiting the exogeneity of reasons for working. Explicitly
address health selection and reverse causation. Rigorously control for years elapsed
since leaving work (time since retirement) as a covariate — in order to distinguish the
effect of working from the cumulative natural depreciation over the period out of work.
H1 is the most expensive to test and the hardest to identify of this paper's hypotheses. As
confirmed in Section 3, the association between work and cognitive function is doubly
contaminated by health selection (the healthy worker effect) and reverse causation (the
prodromal phase of cognitive decline hastens retirement), and, moreover, conclusions can
reverse depending on the choice of instrumental variable. Estimates using international
differences in pension and tax institutions as instruments suggested a large negative effect of
retirement (Rohwedder & Willis 2010), whereas estimates using employer-offered earlyretirement
incentive windows as instruments did not accept the association as causal and, for
blue-collar workers, instead reported a positive relationship between time in retirement and
cognition (Coe et al. 2012). A systematic review summarizes the evidence as mixed (Meng et
al. 2017), and the average reading of recent causal-inference research likewise goes no
further than "most estimates find that cognitive skills decline after retirement, but the effects
are highly heterogeneous by occupation and by the voluntariness of retirement" (van Ours
Working Paper | Ageless Management in the AI Era 139
2022). Testing H1 must therefore not depend on a single instrumental variable; it must
combine multiple identification strategies — the exogenous variation of institutional reforms
and the exogeneity of reasons for working — in a design that reports the discrepancies
among estimates themselves. H1 furthermore contains a mediation hypothesis. To test the
structure of Proposition 7 — that the mediator is not working hours themselves but the
maintenance of cognitive engagement through occupation in Gc-type roles — measurement
not only of whether one works but of the quality of work (the share of Gc-type components in
the role, the presence or absence of AI complementation) is indispensable. On this point, the
data of existing retirement studies cannot substitute, and the longitudinal study and medical
collaboration of Section 9.4 must be awaited.
Hypothesis H2 (Multigenerational Team Outcome Hypothesis)
Under AI use, mixed teams of "experiential audit (Gc) × youthful prototyping (Gf) × midcareer
orchestration" show significantly higher scores than homogeneous teams on both
the novelty and the feasibility of the ideas produced. Blinded external evaluation. In
team assignment, other demographic attributes such as gender, ethnicity, and cultural
background are handled by stratified randomization or covariate control, identifying
the effect of age-and-experience heterogeneity apart from other diversity dimensions.
Analysis plan: the significance of the interaction alone is not counted as support for
Proposition 6; a decomposition of simple main effects is used to confirm an absolute
improvement in the mixed teams' scores (excluding the false positive in which the
interaction arises solely from the deterioration of homogeneous teams through AI
overtrust). Time to decision, verification effort, and the human-to-human
communication time required to reach agreement (time-to-consensus) are recorded as
cost variables, and outcomes are reported as net benefit after deducting them. A positive
net benefit is a requirement for support; superiority in idea quality alone is not counted
as support.
Testing H2 presupposes an honest recognition of the empirical baseline. As confirmed in
Section 6, the average relation between age diversity and team outcomes is near zero (Joshi &
Roh 2009; Schneid et al. 2016; Wallrich et al. 2024), and positive effects appear only under the
conditions of task complexity and creativity and an inclusive climate (Backes-Gellner & Veen
2013; Wegge et al. 2012). H2 is therefore not a claim of average effect — that
multigenerational composition raises outcomes — but a test of Proposition 6's moderator
structure: the effect turns positive only in the presence of AI-mediated complementarity. The
ideal design is a 2 × 2 team-level randomization crossing team composition (mixed/
homogeneous) with AI use (with/without), testing Proposition 6 up to and including the
disappearance or reversal of the effect in mixed teams without AI use. The composition of the
homogeneous control teams is stipulated here. The primary control is a homogeneous team
composed solely of the mid-career generation; where feasible, a multi-arm comparison adds
Working Paper | Ageless Management in the AI Era 140
homogeneous conditions of juniors only and seniors only — because which generation the
homogeneous teams are drawn from changes both the meaning of the hypothesis and the
number of teams required. In assignment to teams, other demographic attributes such as
gender, ethnicity, and cultural background are handled by stratified randomization (or
covariate control where cell sizes are insufficient) — because if the mixed/homogeneous
contrast is confounded with diversity dimensions other than age-and-experience
heterogeneity, the observed effect cannot be attributed to the axis Proposition 6 asserts; the
outline of the assignment procedure is supplemented in Appendix A (A.6). The analysis plan
additionally specifies a decomposition of simple main effects to identify the source of the
interaction's sign. An apparent interaction produced by AI mediation lowering the
performance of homogeneous teams and an interaction produced by raising the performance
of mixed teams are not equivalent as support for Proposition 6. The task is a creative
divergent task such as drafting new-business proposals, and evaluation is by blinded external
evaluators — with team composition and participant attributes concealed — who score
novelty and feasibility independently. The procedure of blinded third-party scoring follows
the standard method of generative-AI experiments (Noy & Zhang 2023; Dell'Acqua et al. 2023).
The identification threats are the endogeneity of team formation (addressed by
randomization), insufficient statistical power from small team numbers (the team is the unit
of analysis, so the required sample is large), behavioral change from participants sensing the
experimental intent, and breach of blinding when evaluators infer team composition from
style or the like. In light of the meta-analysis finding that human-AI combinations on average
fall below the best single agent (Vaccaro, Almaatouq & Malone 2024), there is no guarantee
that the AI-use condition automatically raises outcomes. Moreover, cost-side measurement is
built into the hypothesis's very criterion of support. Generationally heterogeneous teams may
require longer than homogeneous teams to reach agreement, in order to process and
coordinate the points raised, and unless this cost is booked, the outcome that Proposition 6
defined as net benefit (Section 5) is not measured. H2 therefore records, in addition to time to
decision and verification effort, the human-to-human communication time required to reach
agreement (time-to-consensus) as cost variables, and makes a positive net benefit after their
deduction a requirement for support — if only superiority in idea quality is observed and net
benefit after cost deduction is not positive, H2 is not counted as supported. H2 is a design that
can produce results unfavorable to this paper's theory — no mixed-team advantage, no AImediation
effect, a quality advantage eaten up by costs — and it is for that reason that it
deserves the name of a test.
Working Paper | Ageless Management in the AI Era 141
Hypothesis H3 (Direct Test of the Non-Compressibility of Experience)
A 2 × 2 factorial design [experience level: long-term domain experts vs. juniors] × [AI
assistance: with vs. without]. The task is the verification of AI-generated business
proposals and documents with misinformation and contextual risks embedded. So that
the embedded errors do not depend on the experimenters' preconceptions, tasks blindly
generated from an empirical failure-case dataset of accidents, scandals, and failures that
actually occurred are included. Some of the errors are constructed in both a form
conforming to auditors' industry received views (confirmation-conforming) and a form
departing from them (deviating), simultaneously testing boundary condition (i) of
Proposition 4. The fluency of the task documents (high fluency = stylistically polished
output vs. low fluency) is controlled as a factor or covariate, simultaneously testing
boundary condition (iii). Dependent variables: (i) detection rate of surface errors; (ii)
detection rate of contextual and practical risks; (iii) quality of proposed corrections
(blinded evaluation); (iv) over-rejection rate — the rate of erroneously rejecting the
embedded "correct but convention-defying innovative proposals" (measuring the
boundary at which experience turns into over-auditing). Years elapsed since leaving
work are recorded and controlled. Predictions: on (i), the experience gap narrows with
AI assistance (consistent with prior research), but on (ii) and (iii) a main effect of
experience persists and does not narrow even under AI assistance (a direct test of
Proposition 3). No prediction is placed on (iv); it is an exploratory indicator. Appendix A
gives the detailed experimental protocol.
H3 has the highest priority of the three hypotheses. There are three reasons. First, its cost is
the lowest. It is a task experiment with the individual as the unit of analysis, executable
online, requiring neither longitudinal tracking nor team-level randomization. Second, it
strikes directly at the theory's core. The claim that experiential audit capacity (Definition 4) is
real and is not compressed by AI assistance is the keystone of this paper's entire construction
— the supply-side theory of oversight, the design of the multigenerational ecosystem, the
conversion of social problems into resources — and if H3 is rejected, the theory requires
revision from its foundations. Third, it is prior to the interpretation of H1 and H2. If the
existence of experiential audit capacity cannot be confirmed, H2's mixed-team design loses its
basis, and the definition of H1's treatment — a Gc-exercising role — becomes ambiguous.
Verification should proceed from the cheap and decisive, and H3 is the only hypothesis that
meets that condition. The details of the design are given in the next section.
Table 8 The map of verification — verification designs and identification threats for the three
hypotheses
Hypothesis Verification design
Dependent
variables
Identification
threats
Cost and order
of execution
Working Paper | Ageless Management in the AI Era 142
H1 (brain-health
hypothesis)
Corresponds to
Propositions 7 and
8
Quasi-experiment.
Panel analysis using
the exogenous
variation of pension
reforms and
mandatoryretirement
rules as
instrumental
variables, or
comparison
exploiting the
exogeneity of
reasons for working.
Mediation analysis
used in combination
Rate of cognitive
decline (MoCA, etc.).
Cognitive
engagement as
mediating variable.
The share of Gc-type
components in the
role and the
presence or absence
of AI
complementation
included in the
definition of
treatment
Health selection
(healthy worker
effect), reverse
causation,
voluntariness of
retirement. Prior
examples of
conclusions
reversing with
the choice of
instrumental
variable
(Rohwedder &
Willis 2010 vs.
Coe et al. 2012).
Measurement
error in the
mediating
variable
High (requires
years of
longitudinal
tracking and
medical
collaboration).
Third in order
— conducted in
the third and
fourth stages of
Section 9.4
H2
(multigenerational
team outcome
hypothesis)
Corresponds to
Propositions 5 and
6
Team-level
randomized
experiment. 2 × 2 of
team composition
(mixed/
homogeneous) × AI
use (with/without).
In team assignment,
other demographic
attributes such as
gender, ethnicity,
and cultural
background handled
by stratified
randomization or
covariate control.
Creative divergent
task. Blinded
external evaluation.
Decomposition of
simple main effects
specified in the
analysis plan
(excluding spurious
interactions arising
solely from the
deterioration of
homogeneous
teams)
Blinded scores for
the novelty and
feasibility of the
ideas produced. Time
to decision,
verification effort,
and communication
time to reach
agreement (time-toconsensus)
recorded
as cost variables,
with outcomes
reported as net
benefit after their
deduction. A positive
net benefit is a
requirement for
support. Detectionoverlap
rate as
auxiliary measure
(Section 9.3)
Insufficient
power from the
number of teams.
Behavioral
change from
experiment
participation.
Breach of
blinding
(inferring
composition from
style, etc.).
Artificiality of the
task. The
empirical
baseline of a
near-zero average
effect.
Misattribution of
the source of the
interaction's sign
(addressed by
decomposition of
simple main
effects)
Medium
(requires
simultaneous
execution at the
scale of dozens
of teams).
Second in order
— designed in
light of H3's
results
H3 (noncompressibility
of
experience)
Corresponds to
Individual-level 2 ×
2 factorial
experiment.
Experience level ×
(i) Detection rate of
surface errors; (ii)
detection rate of
contextual and
Confounding of
experience with
chronological age
(addressed by
Low (individual
task, executable
in weeks). First
in order —
Working Paper | Ageless Management in the AI Era 143
Propositions 3 and
4
AI assistance.
Verification task on
AI-generated
documents with
embedded errors.
Embedded errors
include tasks blindly
generated from an
empirical failurecase
dataset,
constructed in both
confirmationconforming
and
deviating forms
(simultaneous test of
boundary condition
(i) of Proposition 4).
Fluency of the task
documents (high/
low fluency)
controlled as a
factor or covariate
(simultaneous test of
boundary condition
(iii)). Years since
leaving work
recorded and
controlled.
Executable online
practical risks; (iii)
quality of proposed
corrections (blinded
evaluation); (iv)
over-rejection rate
(exploratory
indicator)
recruiting highage
lowexperience
cells,
etc.). Construct
validity of the
embedded errors
(experimenter
bias mitigated by
blind generation
from the failurecase
dataset).
Artificiality of the
task.
Measurement
error in domain
knowledge
highest priority,
lowest cost
Note: Cost levels are relative assessments. Each hypothesis is designed so that an unsupportive result connects directly
to the refutation condition of the corresponding proposition (Section 5). For details of H1's identification threats see
Section 3.2; for H2's empirical baseline see Section 6.
The three rows of Table 8 are not independent items of verification; they carry a cascade
structure of rejection. If H3 is rejected, H2's mixed-team design and H1's definition of
treatment, both of which depend on the reality of experiential audit capacity, require
reconstruction, and this paper's theory is revised from its core. If H3 is supported and H2 is
rejected, experiential audit capacity is real at the individual level, but the organizational
exploitation of it through team composition (the design theory of Section 6) is in error. If H3
and H2 are supported and H1 is rejected, Ageless Management survives as a theory of
organizational outcomes but must abandon the claim of spillover to brain health (the
individual side of Propositions 7 and 8). Specifying in advance which combination of results
kills which part of the theory — that is what the phrase "the map of verification" means.
9.2 The Experimental Design of H3
The H3 experiment is a 2 × 2 factorial design (Figure 5). The first factor is experience level,
comparing practitioners with long-term experience in the target domain against juniors with
shallow experience in the same domain. The second factor is AI assistance, with one
Working Paper | Ageless Management in the AI Era 144
condition permitting and one condition not permitting the use of generative AI while
performing the verification task. The task is the verification of business proposals and
practical documents for the domain, produced with generative AI, into which the
experimenters have systematically embedded two types of error. The first type is surface
errors — numerical inconsistencies, missing steps, formal breakdowns — detectable by
careful cross-checking even without domain knowledge. The second type is contextual and
practical risks — reproductions of past failure patterns, collisions with regulation and
commercial practice, missing consideration for stakeholders, ethical risks — corresponding to
what Definition 4 specified as the detection targets of experiential audit capacity. The design
of measuring AI users' performance on tasks with deliberately embedded errors adapts the
outside-the-frontier task method of Dell'Acqua et al. (2023) to the measurement of
experiential audit capacity. Two disciplines are imposed on the construction of the embedded
errors. First, so that the errors do not become reflections of the experimenters'
preconceptions, tasks blindly generated — by producers ignorant of the hypotheses — from
an empirical failure-case dataset of accidents, scandals, and failures that actually occurred
are included. Second, some of the errors are constructed in both a form conforming to
auditors' industry received views (confirmation-conforming) and a form departing from
them (deviating) — in order to test simultaneously, in the same experiment, boundary
condition (i) of Proposition 4: that experience-derived priors aid detection for conventiondeviating
errors, whereas for convention-conforming errors detection can instead be
impeded by the synergy of confirmation bias and automation bias (Section 5). In addition,
fluency is controlled as a property of the task documents themselves. Two versions of
identical content, manipulating only stylistic polish — a high-fluency and a low-fluency
version — are prepared, and fluency is treated as a factor or covariate — in order to test
simultaneously, in the same experiment as the main test, boundary condition (iii) of
Proposition 4: that high-fluency output can, through the effect of processing fluency, raise the
threshold of auditors' cognitive sense of unease and impede detection without any deficit in
the auditor's capacity (the manipulation procedure and manipulation checks are in Appendix
A).
There are four dependent variables. (i) The detection rate of surface errors; (ii) the
detection rate of contextual and practical risks; (iii) the quality of the proposed corrections
for the problems detected, scored by blinded evaluators from whom participant attributes
and conditions are concealed. (iv) is the over-rejection rate — the rate at which the "correct
but convention-defying innovative proposals" embedded in the task were erroneously
rejected as problems — measuring the boundary at which experience turns into overauditing
(excessive risk aversion). The predictions are asymmetric. On (i), AI assistance
narrows the experience gap — this is the direction consistent with the compression evidence
organized in Section 4.1, and this paper's theory in fact predicts compression here. On (ii) and
(iii), a main effect of experience persists and does not narrow even under AI assistance. This
is the direct test of Proposition 3. If, under the AI-assisted condition, the gap in contextual
Working Paper | Ageless Management in the AI Era 145
detection rates between long-term experts and juniors disappears or reverses, Proposition 3
is rejected exactly as its refutation condition states, and this paper's theory, resting on
experiential audit capacity, loses its core. Conversely, if the main effect of experience persists
on (ii) and (iii), that is the first direct evidence supporting the existence of experiential audit
capacity. No directional prediction is placed on (iv); it is an exploratory indicator — whether
over-rejection increases with experience, and whether it is distributed as the flip side of
confirmation-conforming misses, is information that demarcates the location of Proposition
4's boundary, and it is reported whatever the results.
Three devices are built into this design. The first is the separation of experience from
chronological age. As argued in Section 4.4, this paper's reattribution thesis asserts that "what
predicts detection capacity is experience, not age" (Proposition 4). Recruitment of participants
therefore deliberately breaks the correlation between experience and age — including older
career-changers and returners with shallow experience in the domain, and young early
specializers with long experience — making it possible to estimate the independent effect of
age controlling for years of experience. Proposition 4's refutation condition (that
chronological age still independently predicts detection capacity after controlling for years of
experience) becomes testable only through this design. The second is the explicit
manipulation of AI assistance as a factor. This transcribes into the measurement design the
divergence of Section 4.3 — that assisted performance does not guarantee unassisted
competence (Bastani et al. 2025) — and is also a requirement for bringing onto the test bench
the relation between BCM's (Kadowaki 2026e) "protected unassisted practice" and H3, which
this paper newly asserts and which is untested. The third is the recording and control of years
elapsed since leaving work. The low performance of participants long away from practice
may reflect not the absence of experiential audit capacity but depreciation over the idle
period, and unless these two are distinguished, the experience-level factor is confounded
with the effect of timing of exit. The same reason that H1 demands rigorous control of years
since retirement applies to H3 at the level of individual differences (eligibility criteria and
controls are detailed in Appendix A).
Working Paper | Ageless Management in the AI Era 146
Figure 5 The 2 × 2 factorial design of Hypothesis H3. Experience level (long-term domain experts /
juniors) is crossed with AI assistance (with / without), and a verification task on AI-generated documents
with embedded errors is imposed. The predictions are asymmetric across dependent variables: on the
detection rate of surface errors (i) the experience gap narrows with AI assistance, but on the detection
rate of contextual and practical risks (ii) and the quality of proposed corrections (iii) the main effect of
experience persists even under AI assistance. This asymmetry constitutes the direct test of Proposition 3
(the non-compressibility of experience). The over-rejection rate (iv) is recorded as an exploratory
indicator with no prediction placed on it. Note: The detailed experimental protocol — participant
requirements, task construction, the typology of embedded errors, evaluator blinding, the testing plan,
and an approximate required sample size — is given in Appendix A.
Only the essentials of execution are noted in the main text. The task domain must match
the participants' domain of experience, and execution in a single domain (for example,
business planning in a specific industry) is the first step. Generalization of the results requires
replication in multiple domains. Evaluator blinding includes a robustness check against the
possibility that experience level is inferred from the style of participants' responses. Because
chance hits contaminate performance on detection tasks, the rate of flagging non-error
locations as errors (the false-alarm rate) is recorded in parallel, and sensitivity d′ and
response criterion c are estimated separately within the framework of signal detection theory
(SDT) — because detection rates alone cannot distinguish true sensitivity from a "suspect
everything" response bias (the details of the analysis plan are in Appendix A). These, together
with participant requirements, task construction, the typology of embedded errors, evaluator
blinding, the testing plan, and the approximate required sample size, are detailed in the
experimental protocol of Appendix A. Furthermore, execution requires preregistration of the
hypotheses and testing plan and publication of the task materials, the list of embedded errors,
and the evaluation criteria. The compression evidence this paper relies on includes
preregistered experiments (Noy & Zhang 2023; Dell'Acqua et al. 2023), and in the present case,
where the theory's proposer may be involved in its verification (see the conflict of interest in
H3: 2×2 factorial design (experience level × AI assistance)
Junior × no AI assistance
Baseline
Junior × AI assistance
Surface detection improves (predicted)
Experienced × no AI assistance
Contextual-detection reference
Experienced × AI assistance
Contextual gap persists (predicted)
Experience
level
DVs: (i) surface-error detection (ii) contextual/practical-risk detection (iii) quality of corrections (blind-rated)
(iv) over-rejection rate — errors span confirmatory/anomalous and high/low fluency — prediction: gap shrinks on (i), persists on (ii)(iii); (iv) Working Paper | Ageless Management in the AI Era 147
Section 10.6), making after-the-fact substitution of hypotheses structurally impossible is
demanded even more strongly than in an ordinary experiment.
9.3 Organization-Level Indicators — Extending K × u to the Ecosystem
In parallel with hypothesis testing, an organization implementing Ageless Management needs
indicators for measuring its own state. This paper extends the framework of brain capital =
stock (K) × utilization rate (u), formalized by BCM (Kadowaki 2026e), from the employees of a
single firm to the participants in a multigenerational ecosystem not limited to the
employment boundary (Definition 7). What Proposition 8 asserts is that an ecosystem
satisfying its conditions increases K and u bidirectionally — restrained depreciation of K
among super-seniors, early formation of K among youth and socially marginalized groups,
and a rise in u for the organization. Making this claim measurable requires participant-level
K indicators, organization-level u indicators, and, as the third dimension this paper adds,
indicators of decorrelation. Furthermore, since these indicators become inputs to role
allocation, the ramparts that keep measurement itself robust (the operational requirements
of Proposition 12) must be built into the operation of the indicators.
First, participants' K-maintenance indicators. As a premise, the structure of K should be
restated. As made explicit in the commentary on Proposition 8 (Section 5), K is not a single
number, nor is it an additive composite score of three dimensions — (1) clinical cognitivefunction
scores (standardized tests such as MoCA), (2) structural and functional brain
measures, and (3) standardized domain-knowledge measures. K is conceptualized as a
hierarchical function K = 1[Kbase ≥ θ] × f(Kbase, Kdomain), gated by the foundational
cognitive function Kbase measured by (1) and (2) exceeding a threshold θ, on top of which the
domain knowledge and operational schemata Kdomain of (3) are multiplied (an inheritance
and refinement of BCM's measurement framework). This structure has two implications for
measurement procedure. The dimensions must be measured and reported independently,
and no additive composite may be constructed in which high domain-knowledge scores offset
declines in clinical scores — that would resurrect, on the measurement side, the substitution
the hierarchical function forbids. Also, the gate judgment on Kbase logically precedes the
measurement of the other dimensions — this threshold is the same gate as the cognitivescreening
lower bound of Proposition 2(b) (the theory's internal consistency), and its
judgment is also the judgment of the boundary for allocation to experiential-audit roles. The
"systematic deterioration of brain-capital indicators" in Proposition 8's refutation condition is
adjudicated by dimension-wise measurement in accordance with this hierarchical structure.
Of these, direct measurement of (1) and (2) should be conducted only under medical
collaboration (Section 9.4); what an organization can handle day to day is limited to proxy
indicators for them. There are three candidates. (1) Measurement of cognitive engagement —
the share of Gc-type components in the role occupied, and the subjective absorption in and
challenge level of role occupation. This is the mediating variable of Proposition 7 and also
connects to the testing of H1. (2) Periodic performance on detection tasks under unassisted
Working Paper | Ageless Management in the AI Era 148
conditions — small H3-type tasks administered periodically without assistance, tracking the
level and trajectory of experiential audit capacity. This can function as an early warning
against the skill-depreciation risk discussed in Section 4.3, but it must be stated repeatedly
that whether unassisted practice is itself effective for maintaining capacity is untested. (3)
Continuity of participation and changes of role — tracking whether exits due to the cognitive
bottleneck (Definition 3) are actually decreasing. All of these are apprehensions of state, not
measurements of the effect of interventions, and the distinction between state and
intervention that BCM made a discipline is maintained here as well.
Second, the organization's u indicators. The utilization rate (u) is operationalized as the
share of the cognitive assets the ecosystem could connect that are actually connected to roles.
Concretely: (1) the fit rate between participants' held assets (measured cognitive
characteristics and domain experience) and their current role requirements; (2) the fill rate
of Gc-type and audit-type roles — how far the supply of experiential audit capacity is
connected to oversight demand (the verification function shown in Section 4 to be growing
scarce); and (3) the distribution of participation and remuneration by contractual form —
data to be read alongside the Brain Safety indicators (Section 8, Table 9) to check that nonemployment
participation is not falling into the protection vacuum (Proposition 9). A rise in
u, if it becomes an end in itself, slides into exploitation (Definition 8(ii)). u indicators must
always be reported paired with indicators of load and remuneration.
Third, the measurement of decorrelation. Proposition 5 (generational decorrelation) can be
measured directly within an organization. The core of the proposed measurement is the
detection-overlap rate. The same set of AI outputs is audited independently by multiple
supervisors, and the sets of missed errors are compared across supervisors. The overlap rate
of misses for generationally homogeneous supervisor pairs is contrasted with that for
heterogeneous pairs, and if the latter is systematically lower, Proposition 5's claim — that the
correlation of blind spots is lower between generations than within them — is supported.
Conversely, if there is no difference in overlap rates, Proposition 5 is rejected exactly as its
refutation condition states. This measurement adds the audit channel as a factor — a
condition in which auditors access the raw output unmediated, and a condition in which they
access it through summarization and pre-screening by the same AI. Proposition 5 is
formulated conditional on unmediated access (Section 5), and the claim of its latter half —
that sharing an AI-mediated channel destroys, at the channel level, the decorrelation supplied
by generational heterogeneity — can be directly tested through this design as the difference
in overlap rates between the unmediated and mediated conditions. For the same reason, a
foundation-model condition is also added as a factor — a condition in which the audited AI
outputs derive from a single foundation model and a condition in which they derive from
multiple models of different architectures and developers. As condition (iii) of Proposition 5
asserts, in a situation where a single model's systematic blind spots dominate all outputs,
generational heterogeneity on the human side should not be able to override them, and this
claim becomes testable as an interaction in which the reduction of overlap rates from
Working Paper | Ageless Management in the AI Era 149
generational heterogeneity shrinks or disappears under the single-model condition. This
measurement is an operationalization of the independence side of WP8's (Kadowaki 2026h)
oversight value = independence × detection probability; the detection-probability side is
measured separately with H3-type tasks — to conflate the two is to commit the error of rating
highly a supervisor pool whose blind spots merely fail to overlap while detecting nothing. It
should be noted that experimental evidence that information exchange broadens and factual
errors decrease in diversely composed groups exists in the context of racial diversity
(Sommers 2006), but replication on the age-and-generation axis is unconfirmed, and the
measurement of detection-overlap rates is precisely an attempt to fill that gap.
Fourth, the robustness of measurement itself — the ramparts against Goodhart's law. Since
this section's indicators, above all detection-task performance and role-fit rates, become
inputs to dynamic role allocation (Definition 1), measurement connects directly to allocation
and is therefore structurally exposed to gaming (Proposition 12). The operational
requirements of Proposition 12 are here given concrete form as measurement procedures.
First, the non-periodic replacement of detection and measurement tasks. Repetition of the
same tasks permits overfitting to the tasks — improvement in test-taking rather than in
experiential audit capacity — so the task pool is refreshed without notice, and discontinuities
in performance across refreshes are monitored as an indicator of overfitting. Second,
backtesting against real-work logs. Whether performance on the verification tasks used as
grounds for role allocation diverges from audit performance in real work — the record of
detections and misses measured by ex-post verification through unmediated audit of realwork
logs — is checked periodically, and any indicator for which divergence is systematically
observed is removed from the grounds of allocation. Only operation accompanied by ex-post
verification, not one-off test scores, can sustain measurement as a ground of allocation. Third,
unannounced, real-time blind in-situ verification (unannounced in-situ verification). What
backtesting verifies is the logs that remain, and as long as how logs are kept is itself under the
control of those subject to allocation, a deeper level of gaming remains — adaptation toward
"ways of keeping logs that backtesting does not detect" (Proposition 12). Accordingly, a
procedure of blindly observing, in real time and without advance notice, randomly selected
audit scenes in real work, and collating the judgments made on the spot with the entries in
the logs, is operated as a pair with backtesting — a double rampart against the two levels of
gaming: adaptation to the test and adaptation of log formation. In addition, monitoring of
each auditor's false-alarm rate is built into the operation of the indicators. Monitoring only
the detection rate (hit rate) cannot distinguish a shift of the response criterion toward
suspecting everything — a criterion shift that raises apparent detection while increasing false
alarms and verification costs — from a true improvement in detection capacity. Sensitivity d′
and response criterion c are separated within the framework of signal detection theory
(Appendix A), and each auditor's false-alarm rate is continuously monitored as an estimate of
the response criterion c. This measurement is the monitoring of the same quantity as the
rampart against gridlock through excessive flagging (Section 8), and it is also a condition for
Working Paper | Ageless Management in the AI Era 150
preserving the interpretability of detection performance as an indicator. These are, however,
ramparts, not guarantees. No means exists to completely prevent divergence between
measurement and true capacity, and the implications of that residual risk are discussed in
Section 10.3.
Finally, the risk of measurement misuse is stated explicitly. This section's indicators are
designed for improving role allocation and monitoring Brain Safety; the moment they are
diverted into instruments for rating individuals, adjudicating exit, or cutting remuneration,
Ageless Management degenerates into a device that merely replaces discrimination by
chronological age with discrimination by measurement. A decline in K-maintenance
indicators is an occasion for support and role redesign, not a ground for exclusion. This
danger is inherent in this paper's theory, and it is confronted head-on in Section 10.3.
9.4 Implementation Roadmap
Verification and implementation cannot be separated. The propositions of Ageless
Management include some that can be tested only inside an implemented multigenerational
ecosystem (Propositions 8 and 11), and conversely, implementation without verification does
not rise above the level of an ideal. This section presents a four-stage roadmap that raises cost
and the strength of causal inference step by step.
The first stage is a pilot. In the small-scale ecosystems of the author's own organization and
its partners, the H3-type verification tasks and the measurement of detection-overlap rates
are trialed, and the collection procedures for the indicators of Section 9.3 — proxy indicators
of K maintenance, u indicators, Brain Safety indicators — are established. The purpose of this
stage is not the estimation of effects but the verification of the feasibility of measurement. For
super-seniors, for whom participation in the tasks is itself a load, applying Brain Safety
(Section 8) — load ceilings and the adequacy of compensation — to the design of the
measurement itself is also among the procedures to be established at this stage.
The second stage is case studies. WP3 of this series (Kadowaki 2026c) collated theory with
observation for the theoretical framework of enterprise redefinition through structured
observation of company cases based on public information. The same observational method
is re-applied to this paper's framework: cases of organizations in the process of implementing
multigenerational role composition, non-employment participation, and AI-mediated
complementarity are described in structured form along the design variables of Section 6 (the
default design of Table 5 and deviations from it, the operation of dynamic role allocation, the
organizational design of decorrelation). Case studies do not settle causation, but they reveal
whether the propositions take observable form in real organizations and whether there are
unanticipated failure modes. The selection bias by which observed cases skew toward
successes is unavoidable here, and its implications must be interpreted together with the
survivorship-bias discussion of Section 10.4.
Working Paper | Ageless Management in the AI Era 151
The third stage is a longitudinal study. Ecosystem participants across multiple organizations
are formed into a panel, and the K and u indicators together with data on roles, contracts, and
load are tracked continuously. Because a single organization lacks both sample size and
exogenous variation, the formation of a multi-organization consortium is a precondition. At
this stage, H1's identification strategy comes into view. Exogenous institutional variation —
reform of Japan's in-work old-age pension offset (Section 7) or changes to mandatoryretirement
rules — is a candidate instrumental variable linking changes in work and role
occupation to cognitive outcomes, and if the panel is maintained across an institutional
change, it can be used as a quasi-experiment. As seen in Section 3, however, this field has a
history of conclusions reversing with the choice of instrument. From the outset, the design
incorporates the discipline of not betting on a single identification strategy and of reporting
including the discrepancies among estimates. The verification tasks at this stage include
estimating the depreciation function of detection capacity with years since leaving work as a
continuous variable. For the stratum long past exit, to which H3's eligibility criterion (within
five years of leaving practice) does not permit extrapolation, the speed at which experiential
audit capacity depreciates with idle time is an unresolved empirical question that directly
demarcates the scope of the resource-conversion claim (Propositions 2 and 11).
The fourth stage is medical collaboration. Measuring the rate of cognitive decline (MoCA,
etc.), H1's dependent variable, requires the ethics review, clinical expertise, and long-term
follow-up infrastructure of medical research, and cannot be carried out by management
studies alone. There is precedent. The effects of structured productive social engagement by
older adults on cognitive function and brain structure have been tested jointly by medicine
and social science in the randomized controlled trials of Experience Corps (Carlson et al.
2008; Carlson et al. 2015). These, however, are effects of fifteen hours per week of structured
volunteering in a limited population (chiefly low-income urban populations in the United
States), and extrapolation to work requires bridging through the superordinate concept of
productive social engagement (Section 3). What corresponds to this in the context of Ageless
Management is the randomized or quasi-experimental evaluation of the intervention of
occupation in Gc-exercising, AI-complemented roles. Including the design of ethically
permissible forms of assignment — such as randomizing the order in which roles are offered
to those wishing to participate — this stage presupposes joint design with medical
researchers.
This roadmap is at the same time a declaration of the provisional character of this paper's
claims. The first and second stages can be undertaken by practitioners including the author's
own organization, but the execution of the third and fourth stages — and above all the
interpretation of results — must be open to execution and replication by independent
researchers. A configuration in which the proposer of a theory monopolizes its verification is
to be avoided, particularly under the conflict of interest disclosed in Section 10.6. The results
of each stage — above all a rejection of H3 — connect directly to the revision or abandonment
of the propositions. Until verified, this paper's propositions remain proposals.
Working Paper | Ageless Management in the AI Era 152
10. Limitations and Self-Critique
This section is not a ritual enumeration of the paper's limitations. The theory of Ageless
Management carries an internal structure by which it can itself turn into a new apparatus of
discrimination, selection, and exclusion, and the proposer of the theory knows the pathways
of that transformation most concretely. This section discusses six limitations — the limits of
the evidence, new stereotyping, the risks of measurement, survivorship bias, economic
presuppositions, and conflict of interest — naming names wherever possible. The paper's
core thesis bears restating. Whereas conventional senior-employment and D&I arguments
have rested on a paradigm of "accommodation and compensation (cost/CSR)," this paper
presents a management model that, through bidirectional complementarity between
heterogeneous cognitive abilities mediated by AI, converts multigenerational and diverse
talent into co-creating agents of the managerial resource of brain capital. The sections that
follow specify the conditions under which this thesis fails, and the conditions under which,
even succeeding, it does harm.
10.1 The Limits of the Evidence — Most of the Propositions Are Untested
Begin with the most basic limitation. Of this paper's twelve propositions, not one has been
directly tested. The empirical work this paper relies on is in every case not a test of the
propositions themselves but peripheral evidence consistent with them. And that peripheral
evidence itself has the following three weaknesses.
The first is misalignment of axes. What the empirical work on compression, oversight, and
depreciation organized in Section 4 measured was tenure, skill, and expertise, not
chronological age, and no study that operationalized age or years of tenure and measured
performance in AI oversight and verification exists within the range of this paper's search
(Section 4.4, search record). Neither the reality of experiential audit capacity (Definition 4)
nor the claim that it is not compressed by AI assistance (Proposition 3) has direct evidence
until Hypothesis H3 is carried out. That the keystone of this paper's theory is placed at the
point where the evidence is currently thinnest should be stated explicitly.
The second is the mixed character of the evidence. The relation between work and brain
health (the background of Propositions 7 and 8) is a field in which the sign of the estimates
has reversed with the choice of instrumental variable (Rohwedder & Willis 2010 vs. Coe et al.
2012), and the summation of the systematic reviews goes no further than "the evidence is
mixed, with large research gaps" (Meng et al. 2017; van Ours 2022). On age diversity and team
outcomes (the background of Propositions 5 and 6), the meta-analytic average effect is near
zero (Joshi & Roh 2009; Schneid et al. 2016; Wallrich et al. 2024), and this paper's Proposition 6
is a hypothesis that stacks an untested moderator — AI mediation — on top of this headwind
baseline. Citing only the favorable side of the estimates would make this paper's claims look
strong, but that is not the state of the evidence.
Working Paper | Ageless Management in the AI Era 153
The third is the mixture of evidence grades. This paper has referred, alongside peerreviewed
research, to preprints, technical reports, surveys by NGOs and membership
organizations, and practitioners' essays, with the grades made explicit. In particular, the
headwind data of Section 4.5 and parts of the description of institutional operation in Section
7 rest on grey literature, and these are not measurements of ability or outcomes. Making
grades explicit is a requirement of honesty, but it does not cure the weakness of the evidence
itself. In addition, there are limits of external validity. Much of the empirical work this paper
cites is data from the United States and Europe, often from a single firm, a single occupation,
or a limited population (noted individually at various points in Sections 3, 4, and 6), and its
transferability to the context of Japanese employment practice, pension institutions, and
elderly employment is itself an assumption requiring verification. The institutional
comparison of Section 7 has done no more than roughly demarcate the conditions of that
transfer. Taken together, the accurate positioning of this paper is not a report of established
facts but a proposal of theory equipped with refutation conditions and a map of verification
(Section 9). This qualification, stated in the abstract, becomes all the more important when
read together with the COI disclosure at the end of this section (Section 10.6).
Alongside the limits of the evidence, the limits of the concepts should also be stated. This
paper's theory is built on the two-axis classification of intelligence into Gf and Gc, but as
Definition 2 made explicit, this classification is a relative weighting on a continuum, not a
binary, and many real tasks are inseparable bundles of both components. If, as the terms Gftype
and Gc-type recur through this paper, they harden in the reader's mind into binary
categories, that is a misreading of Definition 2 — and at the same time a misreading invited
by this paper's own mode of exposition. Moreover, the construct validity of experiential audit
capacity (Definition 4) is not established. Years of experience, used to operationalize longterm
domain experience, is a coarse proxy variable, and there is no guarantee that the same
years of experience form the same detection capacity — the quality of experience (exposure
to failure cases, the density of feedback) is very likely the true explanatory variable, and this
refinement can only be carried out once H3's results are in.
10.2 The Risk of New Stereotyping — A Self-Critique of Table 5
This paper has repeatedly used the role image of "seniors as auditors." Although this image
was introduced in order to dismantle the presumption of ability from age, it carries the risk
of itself hardening into a new presumption. The statement that older people are suited to
auditing carries on its reverse side the implication that older people are not suited to
execution, and it revives the very operation Definition 1 was meant to remove: inferring role
aptitude from chronological age. Even when the direction is celebratory, the structure of
inferring an individual's characteristics from age is isomorphic with age discrimination, and
the fixation of seniors-as-auditors can become age discrimination's inverted twin.
This danger lies not outside this paper but inside it. To name names, it is Table 5 of Section
6. Table 5 presented as a "default design" a role composition assigning experiential audit and
Working Paper | Ageless Management in the AI Era 154
contextual evaluation to super-seniors, rapid prototyping to youth, orchestration to the midcareer
generation, and lived-experience problem perception to socially marginalized groups.
This paper positioned it as a convenience of departure and made deviation through dynamic
role allocation (Proposition 12) the rule, but the danger remains that practice will operate the
table as a norm. Tables circulate faster than theories. If Table 5 alone is excerpted and
transcribed into training materials and HR systems, then older people who want and can
perform Gf-type roles, early specializers with high experiential audit capacity at a young age
(the thought experiment of Section 4.4), and socially marginalized people who want to
participate in execution rather than audit are rendered invisible once again inside the new
template. That low age stereotyping is a condition for the success of age diversity is also
empirically suggested (Wegge et al. 2012), and if this paper's schema supplies new
stereotypes, that is a self-destructive consequence that undermines the paper's own
theoretical premises.
The danger of fixation is not limited to super-seniors. Table 5 assigns youth to rapid
prototyping, but if this role image hardens, youth will keep specializing in AI-assisted
execution, and by the crutch-effect logic seen in Section 4.3 they risk being deprived of the
opportunity to form contextual verification capacity — that is, the acquisition pathway of
future experiential audit capacity. This stands in direct tension with Proposition 8, which
asserts the early formation of K among youth. Unless the role design of the multigenerational
ecosystem includes intergenerational role transition (the staged entry of youth into audit
experience), this paper's schema fixes the present division of labor and dries up the future
supply. Likewise, the description that positions socially marginalized groups as suppliers of
lived-experience problem perception carries a pathway of sliding into the exploitation of
lived experience — tokenism in which opinions alone are solicited while decisions and
rewards are withheld. For the provision of a lived-experience perspective to be respected as a
role means that it carries compensation and decision rights (Section 8), and it must be
distinguished from the staging of symbolic participation.
Three counter-devices are made explicit. First, Definition 1 removes chronological age from
the criteria of role allocation, leaving only measured characteristics, experience, health
status, and the person's own intent. Second, Proposition 12 asserts the inferiority of age-fixed
allocation, on the grounds of intra-individual variation in cognitive characteristics and the
overlap of distributions across age groups, and attaches a refutation condition. Third, the
operating model of Section 6, as the name dynamic role allocation indicates, requires periodic
re-measurement and re-allocation of roles. But the existence of counter-devices does not
mean the disappearance of the danger. What this paper can do extends no further than to
state here in plain words: any reading that cites Table 5 as a norm is a misreading of this
paper.
Deeper still beneath the danger of stereotyping lies the critique of commodification. This
paper's vocabulary — brain capital, stock and utilization rate, the decomposition of job
bundles into Gf-type and Gc-type components, the connection of cognitive assets — can be
Working Paper | Ageless Management in the AI Era 155
read as a schema that disassembles human beings into parts of brain function, selects the
parts with market value, and reuses them as managerial resources, and this reading cannot
be dismissed as a mere misreading. This paper has in fact consistently discussed the
operations of decomposing jobs into components, measuring individuals' cognitive
characteristics, and optimizing their connection, and what the critique points at is not this
paper's periphery but its method itself. This paper's answer to it is placed, as stated in Section
1.4, in the normative anchor that connects its concept of capital to Sen's (1999) capability
approach. The measurement and connection of brain capital can be legitimate only insofar as
it serves not the maximization of organizational output for its own sake, but the return to
individuals of the freedom of participation that the coarse proxy variable of chronological
age has taken from them — the substantive freedom to live a life one has reason to value.
That Definition 1 includes the person's own intent among the allocation criteria, that
Definition 8 makes protection a precondition of participation, and that Propositions 8 and 11
demand the bidirectionality of benefits across all participating strata are the expressions of
this limitation inside the theory. But the existence of the anchor does not extinguish the
danger of the vocabulary. The words capital, asset, and utilization rate can circulate stripped
of the limiting clauses of intent and dignity, and at that point this paper's framework
functions as a lexicon for the instrumentalization of human beings. This danger too, like
Table 5, is one this paper has itself supplied.
10.3 The Risks of Measurement — Transformation into an Apparatus of
Selection
Definition 1 placed measured cognitive characteristics in the vacancy left by removing
chronological age. This substitution is the core of this paper's theory, and at the same time its
most dangerous point. If measurement is diverted into selection for hiring, treatment, and
exit, Ageless Management will not have abolished age discrimination but merely replaced it
with trait discrimination. Discrimination by chronological age is at least visible, and is
established as an object of legal regulation (Section 7). Selection based on the measurement of
cognitive characteristics, clothed in the appearance of objectivity, is that much harder to
make visible and to regulate. This paper cannot deny the possibility that Definition 1 will be
read as the blueprint of a new apparatus of selection.
There are four concrete dangers. The first is mismeasurement. The measurement of
cognitive characteristics is accompanied by measurement error, situational dependence, and
practice effects, and a low score at a single point in time is misread as a permanent deficit of
ability. The second is gaming. If measurement connects directly to treatment, the
optimization of measured performance becomes an end in itself, and the validity of the
measurement itself collapses. The third is cultural and attribute bias. When the language,
format, and context of a test work against particular groups, measurement automates bias
under the appearance of neutrality. The observation that algorithmic hiring tools can
perpetuate discrimination against people with disabilities (El Morr et al. 2024) is a real
Working Paper | Ageless Management in the AI Era 156
instance, in an adjacent field, of the pathway by which a measurement apparatus that
proclaims inclusion reproduces exclusion. The fourth is use beyond purpose. The pathway by
which cognitive-characteristic data collected for role allocation is diverted to the person's
detriment — exit inducement, insurance, credit — is, under the protection vacuum of nonemployment
participation (Proposition 9), especially unclosed.
The safety devices internal to this paper are the refutation condition of Proposition 12 — if
the costs of trait measurement, mismeasurement, and gaming exceed the gains of dynamic
allocation, the superiority of dynamic allocation is rejected — together with the explicit
prohibition of misuse in Section 9.3 and the Brain Safety of Section 8 (a design that restricts
measurement to an instrument of protection). As to gaming, Proposition 12 wrote into the
proposition itself, as operational requirements, the ramparts against Goodhart's law — the
non-periodic replacement of measurement tasks and ex-post verification through
unmediated audit of real-work logs (backtesting) — and Section 9.3 gave these concrete form
as procedures of indicator operation. But ramparts do not extinguish residual risk. As long as
measurement is connected to allocation and treatment, no means exists to completely
prevent divergence between measured performance and true capacity — overfitting to the
tasks, manipulation of the very real-work logs that backtesting targets, the skewing of effort
toward measurable components. The formula that when a measure becomes a target, it
ceases to be a good measure (Strathern 1997) can be mitigated by replacement and
backtesting but not repealed. The superiority of dynamic role allocation (Proposition 12) must
therefore be read as a net-benefit claim that prices in this residual cost of gaming, and that is
why its refutation condition explicitly names the costs of measurement, mismeasurement,
and gaming. On the institutional side, the minimum line is purpose limitation of
measurement (for role allocation, not for exit adjudication), the person's right of access to
results and right to re-measurement, and the guarantee of alternative pathways upon refusal
of measurement. But explicit statement is not prevention. The observation of an operational
reality in which the misuse of measurement exceeds its gains is a legitimate ground for the
normative criticism that the organization in question lacks the standing to claim the benefits
of Ageless Management. This normative statement, however, must not be diverted into a
defense of the propositions. In testing Propositions 8 and 11, the adjudication of whether the
conditions were satisfied is by ex-ante determination prior to the observation of outcomes
(Section 5). Reclassifying failure cases after the fact as "not having been implemented" in
order to protect the propositions is, as the refutation conditions of both propositions state
explicitly, not admitted as a defense, and the moment this epistemic discipline is broken, the
refutability that is this paper's signboard loses its substance.
Before the risk of misuse lies a design-level problem: the legality of the measurement
apparatus itself. Section 7 compared institutions organized around age, but the measurement
apparatus Definition 1 puts in age's place itself intersects with a different lineage of legal
regulation. First, it is a general demand of anti-discrimination law that cognitive testing in the
employment context be job-related, and comprehensive cognitive measurement unconnected
Working Paper | Ageless Management in the AI Era 157
to role requirements may not be justifiable as a selection procedure. Second, using health
status as an allocation criterion can stand in tension with the legal regimes prohibiting
discrimination on the basis of disability — the ADA (Americans with Disabilities Act) in the
United States; in Japan, the Act on Employment Promotion of Persons with Disabilities and
the Act for Eliminating Discrimination against Persons with Disabilities. The pathway by
which restriction of roles based on health status collides with duties to provide reasonable
accommodation and with limits on medical examinations cannot be excluded. Third, data on
cognitive characteristics and health status are of a class that may constitute special carerequired
personal information under Japanese law, subject to heightened discipline in
acquisition and use. That is, the apparatus introduced to remove discrimination by
chronological age can be assessed as unlawful or improper as trait discrimination or healthinformation
discrimination — a problem of the institutional viability of the design itself, prior
to the misuse discussed at the opening of Section 10.3. Furthermore, because years of
experience correlate strongly with chronological age, experience-based role allocation, even
where at the individual level it is reattribution away from age (Section 4.4), can appear at the
group level as a difference in treatment correlated with age — indirect discrimination. The
implementation of Ageless Management presupposes passing these legality reviews at the
design stage. The description in this section is an identification of issues, not legal advice, and
individual assessments depend on the statutes of each jurisdiction and the facts of each case
(the same reservation as at the opening of Section 7).
The handling of data carries its own dangers. In the non-employment participation on
which Ageless Management depends (outsourced engagement, advisory roles, PBL, NPO
partnership), the disciplines protecting workers' personal information and governing health
information that presuppose an employment relationship may not apply, or their application
may be unclear. The data whose collection this paper's framework demands — cognitive
characteristics, health status, load — belong by their nature to the class whose misuse carries
the heaviest consequences. Minimization of collection, limitation of retention periods, giving
substance to the person's consent (under asymmetric bargaining power, consent easily
becomes a formality), and prohibition of data linkage across organizations are the minimum
line (Section 8), but these remain voluntary standards without legal backing. The protection
vacuum Proposition 9 identified extends not only to workers' accident compensation and
remuneration but to data.
10.4 Survivorship Bias — The Re-Invisibilization of Those Who Cannot Work
This paper's theory is, by its very materials, skewed toward survivors. The evidence on
capacity maintenance in Section 2 is data from those who could keep participating in testing;
the association between work and health in Section 3, from those who could keep working;
the evidence on older workers' productivity in Section 6, from those who remained in the
factory (Börsch-Supan & Weiss 2016 corrects statistically for survivorship bias, but complete
removal is impossible). A theory standing on this skew carries the danger that, by speaking of
Working Paper | Ageless Management in the AI Era 158
the possibilities of older people who can work, it renders invisible once again those who
cannot.
One should have a sense of scale. Healthy working life expectancy at age 50 in England —
the years a person can expect to spend both healthy and in work — is estimated at about 9
years, below the years remaining to state pension age (Parker et al. 2020). And this figure has
steep gradients: 10.9 years for men against 8.3 for women, 6.8 years in the most deprived
areas, with regional differences reaching about 4.5 years. In the sample of a large Japanese
longitudinal survey as well, 66.0% of those aged 65 and over are retirees, and 10.2% have no
work experience at all (Takeuchi et al. 2024). The image this paper draws — "super-seniors
exercising experiential audit capacity into their 90s" — is legitimate as a description of
possibility, but if it circulates as a standard image, it turns into a pressure to reinterpret the
condition of the many who cannot participate because of health, caregiving, poverty, or
disability as if it were a matter of will or effort. The pathway by which the vocabulary of
Ageless Management is appropriated to justify policies such as raising the pension eligibility
age or tightening work requirements lies outside this paper's control, but being foreseeable, it
is warned against explicitly here.
Furthermore, this paper's theory has, even regarding change in cognitive function itself,
shone its light on the side that can be maintained — Gc-type abilities and metacognition. But
the changes of cognition that accompany aging do actually include decline to a level at which,
if it progresses, experiential audit capacity itself ceases to hold. For people in states of
dementia or long-term care need, this paper's prescription of participation in audit-type roles
is no answer. What Ageless Management can speak to extends only to role design within the
range where cognitive participation-capability is preserved, and to the provision of the
environment that preserves that capability (Brain Safety, cognitive engagement); dignity and
care after the capability is lost are a separate task outside this paper's theory — and one that
does not rank below it. Talk of "lifelong active engagement" that blurs this boundary runs
counter to this paper's intent.
The second boundary condition of Proposition 2 (Section 5) writes this fact into the theory
in plain terms. AI's complementation of Gf-type components holds only within the range in
which standard cognitive screening does not fall below the threshold of mild cognitive
impairment, and this model cannot complement a state in which attentional resources
themselves are exhausted — this lower bound is the acknowledgment that the
complementation model has a neurological limit point. This paper makes this boundary
explicit not to justify exclusion but to draw the theory's honest boundary. A theory that does
not write down its lower bound would, by the overclaim that AI complementation is possible
for people in every cognitive state, instead blur the distinctiveness of the real support — care,
medicine, income security — needed by those outside the boundary. At the same time, as
Proposition 2 itself states explicitly, this lower bound is a boundary on allocation to
experiential-audit roles, not a constraint on participation in the multigenerational ecosystem
in general. The possibility that a person below the threshold participates in forms other than
Working Paper | Ageless Management in the AI Era 159
audit — the provision of lived-experience problem perception, the handing down of
narrative, loose social connection — remains open outside the boundary, and its concrete
design remains a task this paper has not discharged. The rampart against survivorship bias is
not to hide the existence of the limit point. It is to make the location of the boundary explicit
in measurable form, and to condition the dignity of people on neither side of it on
participation — the distinction drawn in the preceding passage is the minimum line for that
purpose.
The breakwaters within this paper's framework are three. First, Definition 1 includes
health status and the person's own intent among the allocation criteria, and participation is a
right, not a duty. Second, Definition 8 (Brain Safety) requires a design that caps the intensity
and load of participation, and Proposition 10 requires that this protection be designed
symmetrically as a response to asymmetric bargaining power. Third, the socioeconomic
gradient of healthy life expectancy requires that the ecosystem's design object include not
only "the utilization of those who can participate" but "the distribution of participationcapability
itself" — this is also the reason the institutional analysis of Section 7 is a
component of the management theory, not an external condition to it. Even so, this paper
presents no theory of its own about the welfare of people in states where they cannot work.
Ageless Management is a theory of inclusion through work and participation, and it is no
substitute for inclusion that does not pass through participation — income security, care,
dignity without participation. Making this boundary explicit is the limit of the honesty this
paper can offer.
10.5 The Limits of the Economic Presuppositions
This paper's theory places several economic presuppositions implicitly. In domains where
they break down, implementation fails even if the propositions are theoretically correct. The
first is the cost of AI use. This paper took the fall in the marginal cost of Gf-type components
(Proposition 1) as its point of departure, but what an organization bears is not only the unit
price of inference. Systems that present output and grounds in auditable form, work
environments compliant with Brain Safety, and the cost of keeping up with model updates
remain as fixed costs. The second is the cost of task decomposition. The work of decomposing
job bundles into Gf-type and Gc-type components and re-allocating them across AI and
multiple participants itself generates costs of analysis, coordination, and contracting.
Moreover, the boundary of AI's zone of competence is hard to see in advance (the jagged
frontier of Section 4), and errors of decomposition surface only after the fact, as oversight
failures. The third is the transaction cost of running the ecosystem. Managing the multiple
contractual forms of employment, outsourced engagement, advisory roles, PBL, and NPO
partnership; measuring and re-allocating participants; and maintaining a composition that
preserves decorrelation all require administrative cost.
A large part of these costs diminishes with organizational scale, and this paper's design
theory may be tilted in favor of large organizations or consortia. This point must be accepted
Working Paper | Ageless Management in the AI Era 160
explicitly as a critique of external validity. The cost of building and maintaining, firm by firm,
the measurement and governance apparatus this paper demands — measurement and remeasurement
of cognitive characteristics, dual-track auditing, risk-weighted sampling,
unannounced in-situ verification, a layer-separated feedback infrastructure — is enormous,
and this paper cannot exclude the possibility that net benefit turns positive only in large
organizations. The response to this critique rests on two pathways. The first is the platform
pathway presented as a possibility in Section 6.5. If the administrative and measurement
infrastructure is standardized as common infrastructure, the marginal cost per firm can fall,
and if this holds, the scale threshold of net benefit comes down — though this is the
presentation of a possibility, not an assertion that it holds. The second is the net-benefit
requirement of Hypothesis H2. Because H2 records time to decision and verification effort as
cost variables and makes a positive net benefit after their deduction a requirement for
support, if administrative overhead eats up the benefits, Proposition 6 is by design not
supported — that is, the testing of this external-validity critique is built in not as an objection
from outside the theory but as an adjudication criterion inside this paper's verification plan.
Implementability in small, medium-sized, and regional organizations could not be examined
in this paper, and it is an important observation item for the pilot and the case studies in the
roadmap of Section 9. Also, regarding the time lag until the marginal value shift of
Proposition 1 is actually reflected in the labor market's reward structure, and the possibility
that the reward premium of Gc-type tasks is concentrated in particular occupations and
strata, this paper has nothing beyond theoretical prediction. In domains where the economic
presuppositions are not met, the multigenerational ecosystem will remain a cost center, and
pressure to revert to the "accommodation and compensation" paradigm will operate. That is
not a refutation of this paper's thesis but a demarcation of its scope — yet narrowness of
scope directly diminishes the value of a theory.
Finally, the temporal reach (Temporal Robustness) of this paper's theory should be made
clear. As declared in Section 1.4, what this paper's propositions are anchored to is not the
performance profile of the current generation of generative AI — the transient performance
gap that "today's LLMs are poor at contextual verification." The anchor point is the structure
of verification independence formalized by WP8 (Kadowaki 2026h, Propositions 10 and 11).
Namely: for any intelligent agent whatsoever, a verifier whose errors correlate with its own
training distribution cannot supply itself with statistical independence in detecting its own
systematic blind spots. The independence of verification is not supplied by improvements in
capability — however capable the verifier, so long as it shares an error distribution with the
generator, its misses remain correlated with the generator's misses. Therefore, even in a
world where future AI has made self-correction and self-verification highly sophisticated, (a)
the detection of systematic blind spots deriving from the training distribution will still
require decorrelated verifiers, and (b) the sources of their supply will be both humans with
heterogeneous experience and models of different lineages, so that this paper's framework
Working Paper | Ageless Management in the AI Era 161
persists as the allocation problem between the two. The audit-certification pathway shown in
the commentary on Proposition 1 (Section 5) is likewise anchored to this structure.
This structural anchoring, however, does not guarantee the size of the human contribution
in that allocation. Stated honestly, this paper cannot exclude the possibility that the advance
of machine-side decorrelation — cross-verification by models of different architectures and
developers (Proposition 5's model condition) — and the acceleration of verification-skill
acquisition in AI-use environments will progressively shrink the residual niche of humansupplied
decorrelation. Because this shrinkage would appear not as a one-stroke refutation of
the theory but as a gradual narrowing of its scope, it is all the harder to detect. What this
paper designates as its detector is the second sentence of Proposition 3's refutation condition
— longitudinal evidence that the acquisition of verification capacity accelerates in AI-use
environments and substantially substitutes for the effect of accumulated experience
(longitudinal compression). This paper's propositions should therefore be read as dynamic
claims to be re-tested at each point of the technology's development; the results of H3 must
always be annotated with the model generation at the time of data collection; and the
longitudinal measurement in the roadmap of Section 9 must include, among its objects of
observation, the very speed at which this residual niche shrinks.
10.6 Disclosure of Conflict of Interest and Classification of Self-Citations
Finally, this paper discloses the conditions of its own production. The author of this paper is
affiliated with VURA Capital Innovation Holdings, Inc. The company is a business operating in
the domains of longevity and brain capital, and it stands to gain business benefit from the
spread of the ideas of Ageless Management and brain capital management. That is, this paper
is written not by an observer with no stake in the theory's consequences, but by an interested
party who may benefit from the theory's diffusion. Readers should read this paper with this
structure discounted, and this paper's disciplines — the attachment of refutation conditions
to every proposition, the explicit grading of evidence, the explicit presentation of headwind
data (Section 4.5), and this section — are designed on the premise of that discount. It is for the
same reason that the roadmap of Section 9 makes it a requirement that verification from the
third stage onward, and the interpretation of results, be open to execution and replication by
independent researchers. A configuration in which the proposer of a theory monopolizes its
verification is not permissible under this conflict of interest.
The classification of self-citations is also made explicit. This paper has cited eight papers of
the VURA Working Paper Series, and their positioning divides into three types. The first is
inheritance. The inversion of valuation in Future Value Theory (Kadowaki 2026a); the K × u
framework and the state-versus-intervention distinction of Brain Capital Management
(Kadowaki 2026e); and the oversight value = independence × detection probability, error
decorrelation, the solvency condition, and the non-delegable residual of the Human on the
Loop argument (Kadowaki 2026h) were used by this paper as presupposed frameworks. The
second is reference. Enterprise redefinition and its observation (Kadowaki 2026b; 2026c),
Working Paper | Ageless Management in the AI Era 162
purpose-description-based role design (Kadowaki 2026d), the self-defined society (Kadowaki
2026f), and redefinition capitalism (Kadowaki 2026g) were cited solely to indicate position
within the series. The third is what newly becomes an object of verification in this paper. The
relation between BCM's "protected unassisted practice" and Hypothesis H3, and the extension
of HOTL's decorrelation condition to the generational axis (Proposition 5), are new claims of
this paper, and they are untested.
The point of this classification lies in the following discipline. The inherited frameworks
themselves are all unrefereed working papers, and mutual citation within the series is no
substitute for external evidence. This paper does not make claims for which no independent
supporting evidence outside the series exists appear established through a chain of selfcitations.
The strength of this paper's claims is supported not by the internal coherence of the
series but solely by the independent empirical literature cited in Sections 2 through 4 and by
the executability of the verification plan of Section 9. And since the core of that verification
plan (H3) has not yet been executed, what this paper has reached is a proposal put into
refutable form — this sentence is the conclusion of this section, and the premise of the
conclusion of the next.
11. Conclusion
This paper has presented the theory of a management regime — Ageless Management — that
removes the variable of chronological age from decisions on role allocation, evaluation,
participation, and exit, and that dynamically allocates roles, within an ecosystem not limited
to the boundary of employment, on the basis of measured cognitive characteristics,
accumulated domain experience, health status, and the person's own intent. The skeleton of
the theory consists of three operations. First, the operation of reattributing the oversight
value of older workers from age to "the interaction of long-term domain experience × Gc ×
metacognition" — that is, to experiential audit capacity. Second, the operation of formalizing
generational heterogeneity as an organizational source of error decorrelation in AI oversight.
Third, the operation of converting the three social problems of population aging, constrained
youth participation, and the exclusion of socially marginalized groups into untapped sources
of brain capital (K × u), conditional on AI-mediated complementarity and on Brain Safety
operated in an ex-ante verifiable form. These were cast into eight definitions, twelve
propositions each carrying a refutation condition, and three testable hypotheses (Section 9).
This paper's position within the series should also be confirmed. This paper extends Brain
Capital Management (Kadowaki 2026e), which treated brain capital as an asset of a single
firm's employees, vertically into a multigenerational ecosystem not limited to the boundary of
employment, and at the same time bridges the theory of future value (Kadowaki 2026a) and
the theory of oversight structure (Kadowaki 2026h) along the axis of age. WP8, which
demonstrated the failure of oversight, and this paper, which asks about the supply side of
oversight, divide a single question between demand side and supply side — who bears the
Working Paper | Ageless Management in the AI Era 163
residual that cannot be delegated to AI, under what protections, and at what cost. This paper's
answer — experiential audit capacity and generational decorrelation — is not the only
answer to this question, but the first candidate arranged in a testable form.
The meta-structure in the lineage of theory declared in the introduction (Section 1.4) should
also be reconfirmed here. The main axis of this paper is a theory of firms' competitive
advantage. Experiential audit capacity, as a path-dependent product that can be accumulated
only over time through long-term domain experience, is a candidate for a scarce and hard-toimitate
oversight resource (the resource-based view — Barney 1991), and the configurations
that organize it — the oversight portfolio that designs error decorrelation, and dynamic role
allocation — sit at the level of dynamic capabilities that integrate and reconfigure resources
under change (Teece, Pisano & Shuen 1997). By contrast, the institutional argument (Section
7) and the health argument (Section 3) are not parallel claims of this paper but the
institutional context and complementary assets that make the acquisition of competitive
advantage possible. The protective institutions for non-employment participation are the
institutional context that makes connection to oversight resources possible, and Brain Safety
and brain health are the complementary assets that prevent the depreciation of the resource
and sustain its utilization. That this paper — a theory of management — discussed
institutions and health at length is due to this complementary structure; conversely, when the
institutional context and complementary assets are absent, experiential audit capacity
remains unexpressed as a source of competitive advantage.
What ran through this paper's argument is a discipline of conditionality. This paper did not
claim that "humans are good at oversight"; it claimed only that, given that a residual of
oversight not delegable to AI exists, the supply side of oversight — a scarce and consumable
resource — must be interrogated. It did not claim that "working makes you healthy"; it
presented a mediation hypothesis (H1) conditional on the quality of work. It did not claim
that "multigenerational means more value"; on top of the empirical baseline of a near-zero
average effect, it placed a conditional proposition with AI-mediated complementarity as the
moderator. Nor did it unconditionally claim that "different generations mean fewer blind
spots"; it conditioned the supply of decorrelation on three conditions (Proposition 5) —
unmediated audit channels, flattened power gradients, and cross-validation across multiple
foundation models. On the audit value of experience itself, it drew the boundaries of
impairment under confirmation-congruent errors and the half-life of knowledge (Proposition
4), and the neurological boundary of a cognitive-screening floor (Proposition 2). And it offered
the keystone of the theory itself — the existence of experiential audit capacity — as the object
of the highest-priority, lowest-cost experiment (H3). Most of the propositions are untested,
and this paper is not a report of established facts. In addition, this paper's theory contains
within it pathways of misuse — a new stereotyping of seniors as auditors, the conversion of
measurement into an apparatus of selection, and the re-invisibilization of those who cannot
work — and Section 10 identified them by name.
Working Paper | Ageless Management in the AI Era 164
Even so, the reason this paper presents this theory lies in one point inherited from the
original draft. Ageless Management is not an improvement of measures for making older
people work longer. What it aims at is to dismantle the institutions and notions by which
human beings are uniformly severed from the front line of social participation by the single
variable of chronological age, and to prepare a structure in which human beings carrying
cognitive change can, together with the complementary apparatus of AI, participate in the
creation of value across the lifespan — a state in which human dignity and organizational
value creation are compatible (Human Flourishing). The linear three-stage model of life —
education, work, and retirement — is an institutional product of an era when average
lifespans were shorter than they are now, and the extension of lifespans is already
invalidating this model's premises. The dynamic role allocation this paper has presented is an
attempt to concretize, at the level of management, one candidate for what should come after
this model — a configuration in which roles are continually reallocated across the lifespan
not as a function of age but as a function of measured characteristics, experience, and intent.
Longevity is experienced as a burden when it lacks a structure to support it, and as a
resource when it gains one. Which it becomes is a matter not of biology but of design, and
among its design variables, as this paper has argued, not a few lie within the reach of
management and institutions. And that design can be legitimate only inside the boundary
line of Section 10 — that the dignity of people who are in a state of being unable to work is
not conditioned on participation.
The implications of this paper are stated in minimal form for each type of reader. What this
paper offers to management practice is not a collection of measures whose success is
guaranteed but an inventory of design variables — the decomposition of job bundles into Gftype
and Gc-type components, the measurement of experiential audit capacity and its
connection to roles, generationally heterogeneous supervisor configurations (including the
three conditions of Proposition 5 — raw-audit slots, flattened power gradients, and crossvalidation
across multiple foundation models), state monitoring of K × u, and Brain Safety as
bidirectional protection against both overload and exploitation. A partial adoption lacking
any one of these — above all, utilization without protection — is not recognized as an
implementation within this paper's framework. The implication for policy, as argued in
Section 7, is that institutional development filling the protection gap for non-employment
participation (Proposition 9) is a precondition of Ageless Management. Flexibilization without
protection is a passage to exploitation, and this paper refuses to be cited as an argument for
flexibilization. For research, the map of verification in Section 9 is, as it stands, an inventory
of research problems. H3 in particular can directly test the core of the theory at low cost, and
whichever way the result falls — as the first direct evidence of the non-compressibility of
experience, or as a refutation of the core of this paper's theory — it constitutes a contribution.
Within the scope of this paper's search, no study was found that operationalized age or
tenure and measured performance in AI oversight (Section 4), and this gap is worth filling
beyond the question of whether this paper's theory is right.
Working Paper | Ageless Management in the AI Era 165
Yet the sentences that speak of this prospect must be chosen carefully. What this paper has
shown is not a realized achievement but a structure of possibility. That complementation of
Gf-type components can remove the cognitive bottleneck, that experiential audit capacity can
become a scarce input to oversight, that generational decorrelation can widen an
organization's detection set, that the three social problems can be converted into sources of
brain capital — all of these are written in the modality of "can," and their realization is not
automatic. What is required is the execution of the verification program of Section 9, the
revision or abandonment of the propositions according to its results, and the institutional
design that fills the protection gaps identified by Propositions 9 and 10. Absent verification,
this paper remains a system of aspirations; absent institutional design, implementation can
degenerate into exploitation. What this paper has shown is a structure of possibility, and its
realization is conditional on the verification of the propositions and on institutional design.
Appendix A Experimental Protocol for Hypothesis H3
This appendix details the experimental protocol for Hypothesis H3 (the direct test of the noncompressibility
of experience) presented in Section 9. The design skeleton of H3 is as follows.
A 2×2 factorial design [experience level: long-term domain-experienced practitioners vs.
junior participants] × [AI assistance: present vs. absent]. The task is the verification of AIgenerated
business proposals and documents in which misinformation and contextual risks
have been embedded. So that the embedded errors do not depend on the experimenters'
preconceptions, the tasks include items blind-generated from an empirical failure-case
dataset of accidents, scandals, and business failures that actually occurred (A.2), and part of
the errors are constructed in both a form congruent with auditors' industry conventional
wisdom (confirmation-congruent) and a form contrary to that wisdom (deviating), so as to
simultaneously test boundary condition (i) of Proposition 4. The task documents are prepared
in two versions — a high-fluency version and a low-fluency version, identical in content and
differing only in stylistic polish — and fluency is controlled as a factor or covariate, so as to
simultaneously test boundary condition (iii) of Proposition 4 (A.2). The dependent variables
are four: (i) the detection rate of surface errors, (ii) the detection rate of contextual and
practical risks, (iii) the quality of correction proposals (blind evaluation), and (iv) the overrejection
rate — the rate at which embedded "correct but counter-conventional innovative
proposals" were erroneously rejected. The prediction is that for (i) the experience gap shrinks
under AI assistance (consistent with prior research), whereas for (ii) and (iii) the main effect
of experience persists and does not shrink even under AI assistance (the direct test of
Proposition 3). No prediction is placed on (iv); it serves as an exploratory indicator of the
boundary at which experience turns into over-auditing. The following records, in order,
participant requirements (A.1), task construction (A.2), procedure (A.3), evaluator blinding (A.
4), and the testing plan (A.5), and appends at the end the team-assignment procedure
(stratified randomization) for Hypothesis H2 (A.6). Implementation presupposes
Working Paper | Ageless Management in the AI Era 166
preregistration (registration of hypotheses, exclusion criteria, and the analysis plan) and
research ethics review.
A.1 Participant Requirements and Operational Definitions
Experience level cannot be randomly assigned. The experience factor of this experiment is
therefore a measured quasi-experimental factor, and only the AI-assistance factor is
randomized. With this asymmetry made explicit, experience level is operationally defined as
follows.
Long-term domain-experienced practitioners are those with 15 or more cumulative years
of practical experience in the task domain (the industry to which the business proposals
described below belong), and who have been away from practice in that domain for no more
than 5 years. Junior participants are those with fewer than 5 cumulative years of practical
experience in that domain. Years of experience are measured by self-report combined with
employment-history verification (documentary confirmation of periods of tenure or
certification by the affiliated organization). Those with 5 or more but fewer than 15
cumulative years are excluded from the main analysis (including the boundary-ambiguous
intermediate stratum would blunt the contrast of the factor; the exclusion criterion is stated
explicitly in the preregistration).
The eligibility criterion "away from practice for no more than 5 years" imposes an
important limit on generalizability. This criterion excludes from the experiment the superseniors
long after exit — the stratum well beyond 5 years away from practice — whom this
paper's management model envisions as the principal source of supply. Whatever the results
of H3, therefore, they cannot be extrapolated to that stratum. Even if H3 is supported, what it
shows is non-compressibility among practitioners who are active or recently exited; for the
long-exited stratum, estimating the depreciation function of detection capacity with years
since leaving practice as a continuous variable remains an independent verification task
(Section 9.4).
At the same time, even within this eligibility range (0–5 years after leaving practice), years
elapsed since exit vary and can confound the experience factor through the depreciation of
unexercised abilities (Section 4.3). Accordingly, years elapsed since leaving practice (set to 0
for those currently working) are recorded in months for all participants and controlled as a
covariate (A.5). This is the H3-side response to the same confound — the conflation of the
effects of work and experience with the accumulated natural depreciation over the period
since exit — for which Hypothesis H1 demands strict control of "Time Since Retirement."
Chronological age is not used as an assignment condition. This is a design choice
corresponding to this paper's core distinction (Section 4.4). Age is recorded and used as a
covariate and in auxiliary analyses (A.5 below). Furthermore, to mitigate the confounding of
experience and age to the extent possible, recruitment deliberately strives to include older
persons with little experience (career changers and returners) and younger persons with long
Working Paper | Ageless Management in the AI Era 167
experience (early specializers). If filling both cells proves difficult, this is reported and stated
explicitly as a limit on the separability of experience and age. In addition, for all participants,
a proxy measure of crystallized intelligence (Gc) (vocabulary and knowledge tests), a selfreport
scale of metacognition, and experience using generative AI (frequency and purposes)
are measured in advance and used as covariates. AI-use experience can correlate with
experience level (the usage gap of Section 4.5) and is therefore especially important as a
control variable.
A.2 Task Construction and Error Embedding
The task materials are business proposals and business documents created with generative AI
(market analyses, business plans, customer proposals, and the like), formatted to match the
practice of the task domain. In each document, errors and risks known to the experimenters
are embedded by type. Four types are embedded.
First, factual errors. These are errors verifiable by checking against information external to
the document, such as misstated figures, references to nonexistent sources, and confusions of
proper nouns. Second, contextual risks. These are contents that are coherent on the face of
the document but that cause problems in light of the practical context of the domain —
transaction terms contrary to industry practice, premises that do not hold for the customer
segment in question, measures whose same-pattern failures are known from the past, and the
like. Third, ethical risks. These are problems requiring value judgment, such as potential
conflicts with laws and norms, unjust burdens on stakeholders, and overlooked conflicts of
interest. Fourth, anachronisms. These are premises that were once valid but no longer hold —
reliance on abolished institutions, designs premised on the specifications of a previous
generation of technology, outdated images of changed consumer behavior. The detection of
anachronisms is a window for observing in which direction the experience of different
historical environments (the foundation of the generational decorrelation of Definition 5)
operates; because old experience may aid detection and, conversely, old knowledge may
generate misplaced confidence about the present, the direction is measured without being
fixed in advance.
Blind Generation of Embedded Errors — Eliminating Experimenter Bias
The creation of embedded errors is subjected to procedures that eliminate the experimenters'
preconception bias. As organized in Section 4.2, contextual information can systematically
distort the judgments of experts themselves (Dror, Charlton & Péron 2006). This finding
applies also to the experimenters who create the tasks: errors devised at the experimenters'
desks would be biased toward types congruent with the experimenters' own views of the
industry and of failure, and that bias could work either for or against the experienced
participants. The response consists of two steps.
First, part of the errors are incorporated as tasks blind-generated from an empirical failurecase
dataset of accidents, scandals, and business failures that actually occurred — accident
Working Paper | Ageless Management in the AI Era 168
investigation reports, administrative sanctions and recall cases, public records of
bankruptcies and withdrawals, and the like. That is, a generation team independent of the
experimenters converts the structure of the failure cases (what was believed to be sound, and
what in fact broke down) into errors and risks within the task documents. The generation
process and the conduct and analysis of the experiment are separated in personnel, and the
experimenters and evaluators are not informed which error derives from which case until
analysis is complete. Second, the full list of errors, the generation procedure, and the
correspondence table to the source cases are preregistered with timestamps before the
experiment begins (a sealed registration kept nonpublic until analysis is complete). This
makes ex-post substitution of errors or convenient reinterpretation procedurally impossible.
Confirmation-Congruent and Deviating Types — Simultaneous Test of Boundary
Condition (i) of Proposition 4
Furthermore, part of the embedded errors — above all the contextual risks and
anachronisms — are constructed in two classes according to their congruence with auditors'
industry conventional wisdom. The confirmation-congruent type consists of plausible errors
in the direction of industry conventional wisdom or the success experiences of the auditors'
generation (contents that follow the conventional wisdom but break down in the given
context); the deviating type consists of errors in the direction contrary to that wisdom. The
validity of the classification is confirmed by having a panel of practitioners in the domain
who do not participate in the experiment independently rate, for each statement, "whether it
is natural in light of industry conventional wisdom," and by its meeting a preregistered
agreement criterion.
This classification is the operation for the simultaneous test of boundary condition (i) of
Proposition 4 — that when AI output is congruent with the auditor's own past success
experiences and industry conventional wisdom, experience can, through the synergy of
confirmation bias and automation bias, actually impair detection. If boundary condition (i) is
real, the detection-rate advantage of the experienced will be observed for deviating errors
and will shrink, vanish, or reverse for confirmation-congruent errors. Conversely, if even for
confirmation-congruent errors the detection rate of the experienced does not fall below that
of the inexperienced, boundary condition (i) is empty and the body of Proposition 4 is
supported in a stronger form (the asymmetric structure of Proposition 4's refutation
condition). This experience × congruence interaction test is preregistered as a secondary
analysis (A.5).
Embedding Innovative Proposals — Measuring the Over-Rejection Rate
In addition to errors, a small number of "correct but counter-conventional innovative
proposals" are embedded in the task documents. These are non-error items for measuring the
boundary at which experience goes beyond the detection of errors and turns into the
rejection of sound deviation (over-auditing), and they are the measurement target of
Working Paper | Ageless Management in the AI Era 169
dependent variable (iv), the over-rejection rate — the rate at which these proposals were
erroneously rejected as problems.
The certification of "innovative but correct" is made, by prior agreement before the
experiment begins, by an expert panel independent of both the experimenters and the
outcome evaluators. The panel certifies, under a preregistered agreement criterion
(unanimity or a prespecified supermajority), that each candidate proposal both (a) deviates
from industry conventional wisdom and (b) is nonetheless sound from the standpoints of
fact, logic, and practice, and candidates that fail the criterion are not used. To the extent
possible, the proposals are constructed on the basis of real instances that were initially
dismissed as contrary to conventional wisdom but whose validity was later established, and
the corresponding provenance table is also included in the sealed registration. No theoretical
prediction is placed on the over-rejection rate; it is treated as an exploratory indicator
(Section 9, Hypothesis H3).
The Fluency Manipulation — Simultaneous Test of Boundary Condition (iii) of
Proposition 4
The task documents are prepared in two versions, a high-fluency version and a low-fluency
version, identical in content and differing only in stylistic polish. The two versions keep the
propositional content — including the embedded errors and the innovative proposals (facts,
figures, logical structure, and the location of risks) — completely identical, and the difference
is confined to the level of style. The high-fluency version is a low-processing-resistance style:
grammatically well-formed, with plain vocabulary and syntax and smooth paragraph
transitions. The low-fluency version is a style that, while preserving the propositional
content, raises processing resistance through more complex syntax, uneven paragraph
construction, and redundant or stilted expression. The versions are produced by combining
generative-AI style transformation with human copyediting, and the identity of the two
versions' propositional content is confirmed by an independent checker through collation
against a correspondence table.
As a manipulation check, a panel of raters who do not participate in the experiment rates
the subjective fluency (readability and ease of processing) of each document, confirming that
the high-fluency version exceeds the low-fluency version by at least a preregistered criterion.
Objective readability indicators such as sentence length and syntactic complexity are also
reported, and document pairs that fail the manipulation check are replaced. In assignment,
fluency is treated as a factor (between-participants or within-participants) or as a covariate,
and combinations of documents and versions are balanced across conditions. The theoretical
grounding of this manipulation lies in the cognitive-psychology finding that processing
fluency inflates judgments of truth (Reber & Schwarz 1999; Alter & Oppenheimer 2009). If
boundary condition (iii) of Proposition 4 — that high-fluency output can, through the effect of
processing fluency, raise the threshold of the auditor's cognitive sense of unease and lower
detection performance — is real, detection rates will fall in the high-fluency version, with the
Working Paper | Ageless Management in the AI Era 170
fall predicted above all in the detection of contextual risks. Conversely, a result in which
detection performance does not fall even when fluency is manipulated empties boundary
condition (iii) and leaves the body of Proposition 4 standing in a stronger form (Proposition
4's refutation condition). The corresponding tests are preregistered as secondary analyses (A.
5).
In embedding, the number of errors of each type and the ratio of confirmation-congruent to
deviating errors are balanced across documents, and placement is randomized so that errors
cannot be detected from cues of position or formatting. Sufficient sound passages containing
no errors are also secured, so that false alarms — participants flagging correct passages as
errors — can be measured. This separates the detection rate (hit rate) from the false-alarm
rate and enables analysis within the framework of signal detection theory (SDT) — the
separate estimation of sensitivity d′ and response criterion c (A.5). The validity of the
documents and the embedded errors is confirmed by prior review by a panel of practitioners
in the domain who do not participate in the experiment, and embedded errors not detected
in that review are replaced.
Table A1 Correspondence between embedded error types and the abilities required for
detection (theoretical predictions)
Error type Definition
Ability required for detection
(theoretical correspondence)
Predicted compression
under AI assistance
Factual error Errors verifiable by
checking against
information external to
the document (figures,
sources, proper nouns)
Procedural skills of collation and
search. Relatively large Gf-type
component
Compressed (AI
assistance takes over
search and collation, and
the experience gap
shrinks)
Contextual risk Contents coherent
within the document
but problematic in light
of the practical context
Interaction of long-term domain
experience × Gc-type abilities
(the formation mechanism that
Proposition 4 asserts for the
detection capacity of Definition
4)
Not compressed (the
main effect of
experience persists even
under assistance) — the
direct test point of
Proposition 3
Ethical risk Potential conflicts with
laws and norms, unjust
allocation of interests,
overlooked conflicts of
interest
Gc-type abilities, metacognition,
and value judgment. Experience
contributes as case knowledge
Not compressed (though
predicted to be less
experience-specific than
contextual risk)
Anachronism Reliance on premises
once valid but no longer
holding
Experience of changing
historical environments (the
foundation of Definition 5).
However, overconfidence in old
knowledge may operate in
reverse
Direction not specified in
advance (exploratory
measurement) — the
observation window for
generational
decorrelation
Proposals deviating
from industry
conventional wisdom
The ability to judge the validity
of content beyond conformity to
conventional wisdom. The
No prediction placed
(exploratory
measurement) — the
Working Paper | Ageless Management in the AI Era 171
Innovative
proposal (nonerror
item)
but certified as sound
by an independent
expert panel by prior
agreement
observation window for the
boundary at which experience
turns into over-auditing
measurement target of
dependent variable (iv),
the over-rejection rate
Note: The correspondence in "ability required for detection" is this paper's theoretical prediction and is itself the object
of the experimental test. Types for which the prediction fails are reported as information delimiting the scope of
application of Definition 4. Contextual-risk and anachronism errors are constructed in two classes: confirmationcongruent
(errors in line with the auditors' industry conventional wisdom) and deviating (errors contrary to that
wisdom) (see the main text). In addition, all task documents are prepared in two versions, a high-fluency version and a
low-fluency version, identical in content and differing only in stylistic polish (see "The Fluency Manipulation" in the
main text).
A.3 Procedure
After determination of experience level, participants are randomly assigned to either the AIassisted
condition or the unassisted condition (stratified randomization: stratified by
experience level, age band, and AI-use experience). Participants in the assisted condition may
freely use generative-AI tools during the verification work, and their usage logs (prompts and
responses) are recorded. Participants in the unassisted condition work without AI tools, with
ordinary reference materials (whether search is included depends on the design of factualerror
detection and is fixed in the preregistration).
The task consists of multiple documents, and the order of document presentation is
randomized. The instruction to participants is: "This document may contain errors or
problems. Point out all problems you find, and attach a correction proposal to each." The
types and number of embedded errors are not disclosed. Working time is recorded with an
upper limit, and time itself is treated as a dependent variable (a proxy for the effort invested
in verification). After all tasks are completed, confidence ratings (confidence for each flag) are
collected, and in the AI-assisted condition a self-assessment of dependence on the assistance.
The correspondence between confidence and correctness serves as an indicator of
metacognitive calibration and is used in examining the interaction term (metacognition) of
Proposition 4.
A.4 Evaluator Blinding
The detection rates of dependent variables (i) and (ii) can be scored mechanically because the
embedded errors are known (only the matching of flags to embedded locations is judged,
independently, by two blinded judges, with disagreements resolved by discussion). The overrejection
rate of dependent variable (iv) can likewise be scored mechanically by whether a
"problem" flag was placed on the location of an embedded innovative proposal (the matching
procedure is shared with (i) and (ii)). The quality of correction proposals, dependent variable
(iii), is assessed by blinded external evaluation. The evaluators are external experts with
practical experience in the domain, and they are presented only with the proposal texts, with
the participants' experience level, age, AI-assistance condition, and names concealed. To
prevent conditions from being inferred from style and the like, the proposal texts are
Working Paper | Ageless Management in the AI Era 172
presented after standardization of presentation (correction of typographical errors and
unification of formatting). Each proposal is rated independently by two or more evaluators,
and inter-rater reliability (intraclass correlation coefficient) is reported. If reliability falls
below the preregistered criterion, the rating procedure is revised and re-rating is conducted.
The evaluation axes are three — accuracy of problem apprehension, feasibility of the
correction, and attention to side effects — and the definition and rating scale of each axis are
stated explicitly in the preregistration.
A.5 Testing Plan
The main analysis conducts, for each of dependent variables (i), (ii), and (iii), an analysis of
variance (ANOVA) of experience level (2) × AI assistance (2), testing main effects and the
interaction. The core of the hypothesis is the contrast of interactions. The predictions are as
follows. For the detection rate of surface factual errors (i), the experience × AI-assistance
interaction is significant and the experience gap shrinks under AI assistance (consistent with
the compression evidence of Section 4.1). For the detection rate of contextual and practical
risks (ii) and the quality of correction proposals (iii), the main effect of experience persists
and no interaction-driven shrinkage is observed — that is, the "absence or smallness of the
interaction" in (ii) and (iii) constitutes the supporting evidence for Proposition 3. Because a
claim of an absent interaction depends on statistical power, mere non-significance is not
relied on; an equivalence test against a preregistered smallest effect size of interest (or a
report of estimation precision) is used in combination. In addition, the analysis plan states
explicitly the application of signal detection theory (SDT). An analysis relying on the detection
rate (hit rate) alone cannot, in principle, distinguish differences in true sensitivity from
differences in response bias — a shift of the judgment criterion toward suspecting everything.
The apparently higher detection rate of the experienced may reflect higher discriminability,
or merely a lower threshold for flagging. Therefore, an SDT analysis incorporating the falsealarm
rate estimates sensitivity d′ and response criterion c separately (Macmillan & Creelman
2005; Hautus, Macmillan & Creelman 2022). Whether the advantage of the experienced is due
to detection sensitivity (d′) or to flagging assertiveness (a low c) is discriminated by this
separate estimation.
Dependent variable (iv), the over-rejection rate, is not subjected to hypothesis testing; it is
an exploratory analysis descriptively reporting the rate by condition (experience level × AI
assistance) with confidence intervals. The over-rejection rate is conceptually a kind of false
alarm, but unlike false alarms on sound passages in general, it is confined to rejections of
non-error items defined on the content dimension of deviation from conventional wisdom.
The two are reported separately, so that the general strictness of the experienced participants'
response criterion can be distinguished from selective rejection of deviation from
conventional wisdom. Within the SDT framework, dependent variable (iv), the over-rejection
rate, is not an independent phenomenon but a manifestation of the response criterion c
shifting toward the conservative (suspecting) side, and it is analyzed in the same framework
Working Paper | Ageless Management in the AI Era 173
as the separate estimation of sensitivity and criterion for (i) and (ii). That is, the general bias
of the criterion estimated from false alarms on sound passages in general and the increment
of selective rejection on deviation items are contrasted within the same model. This
increment is the measured quantity of the boundary at which experience turns into overauditing.
As the secondary analysis corresponding to boundary condition (i) of Proposition 4, the
congruence classification of errors (confirmation-congruent / deviating) is added as a withinparticipants
factor, and the interaction of experience level × congruence (and the three-way
interaction adding AI assistance) is tested. The contrast of interest is the difference in
detection rates between the experienced and the inexperienced for confirmation-congruent
errors. That this difference is significantly smaller than that for deviating errors (including
vanishing or reversal) is consistent with boundary condition (i); that the detection rate of the
experienced does not fall below that of the inexperienced even for confirmation-congruent
errors empties boundary condition (i) (Proposition 4's refutation condition). This contrast and
its adjudication criteria are stated explicitly in the preregistration.
As the secondary analysis corresponding to boundary condition (iii) of Proposition 4, the
fluency of the task documents (high-fluency / low-fluency) is added as a factor (or, in designs
treating it as a covariate, its coefficient), and the main effect of fluency and the interaction of
experience level × fluency (and the three-way interaction adding AI assistance) are tested.
Within the SDT framework, whether the effect of fluency appears as a decline in sensitivity d′
or as a conservative shift of the response criterion c — a rise in the threshold of suspicion —
is reported separately. The mechanism assumed by boundary condition (iii) (a rise in the
threshold of cognitive unease) is predicted to appear primarily as a movement of c, but this
prediction is itself an object of the test. A result in which detection performance does not
decline even when fluency is manipulated empties boundary condition (iii) (Proposition 4's
refutation condition). The adjudication criteria are stated explicitly in the preregistration.
As covariate analysis, an analysis of covariance is conducted entering the Gc proxy
measure, the metacognition scale, AI-use experience, and years elapsed since leaving practice
(A.1). In the auxiliary analysis corresponding to Proposition 4, years of experience and
chronological age are entered into the same model, and the partial effect of chronological age
after controlling for years of experience is estimated. The prediction of Proposition 4 is that
this partial effect is indistinguishable from zero (and that the effect of years of experience
persists). Conversely, if chronological age independently predicts detection performance even
after controlling for experience, this paper's reattribution thesis (reattribution from age to
experience) is refuted. Because this analysis depends on the success of the experience–age
confound mitigation described in A.1, the correlation between the two variables and the
variance inflation factor are always reported.
The required sample size is fixed by a power analysis based on the smallest effect size of
interest (set in advance from the range of effect sizes in prior compression studies and from
Working Paper | Ageless Management in the AI Era 174
practically meaningful detection-rate differences) and on the power requirements of the
interaction test. Detecting an interaction typically requires a larger sample than detecting a
main effect, and this point is treated explicitly in the power analysis. This appendix does not
provisionally fix a specific sample size. The number is fixed at preregistration as the result of
the power analysis and included in the registered content.
Finally, the limitations of this protocol are recorded. First, the experience factor is a quasiexperimental
factor, and differences between the experienced and junior participants may be
contaminated by selection and cohort factors other than experience. Covariate control
removes this only partially. Second, the embedded errors are errors known to the
experimenter side. Blind generation from the empirical failure-case dataset and sealed
registration (A.2) mitigate the bias of error selection toward the experimenters'
preconceptions, but the detection of the errors most dangerous in practice — "errors no one
had anticipated," never once manifested in the past — still cannot be measured. Third, results
from implementation in a single domain do not immediately generalize to other domains,
and replication with changed domains is necessary. Fourth, because this design is a crosssectional
2×2 factorial design, it can detect only compression by contemporaneous assistance,
and it cannot, in principle, detect longitudinal acquisition acceleration — the pathway in
which the acquisition of verification capacity accelerates under an AI-use environment and
effectively substitutes for the effect of accumulated experience (the second clause of
Proposition 3's refutation condition). Testing this pathway requires separate longitudinal
measurement (Section 9.4). These limitations are stated alongside the results report as the
frame for interpreting the experimental results.
A.6 Supplementary Note — Team Assignment for Hypothesis H2 (Stratified
Randomization)
The main object of this appendix is H3, but the skeleton of the assignment procedure for the
team-level randomization of Hypothesis H2 (the multigenerational team-outcome hypothesis)
is also appended here (Section 9.1). What H2 seeks to identify is the effect of heterogeneity in
age and experience; if the composition of other demographic attributes — gender, ethnicity,
cultural background, and the like — is skewed across teams, the mixed/homogeneous contrast
is confounded with these diversity dimensions, and the observed effect can no longer be
attributed to the axis of age and experience. Participants are therefore stratified by attributes
such as gender, ethnicity, and cultural background, and within each stratum randomly
assigned to the team-composition condition (mixed / homogeneous) and the AI-use condition
(with / without), thereby balancing the composition of these attributes across teams and
conditions (stratified randomization). For attributes for which sample-size constraints leave
strata underfilled and full stratification infeasible, the attribute in question is controlled at
the analysis stage as a covariate. In either case, post-assignment attribute balance
(standardized differences across conditions) is reported. The list and definitions of the
attributes used for stratification, the criteria for switching from stratification to covariate
Working Paper | Ageless Management in the AI Era 175
control, and the reporting format for balance indicators are stated explicitly in the
preregistration.
Appendix B Detailed Table of the Institutional Comparison
Table A2 lists the details of the institutions compared in Section 7 (Table 6) — official statute
names, statute numbers, effective dates, key points of content, and sources. Entries are
limited to what was checked against primary statutes and administrative materials or
materials of public research institutions. This table is a description of institutions, not legal
advice, and details of application are governed by the latest content of each statute and
circular.
Table A2 Details of institutions related to older-age and youth work (by jurisdiction, as of August
2026)
Jurisdiction
and
institution
Official statute
name and basis
Effective
date, etc. Key points of content Sources
Japan — Act
on
Stabilization
of
Employment
of Elderly
Persons
Act on Stabilization
of Employment of
Elderly Persons,
etc. (Kลnenreishatล
no Koyล no
Antei-tล ni kansuru
Hลritsu; Act No. 68
of 1971). The
age-70 measures
were newly
established by the
amendment under
Act No. 14 of 2020
The measures
for securing
work
opportunities
up to age 70
(Article 10-2)
took effect on
April 1, 2021
Article 8: the mandatory
retirement age (teinen) may
not be set below 60 (with
exceptions). Article 9: legal
obligation to take
employment-securing
measures up to age 65
(raising the retirement age,
continued employment, or
abolition of the retirement
age). Article 10-2: duty to
make efforts to secure work
up to age 70. Of the five
options, outsourcing
contracts (option 4) and
social-contribution projects
(option 5) are the
entrepreneurship-support
measures (Article 10-2) (nonemployment
type).
Introduction requires
consent procedures with a
majority labor union or
equivalent
Ministry of
Health, Labour
and Welfare
(amendment
overview;
statutes
database)
Japan —
revision of the
in-work oldage
pension
offset
Act Partially
Amending the
National Pension
Act, etc. for
Strengthening the
Functions of the
Pension System in
Effective April
1, 2026
Suspended amount = (basic
monthly amount + totalremuneration-
equivalent
monthly amount − threshold)
÷ 2. The threshold is raised
from 510,000 yen (the actual
FY2025 figure) to 620,000 yen
Japan Pension
Service (special
page;
calculation
method),
Ministry of
Health, Labour
Working Paper | Ageless Management in the AI Era 176
Light of Social and
Economic Changes
(Act No. 74 of 2025;
enacted June 2025)
(the statutory value at 2025
wage levels). Through wage
indexation, the actual figure
applied in FY2026 is 650,000
yen. Employed persons aged
70 and over bear no
premiums (loss of insured
status), but the benefit
suspension (offset) continues
to apply
and Welfare (act
overview)
Japan —
minors
provisions of
the Labor
Standards Act
Labor Standards
Act (Rลdล Kijun Hล;
Act No. 49 of 1947),
Chapter 6
"Minors" (Articles
56–63)
Enacted in
1947 (the
current
provisions
reflect
subsequent
amendments)
Article 56: prohibition of
employing children until the
end of the first March 31 after
they reach age 15 (exceptions
for non-industrial businesses:
age 13 or over, light labor,
permission of the
administrative agency, etc.;
for film and theater, even
under 13 on a permit basis).
Article 57: certificates of age,
etc. Article 58: prohibition of
labor contracts concluded on
behalf of minors by persons
with parental authority.
Article 59: minors'
independent claim to wages.
Article 60: prohibition in
principle of overtime and
holiday work. Article 61:
prohibition in principle of
night work (10 p.m. to 5 a.m.).
Article 62: restrictions on
employment in dangerous
and harmful work. Article 63:
prohibition of underground
labor
Hyogo Labour
Bureau
(commentary on
the minors
provisions)
Japan —
employee
status in
internships
Article 9 of the
Labor Standards
Act (definition of
"worker");
administrative
circular Kihatsu
No. 636 of
September 18, 1997
The circular
was issued in
1997
Where the profit or effect of
the work accrues to the
establishment and a
relationship of use and
subordination is recognized,
students also qualify as
workers (the Labor Standards
Act, the Minimum Wage Act,
and workers' accident
compensation insurance
apply). Where participation is
observational or experiential
and not subject to direction
and control, the student does
not qualify as a worker. The
judgment turns on substance
Nagano Labour
Bureau
materials
Working Paper | Ageless Management in the AI Era 177
(direction and control,
attribution of output,
attendance management,
etc.), not on the label
Japan —
Freelance Act
Act on Ensuring
Proper
Transactions
Involving Specified
Entrusted Business
Operators (Tokutei
Jutaku Jigyลsha ni
kakaru Torihiki no
Tekiseika-tล ni
kansuru Hลritsu;
Act No. 25 of 2023)
Effective
November 1,
2024
Obligations of commissioning
businesses: clear indication
of transaction terms in
writing or equivalent; setting
a remuneration due date
(within 60 days of receipt of
the deliverable) and payment
by that date; prohibition, in
continuing outsourcing, of
refusal of receipt, reduction
of remuneration, returns,
beating down of prices, etc.;
accurate display of
recruitment information;
accommodation for
balancing work with
childcare and caregiving;
establishment of harassmentresponse
systems; 30 days'
advance notice of mid-term
termination. Jurisdiction: the
Japan Fair Trade
Commission, the Small and
Medium Enterprise Agency,
and the Ministry of Health,
Labour and Welfare
Cabinet
Secretariat
(policy portal),
Japan Fair Trade
Commission,
Small and
Medium
Enterprise
Agency
Japan —
expansion of
special
enrollment in
workers'
accident
compensation
insurance
Article 33 et seq. of
the Workers'
Accident
Compensation
Insurance Act
(special
enrollment). The
expansion is by a
ministerial order
partially amending
the Enforcement
Regulations of the
Act, etc.
Effective
November 1,
2024 (the
same day as
the Freelance
Act)
Freelancers engaged in
specified entrusted business
(BtoB outsourcing) become
eligible for special
enrollment (BtoC work is also
covered). In effect, virtually
all freelancers may enroll
voluntarily. Premiums are
borne entirely by the
individual; the Class II special
enrollment premium rate is
the basic daily benefit
amount × 365 × 3/1000 (0.3%).
The basic daily benefit
amount is chosen from 16
steps between 3,500 and
25,000 yen. Procedures run
through special enrollment
associations
Ministry of
Health, Labour
and Welfare
(leaflet;
premium rate
table)
United States
— earnings
test reform
Senior Citizens'
Freedom to Work
Signed April 7,
2000
Abolished the earnings test
for work at and after FRA
(full retirement age). It
U.S. Congress
(statute text),
SSA (official
Working Paper | Ageless Management in the AI Era 178
Act of 2000 (Public
Law 106-182)
remains before FRA: in years
before the year of reaching
FRA, $1 is withheld for every
$2 above the lower exempt
amount; in the year of
reaching FRA (up to the
month before the month of
attainment), $1 for every $3
above the upper exempt
amount. The 2026 exempt
amounts are $24,480 per year
(lower) and $65,160 per year
(upper). Withheld amounts
are effectively recovered
through benefit
recomputation after FRA
exempt-amount
table)
United States
— ADEA
Age Discrimination
in Employment Act
of 1967 (29 U.S.C.
§621 et seq.)
Enacted in
1967. The
1986
amendment
removed the
upper age
limit of
protection
Prohibits age discrimination
in employment (hiring,
discharge, pay, promotion,
etc.) against workers aged 40
and over. With the removal of
the upper age limit,
mandatory retirement is
unlawful in most occupations
(with limited exceptions such
as pilots)
EEOC
Germany —
abolition of
the earnings
limit
Eighth Act
Amending Book IV
of the Social Code
(Achtes Gesetz zur
Änderung des
Vierten Buches
Sozialgesetzbuch
— 8. SGB IV-ÄndG)
Effective
January 1,
2023
Completely abolished the
earnings limit
(Hinzuverdienstgrenze)
during receipt of early oldage
pensions (a permanent
measure applying to all
recipients). Before abolition,
the 2022 limit was 46,060
euros per year (the COVID
special level). For disability
pensions, not abolition but a
transition to a dynamic wagelinked
limit
Official FAQ of
the German
Pension
Insurance (DRV)
EU —
Employment
Equality
Directive
Council Directive
2000/78/EC (the
general framework
directive for equal
treatment in
employment and
occupation)
Adopted
November 27,
2000
Prohibits direct and indirect
discrimination in
employment and occupation
on grounds of religion or
belief, disability, age, or
sexual orientation. Article 6
permits justification of
differences of treatment on
grounds of age, conditional
on legitimate aims such as
employment policy and on
appropriate and necessary
means (an exception for age
EUR-Lex,
European
Commission
Working Paper | Ageless Management in the AI Era 179
alone; the basis provision for
the case law of the Court of
Justice of the EU on the
permissibility of mandatory
retirement)
EU — Young
Workers
Directive
Council Directive
94/33/EC (directive
on the protection
of young people at
work)
Adopted June
22, 1994
Covers persons under 18 and
prohibits in principle the
labor of children (under 15 or
in compulsory education).
Exceptions: cultural, artistic,
sporting, and advertising
activities (permit-based);
combined work/training-type
engagement for those 14 and
over; light work for those 14
and over; light work for
limited weekly hours for
those 13 and over. Regulates
the prohibition of dangerous
work for young people,
working time, night work,
and rest
EUR-Lex
(summary), EUOSHA
ILO —
Convention
No. 138
Minimum Age
Convention, 1973
(No. 138)
Adopted in
1973. Ratified
by Japan in
2000
The minimum age shall be
not less than the age of
completion of compulsory
schooling and, in any case,
not less than 15 (Article 2(3)).
The developing-country
exception was initially 14
(Article 2(4)). Light work at
ages 13–15 may be permitted
by national law (Article 7;
developing-country exception
12–14). Hazardous work at 18
(Article 3; an exception at 16
conditional on protection and
training = Article 3(3))
ILO (convention
text; ILO Office
in Japan)
ILO —
Convention
No. 182
Worst Forms of
Child Labour
Convention, 1999
(No. 182)
Adopted in
1999.
Universal
ratification by
all member
states
achieved on
August 4,
2020. Ratified
by Japan in
2001
Obligates immediate
measures for the prohibition
and elimination of the "worst
forms of child labour" for
those under 18 (slavery,
forced labor, and trafficking;
child soldiers; sexual
exploitation; illicit activities;
hazardous work). With
Tonga's ratification (the
187th), the first universal
ratification in ILO history.
The United States has not
ratified No. 138 but has
ratified No. 182
ILO (official
announcement)
Working Paper | Ageless Management in the AI Era 180
Singapore —
RRA
Retirement and Reemployment
Act
63/68 from
July 1, 2022;
64/69 from
July 1, 2026
A two-tier structure of a
statutory retirement age
(prohibition of compelled
retirement below that age)
and a re-employment-offer
obligation age. The
government has stated a
policy of raising these to
65/70 by 2030. An obligation
to offer re-employment, not
an obligation to retain
employment
Law firms and
HR professional
media (final
confirmation
against MOM
primary sources
remains
outstanding)
Korea —
retirementage
extension
act
Act on Prohibition
of Age
Discrimination in
Employment and
Elderly
Employment
Promotion
(amended 2013;
the amendment
passed in June
2013)
Effective 2016
for
workplaces
with 300 or
more
employees
and public
institutions;
2017 for those
with fewer
than 300
Obligates setting the
mandatory retirement age at
60 or above. The same act
also prohibits employment
discrimination on grounds of
age. Extension of the
statutory retirement age to 65
is under discussion between
labor and management and
in the National Assembly
JILPT, JILAF
Note: Entries reflect what was checked as of August 2026. Amounts, ages, and the like may change with revisions to
each institution. It is stated explicitly that, for the 2026 amendment of Singapore's RRA alone, final confirmation
against primary sources (the competent ministry) remains outstanding. This table is a description of institutions, not
legal advice.
References
These references include the grey literature and industry surveys (survey reports by NGOs,
membership organizations, and companies, etc.) whose evidence grade is stated explicitly in
the main text. Unrefereed preprints and technical reports (including working papers) are
marked as such at the end of each entry. Japanese statutes and Japanese-language official
materials are collected under the subheading "Primary Legal Sources and Official Materials"
at the end.
AARP Research (Perron, R.). (2026). Foresight 50+ Omnibus Survey, Wave 3: AI and workers age 50-plus
(fielded March 12–16, 2026; 1,015 U.S. workers aged 50 and over). Washington, DC: AARP. (grey
literature; membership-organization survey)
Accenture, Disability:IN, & AAPD. (2018). Getting to Equal: The Disability Inclusion Advantage. Accenture.
(grey literature; corporate survey)
Aigner, D. J., & Cain, G. G. (1977). Statistical theories of discrimination in labor markets. Industrial and
Labor Relations Review, 30(2), 175–187.
Akerlof, G. A. (1970). The market for "lemons": Quality uncertainty and the market mechanism. Quarterly
Journal of Economics, 84(3), 488–500. doi:10.2307/1879431
Alter, A. L., & Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation.
Personality and Social Psychology Review, 13(3), 219–235. doi:10.1177/1088868309341564
Working Paper | Ageless Management in the AI Era 181
Ayalon, L. (2026). Intergenerational relations in the workforce in the age of artificial intelligence: Where
do we go from here? International Psychogeriatrics, Article 100235. doi:10.1016/j.inpsyc.2026.100235
Backes-Gellner, U., & Veen, S. (2013). Positive effects of ageing and age diversity in innovative companies –
large-scale empirical evidence on company productivity. Human Resource Management Journal, 23(3),
279–295. doi:10.1111/1748-8583.12011
Baltes, P. B., & Baltes, M. M. (1990). Psychological perspectives on successful aging: The model of selective
optimization with compensation. In P. B. Baltes & M. M. Baltes (Eds.), Successful Aging: Perspectives
from the Behavioral Sciences (pp. 1–34). Cambridge University Press. doi:10.1017/
CBO9780511665684.003
Barney, J. (1991). Firm resources and sustained competitive advantage. Journal of Management, 17(1), 99–
120.
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcฤฑ, Ö., & Mariman, R. (2025). Generative AI without
guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National
Academy of Sciences, 122(26), e2422633122. doi:10.1073/pnas.2422633122 (Correction: doi:10.1073/pnas.
2518204122 — correction of author affiliation only)
Bick, A., Blandin, A., & Deming, D. J. (2024). The rapid adoption of generative AI. NBER Working Paper No.
32966. Cambridge, MA: National Bureau of Economic Research. (unrefereed working paper; updated
figures published by the Federal Reserve Bank of St. Louis, On the Economy, November 2025)
Bonsang, E., Adam, S., & Perelman, S. (2012). Does retirement affect cognitive functioning? Journal of
Health Economics, 31(3), 490–501. doi:10.1016/j.jhealeco.2012.03.005
Börsch-Supan, A., & Weiss, M. (2016). Productivity and age: Evidence from work teams at the assembly
line. The Journal of the Economics of Ageing, 7, 30–42. doi:10.1016/j.jeoa.2015.12.001
Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at work. Quarterly Journal of Economics,
140(2), 889–942. doi:10.1093/qje/qjae044
Budzyล, K., et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy:
A multicentre, observational study. The Lancet Gastroenterology & Hepatology, 10(10). doi:10.1016/
S2468-1253(25)00133-5
Burmeister, A., Wang, M., & Hirschi, A. (2020). Understanding the motivational benefits of knowledge
transfer for older and younger workers in age-diverse coworker dyads: An actor–partner
interdependence model. Journal of Applied Psychology, 105(7), 748–759. doi:10.1037/apl0000466
Carlson, M. C., et al. (2008). Exploring the effects of an "everyday" activity program on executive function
and memory in older adults: Experience Corps. The Gerontologist, 48(6), 793–801.
Carlson, M. C., et al. (2009). Evidence for neurocognitive plasticity in at-risk older adults: The Experience
Corps program. Journals of Gerontology: Series A, Medical Sciences, 64A(12), 1275–1282.
Carlson, M. C., et al. (2015). Impact of the Baltimore Experience Corps Trial on cortical and hippocampal
volumes. Alzheimer's & Dementia, 11(11), 1340–1348.
Carstensen, L. L., Isaacowitz, D. M., & Charles, S. T. (1999). Taking time seriously: A theory of
socioemotional selectivity. American Psychologist, 54(3), 165–181. doi:10.1037/0003-066X.54.3.165
Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of
Educational Psychology, 54(1), 1–22. doi:10.1037/h0046743
Clark, A., & Chalmers, D. J. (1998). The extended mind. Analysis, 58(1), 7–19. doi:10.1093/analys/58.1.7
Coe, N. B., von Gaudecker, H.-M., Lindeboom, M., & Maurer, J. (2012). The effect of retirement on cognitive
functioning. Health Economics, 21(8), 913–927. doi:10.1002/hec.1771
Dawson, W. D., Smith, E., Booi, L., Mosse, M., Lavretsky, H., Reynolds, C. F., III, ... Eyre, H. A. (2022). Investing
in late-life brain capital. Innovation in Aging, 6(3), igac016. doi:10.1093/geroni/igac016
Working Paper | Ageless Management in the AI Era 182
Dell'Acqua, F. (2022–2023). Falling Asleep at the Wheel: Human/AI Collaboration in a Field Experiment on
HR Recruiters. Working paper, Harvard Business School / Laboratory for Innovation Science.
(unrefereed working paper; not published in a refereed journal, checked against a mirror-distributed
version)
Dell'Acqua, F., McFowland, E., III, Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L.,
Candelon, F., & Lakhani, K. (2023). Navigating the jagged technological frontier: Field experimental
evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School
Working Paper 24-013 / SSRN 4573321. (unrefereed working paper; the figures in this paper are from
the WP version; the Organization Science version, doi:10.1287/orsc.2025.21838, was confirmed for
existence only)
Dror, I. E., Charlton, D., & Péron, A. E. (2006). Contextual information renders experts vulnerable to making
erroneous identifications. Forensic Science International, 156(1), 74–78. doi:10.1016/j.forsciint.
2005.10.017
Dufouil, C., Pereira, E., Chêne, G., Glymour, M. M., Alpérovitch, A., Saubusse, E., Risse-Fleury, M., Heuls, B.,
Salord, J.-C., Brieu, M.-A., & Forette, F. (2014). Older age at retirement is associated with decreased risk
of dementia. European Journal of Epidemiology, 29(5), 353–361. doi:10.1007/s10654-014-9906-3
Eibich, P. (2015). Understanding the effect of retirement on health: Mechanisms and heterogeneity. Journal
of Health Economics, 43, 1–12.
El Morr, C., Kundi, B., Mobeen, F., Taleghani, S., El-Lahib, Y., & Gorman, R. (2024). AI and disability: A
systematic scoping review. Health Informatics Journal, 30(3). doi:10.1177/14604582241285743
Fried, L. P., et al. (2004). A social model for health promotion for an aging population: Initial evidence on
the Experience Corps model. Journal of Urban Health, 81(1), 64–78.
Fujiwara, Y., et al. (2023). The relationship between working status in old age and cause-specific disability
in Japanese community-dwelling older adults with or without frailty: A 3.6-year prospective study.
Geriatrics & Gerontology International.
Generation. (2024). The AI Divide: Survey of workers aged 45 and over and employers in France, Ireland,
Spain, the United Kingdom, and the United States (released October 8, 2024). generation.org. (grey
literature; NGO-commissioned survey)
Gerpott, F. H., Lehmann-Willenbrock, N., & Voelpel, S. C. (2017). A phase model of intergenerational
learning in organizations. Academy of Management Learning & Education, 16(2), 193–216. doi:10.5465/
amle.2015.0185
Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect
mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127.
doi:10.1136/amiajnl-2011-000089
Grossmann, I., Na, J., Varnum, M. E. W., Park, D. C., Kitayama, S., & Nisbett, R. E. (2010). Reasoning about
social conflicts improves into old age. Proceedings of the National Academy of Sciences, 107, 7246–7250.
doi:10.1073/pnas.1001715107
Hartshorne, J. K., & Germine, L. T. (2015). When does cognitive functioning peak? The asynchronous rise
and fall of different cognitive abilities across the life span. Psychological Science, 26(4), 433–443. doi:
10.1177/0956797614567339
Hautus, M. J., Macmillan, N. A., & Creelman, C. D. (2022). Detection Theory: A User's Guide (3rd ed.). New
York: Routledge.
Hong, L., & Page, S. E. (2004). Groups of diverse problem solvers can outperform groups of high-ability
problem solvers. Proceedings of the National Academy of Sciences, 101(46), 16385–16389.
Horn, J. L., & Cattell, R. B. (1967). Age differences in fluid and crystallized intelligence. Acta Psychologica,
26(2), 107–129. doi:10.1016/0001-6918(67)90011-X
Working Paper | Ageless Management in the AI Era 183
Insler, M. (2014). The health consequences of retirement. Journal of Human Resources, 49(1), 195–233. doi:
10.3368/jhr.49.1.195
Joshi, A., & Roh, H. (2009). The role of context in work team diversity research: A meta-analytic review.
Academy of Management Journal, 52(3), 599–627. doi:10.5465/amj.2009.41331491
Kadowaki, N. (2026a). Future Value Theory(ๆชๆฅไพกๅค็่ซ). VURA Working Paper Series No.1. ใใฅใผใฉใญใฃใ
ใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026b). Enterprise Redefinition(ไผๆฅญๅๅฎ็พฉ). VURA Working Paper Series No.2. ใใฅใผใฉใญใฃ
ใใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026c). Enterprise Redefinition Observed. VURA Working Paper Series No.3. ใใฅใผใฉใญใฃใ
ใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026d). From Job Description to Purpose Description. VURA Working Paper Series No.4.
ใใฅใผใฉใญใฃใใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026e). Brain Capital Management(่ณ่ณๆฌ็ตๅถ). VURA Working Paper Series No.5. ใใฅใผใฉ
ใญใฃใใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026f). Self-Defined Society(่ชๅทฑๅฎ็พฉๅ็คพไผ). VURA Working Paper Series No.6. ใใฅใผใฉใญใฃ
ใใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026g). Redefinition Capitalism(ๅๅฎ็พฉ่ณๆฌไธป็พฉ). VURA Working Paper Series No.7. ใใฅใผใฉ
ใญใฃใใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kadowaki, N. (2026h). Human on the Loop(ไบบใฎใฌใใใณในใฏใชใ็ ด็ถปใใใฎใ). VURA Working Paper Series
No.8. ใใฅใผใฉใญใฃใใฟใซใคใใใผใทใงใณใใผใซใใฃใณใฐในๆ ชๅผไผ็คพ.
Kaše, R., Saksida, T., & Miheliฤ, K. K. (2019). Skill development in reverse mentoring: Motivational
processes of mentors and learners. Human Resource Management, 58(1), 57–69. doi:10.1002/hrm.21932
Kirsh, B., Stergiou-Kita, M., Gewurtz, R., Dawson, D., Krupa, T., Lysaght, R., & Shaw, L. (2009). From margins
to mainstream: What do we know about work integration for persons with brain injury, mental illness
and intellectual disability? WORK, 32(4). doi:10.3233/WOR-2009-0851
Krzeminska, A., Austin, R. D., Bruyère, S. M., & Hedley, D. (2019). The advantages and challenges of
neurodiversity employment in organizations. Journal of Management & Organization, 25(4), 453–463.
doi:10.1017/jmo.2019.58
Li, Y., Gong, Y., Burmeister, A., Wang, M., Alterman, V., Alonso, A., & Robinson, S. (2021). Leveraging age
diversity for organizational performance: An intellectual capital perspective. Journal of Applied
Psychology, 106(1), 71–91.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle:
How Language Models Use Long Contexts. Transactions of the Association for Computational
Linguistics, 12, 157–173. doi:10.1162/tacl_a_00638
Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of
the American Medical Informatics Association, 24(2), 423–431. doi:10.1093/jamia/ocw105
Macmillan, N. A., & Creelman, C. D. (2005). Detection Theory: A User's Guide (2nd ed.). Mahwah, NJ:
Lawrence Erlbaum Associates.
Marinaci, T., Russo, C., Savarese, G., Stornaiuolo, G., Faiella, F., Carpinelli, L., Navarra, M., Marsico, G., &
Mollo, M. (2023). An inclusive workplace approach to disability through assistive technologies: A
systematic review and thematic analysis of the literature. Societies, 13(11), 231. doi:10.3390/
soc13110231
Meng, A., Nexø, M. A., & Borg, V. (2017). The impact of retirement on age related cognitive decline – a
systematic review. BMC Geriatrics, 17, Article 160. doi:10.1186/s12877-017-0556-7
National Academies of Sciences, Engineering, and Medicine. (2022). Global Roadmap for Healthy
Longevity. Washington, DC: The National Academies Press. (institutional report)
Working Paper | Ageless Management in the AI Era 184
Nielsen, J. (2023). Generative AI enhances old users' intellectual performance through wise winnowing.
UXTigers (published June 28, 2023; updated January 1, 2026). (essay; no empirical data)
Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial
intelligence. Science, 381(6654), 187–192. doi:10.1126/science.adh2586
Omri, A., Omri, H., & Afi, H. (2025). Exploring the impact of AI on unemployment for people with
disabilities: Do educational attainment and governance matter? Frontiers in Public Health, 13, 1559101.
doi:10.3389/fpubh.2025.1559101
Parker, M., Bucknall, M., Jagger, C., & Wilkie, R. (2020). Population-based estimates of healthy working life
expectancy in England at age 50 years: Analysis of data from the English Longitudinal Study of Ageing.
The Lancet Public Health, 5(7), e395–e403.
Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity:
Evidence from GitHub Copilot. arXiv:2302.06590. (unrefereed preprint; vendor-affiliated research)
Phelps, E. S. (1972). The statistical theory of racism and sexism. American Economic Review, 62(4), 659–
661.
Pizzinelli, C., & Tavares, M. M. (2026). Artificial Intelligence and Older Workers: Opportunity, Risk, and
Policy Tradeoffs. Pension Research Council, Wharton School (May 21, 2026, blog essay; related working
paper: AI and the Future of Work in an Aging Economy, SSRN 5345347). (policy essay; grey literature)
Reber, R., & Schwarz, N. (1999). Effects of perceptual fluency on judgments of truth. Consciousness and
Cognition, 8(3), 338–342. doi:10.1006/ccog.1999.0386
Rohwedder, S., & Willis, R. J. (2010). Mental retirement. Journal of Economic Perspectives, 24(1), 119–138.
doi:10.1257/jep.24.1.119
Salthouse, T. A. (2006). Mental exercise and mental aging: Evaluating the validity of the "use it or lose it"
hypothesis. Perspectives on Psychological Science, 1(1), 68–87. doi:10.1111/j.1745-6916.2006.00005.x
Salthouse, T. A. (2009). When does age-related cognitive decline begin? Neurobiology of Aging, 30(4), 507–
514. doi:10.1016/j.neurobiolaging.2008.09.023
Schaie, K. W. (2009). "When does age-related cognitive decline begin?" Salthouse again reifies the "crosssectional
fallacy". Neurobiology of Aging, 30(4), 528–533. doi:10.1016/j.neurobiolaging.2008.12.012
Schneid, M., Isidor, R., Steinmetz, H., & Kabst, R. (2016). Age diversity and team outcomes: A quantitative
review. Journal of Managerial Psychology, 31(1), 2–17. doi:10.1108/JMP-07-2012-0228
Schooler, C. (2007). Use it—and keep it, longer, probably: A reply to Salthouse (2006). Perspectives on
Psychological Science. doi:10.1111/j.1745-6916.2007.00026.x
Sen, A. (1999). Development as Freedom. New York: Alfred A. Knopf.
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when
trained on recursively generated data. Nature, 631(8022), 755–759. doi:10.1038/s41586-024-07566-y
Smith, P., & Smith, L. (2021). Artificial intelligence and disability: Too much promise, yet too little
substance? AI and Ethics, 1, 81–86. doi:10.1007/s43681-020-00004-5
Sommers, S. R. (2006). On racial diversity and group decision making: Identifying multiple effects of racial
composition on jury deliberations. Journal of Personality and Social Psychology, 90(4), 597–612.
Spence, M. (1973). Job market signaling. Quarterly Journal of Economics, 87(3), 355–374. doi:
10.2307/1882010
Staff, R. T., Murray, A. D., Deary, I. J., & Whalley, L. J. (2004). What provides cerebral reserve? Brain, 127(5),
1191–1199. doi:10.1093/brain/awh144
Stern, Y. (2002). What is cognitive reserve? Theory and research application of the reserve concept. Journal
of the International Neuropsychological Society, 8, 448–460.
Working Paper | Ageless Management in the AI Era 185
Stern, Y. (2012). Cognitive reserve in ageing and Alzheimer's disease. Lancet Neurology, 11, 1006–1012. doi:
10.1016/S1474-4422(12)70191-6
Strathern, M. (1997). 'Improving ratings': Audit in the British University system. European Review, 5(3),
305–321. doi:10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4
Takeuchi, H., Ide, K., Wang, H., Tamura, M., & Kondo, K. (2024). The association of agricultural and nonagricultural
work on the healthy ageing of older adults in Japan: A 6-year longitudinal study from the
Japan Gerontological Evaluation Study. Preventive Medicine Reports, 49, 102949.
Teece, D. J., Pisano, G., & Shuen, A. (1997). Dynamic capabilities and strategic management. Strategic
Management Journal, 18(7), 509–533.
Thompson, A. (2014). Does diversity trump ability? An example of the misuse of mathematics in the social
sciences. Notices of the American Mathematical Society, 61(9), 1024–1030.
Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A
systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. doi:10.1038/
s41562-024-02024-1
van Ours, J. C. (2022). How retirement affects mental health, cognitive skills and mortality; An overview of
recent empirical evidence. De Economist, 170(3), 375–400. doi:10.1007/s10645-022-09410-y
Wallrich, L., Opara, V., Wesoลowska, M., Barnoth, D., & Yousefi, S. (2024). The relationship between team
diversity and team performance: Reconciling promise and reality through a comprehensive metaanalysis
registered report. Journal of Business and Psychology, 39, 1303–1354. doi:10.1007/
s10869-024-09977-0
Wegge, J., Jungmann, F., Liebermann, S., Shemla, M., Ries, B. C., Diestel, S., & Schmidt, K.-H. (2012). What
makes age diverse teams effective? Results from a six-year research program. Work, 41(Suppl 1), 5145–
5151. doi:10.3233/WOR-2012-0084-5145
WHO. (2020). Decade of Healthy Ageing: Plan of Action 2021–2030. Geneva: World Health Organization
(declaration adopted by the UN General Assembly on December 14, 2020). (international organization
document)
Xue, B., Cadar, D., Fleischmann, M., Stansfeld, S., Carr, E., Kivimäki, M., McMunn, A., & Head, J. (2018). Effect
of retirement on cognitive function: The Whitehall II cohort study. European Journal of Epidemiology,
33, 989–1001.
Zélity, B. (2023). Age diversity and aggregate productivity. Journal of Population Economics, 36(3), 1863–
1899. doi:10.1007/s00148-022-00911-3
Primary Legal Sources and Official Materials
Act on Stabilization of Employment of Elderly Persons, etc. (Kลnenreisha-tล no Koyล no Antei-tล ni
kansuru Hลritsu; Act No. 68 of 1971). Ministry of Health, Labour and Welfare statutes database. https://
www.mhlw.go.jp/web/t_doc?dataId=75049000&dataType=0&pageNo=1 (primary legal source)
Labor Standards Act (Rลdล Kijun Hล; Act No. 49 of 1947), Chapter 6 "Minors" (Articles 56–63). Hyogo
Labour Bureau commentary. https://jsite.mhlw.go.jp/hyogo-roudoukyoku/hourei_seido_tetsuzuki/
roudoukijun_keiyaku/nensyousya.html (primary legal source and administrative commentary)
Act on Ensuring Proper Transactions Involving Specified Entrusted Business Operators (Tokutei Jutaku
Jigyลsha ni kakaru Torihiki no Tekiseika-tล ni kansuru Hลritsu; Act No. 25 of 2023). Cabinet Secretariat
policy portal. https://www.cas.go.jp/jp/seisaku/atarashii_sihonsyugi/freelance/index.html (primary legal
and administrative source)
Overview of the Act Partially Amending the National Pension Act, etc. for Strengthening the Functions of
the Pension System in Light of Social and Economic Changes (Act No. 74 of 2025). Ministry of Health,
Labour and Welfare. https://www.mhlw.go.jp/content/12401000/001523466.pdf (primary administrative
source)
Working Paper | Ageless Management in the AI Era 186
Ministry of Health, Labour and Welfare. Overview of the amendment to the Act on Stabilization of
Employment of Elderly Persons (securing work opportunities up to age 70) [Kลnenreisha koyล antei hล
no kaisei (70-sai made no shลซgyล kikai kakuho) no gaiyล]. https://www.mhlw.go.jp/stf/seisakunitsuite/
bunya/koyou_roudou/koyou/koureisha/topics/tp120903-1_00001.html (primary administrative source)
Ministry of Health, Labour and Welfare. Leaflet on the expansion of special enrollment in workers'
accident compensation insurance (specified entrusted business operators). https://www.mhlw.go.jp/
content/001250166.pdf (primary administrative source)
Ministry of Health, Labour and Welfare. Table of Class II special enrollment insurance premium rates
(effective November 1, 2024). https://www.mhlw.go.jp/content/
tokubetsukanyuuhokenryouritsu_R0504.pdf (primary administrative source)
Ministry of Health, Labour and Welfare (Labour Standards Bureau). Administrative circular on the
employee status of students in internships (Circular Kihatsu No. 636 of September 18, 1997). Nagano
Labour Bureau materials. https://jsite.mhlw.go.jp/nagano-roudoukyoku/library/nagano-roudoukyoku/
_new-hp/2hourei_seido/roudoukijun/internship291006.pdf (administrative circular; Labour Bureau
commentary)
Report of the Labor Standards Act Study Group, "On the Criteria for Determining 'Worker' Status under the
Labor Standards Act" (Rลdล Kijun Hล Kenkyลซkai hลkoku; December 19, 1985). In: Ministry of Health,
Labour and Welfare, Labour Standards Bureau, "Reference Materials on the Determination of
Employee Status under the Labor Standards Act" (as of October 2024). https://www.mhlw.go.jp/content/
001462701.pdf (administrative source)
Japan Pension Service. Special page on the revision of the in-work old-age pension offset (zaishoku rลrei
nenkin) (effective April 2026). https://www.nenkin.go.jp/tokusetsu/zairoukaisei.html (primary
administrative source)
Japan Pension Service. Pensions while working (calculation method of the in-work old-age pension offset).
https://www.nenkin.go.jp/service/jukyu/seido/roureinenkin/zaishoku/20150401-01.html (primary
administrative source)
Japan Institute for Labour Policy and Training (JILPT). Korea: Overview of the retirement-age extension
act (2013). https://www.jil.go.jp/foreign/jihou/2013_6/korea_02.html (public research institute source)
Age Discrimination in Employment Act of 1967 (ADEA), 29 U.S.C. §621 et seq. EEOC. https://www.eeoc.gov/
statutes/age-discrimination-employment-act-1967 (primary legal source)
Council Directive 94/33/EC of 22 June 1994 on the protection of young people at work. EUR-Lex. https://eurlex.
europa.eu/legal-content/EN/LSU/?uri=celex:31994L0033 (primary legal source)
Council Directive 2000/78/EC of 27 November 2000 establishing a general framework for equal treatment
in employment and occupation. EUR-Lex. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?
uri=CELEX%3A32000L0078 (primary legal source)
Deutsche Rentenversicherung. FAQ: Änderungen beim Hinzuverdienst (8. SGB IV-ÄndG, effective January
1, 2023). https://www.deutsche-rentenversicherung.de/DRV/DE/Rente/Allgemeine-Informationen/
Wissenswertes-zur-Rente/FAQs/Rente/Hinzuverdienst_und_Einkommensanrechnung/
aenderungen_hinzuverdienst_liste.html (primary administrative source)
ILO. Minimum Age Convention, 1973 (No. 138). Convention text (hosted by OHCHR). https://www.ohchr.org/
en/instruments-mechanisms/instruments/minimum-age-convention-1973-no-138 (primary treaty)
ILO. Worst Forms of Child Labour Convention, 1999 (No. 182) — universal ratification (August 4, 2020).
https://www.ilo.org/resource/news/ilo-child-labour-convention-achieves-universal-ratification (primary
source; official ILO announcement)
L&E Global. Singapore: Retirement age and re-employment age to be raised on 1 July 2026. https://
leglobal.law/2026/03/24/singapore-retirement-age-and-re-employment-age-to-be-raised-on-1-july-2026-
Working Paper | Ageless Management in the AI Era 187
and-other-related-changes/ (professional commentary; final confirmation against MOM primary
sources remains outstanding)
Senior Citizens' Freedom to Work Act of 2000, Public Law 106-182. https://www.congress.gov/106/plaws/
publ182/PLAW-106publ182.htm (primary legal source)
Social Security Administration. Exempt Amounts Under the Earnings Test. https://www.ssa.gov/oact/cola/
rtea.html (primary administrative source)
Working Paper | Ageless Management in the AI Era 188

