Ask what AI is doing to human cognition and most answers reach immediately for metaphor — exoskeletons, calculators for the mind, a new nervous system. This article is the opening piece in a series on human–AI coevolution, and its job is narrower and less exciting than a metaphor: to lay out the actual documented mechanics before anyone is allowed to speculate about where they lead. Three bodies of work do the load-bearing work. Philosophy of mind supplies a functional definition of when an external tool counts as part of a cognitive process rather than merely an aid to one [1]. Cognitive psychology supplies decades of experiments measuring when and how people actually offload mental work onto tools, now including AI systems specifically [2] [3] [4]. And cultural-evolution theory supplies a documented account of how new tools for storing and transmitting information have repeatedly reorganized human institutions, long before anyone had heard of a transformer model [5]. None of this is speculative. All of it has been measured, published, and in the case of the philosophical framework, argued over for a quarter century. What is speculative is what happens next — and this piece is careful to say so only where it actually turns to forecasting.

Extended cognition: a real argument, not a slogan

“The mind extends beyond the skull” sounds like a rhetorical flourish. In its original form it is a specific, testable philosophical claim. Andy Clark and David Chalmers proposed what is now called the parity principle: if a process happening outside the body would count as a cognitive process if it happened inside the head, and it is reliably available, easily accessed, and automatically endorsed once retrieved, then the location of the process — inside or outside the skull — should not matter to whether it counts as cognition [1]. Their canonical case is a person with memory impairment who keeps a notebook of addresses and facts, consulting it as automatically and unreflectively as someone else consults biological memory. The notebook, on their argument, is functionally part of that person’s memory system, not an aid standing outside it.

This is a claim about function, not a claim about what feels like “real” thinking. It does not say that a smartphone is conscious, or that offloading a task to a tool is cognitively free, or that every external aid qualifies — Clark and Chalmers are explicit that the coupling has to be reliable, trusted, and easily used, which rules out a reference book you rarely open or don’t trust. The parity principle has been contested since 1998, and the philosophical literature debating where its boundaries lie is itself large; this article is not adjudicating that debate. What matters for a series on human–AI coevolution is narrower: the extended-mind framework is the reason cognitive scientists studying tool use ask a specific question — not “is this tool helpful” but “has part of the cognitive process moved into the tool” — and that question is exactly the one the empirical offloading literature goes on to measure.

ADVERTISEMENT

It is also worth being precise about what the thesis does not claim, since a great deal of loose talk about AI “becoming part of us” borrows the extended-mind label without the argument’s actual conditions attached. Clark and Chalmers’s criteria are functional and falsifiable: the resource must be constantly available, its output must be accepted more or less automatically rather than scrutinized as an outside opinion each time, and the information must have been placed there, or retrieved from there, as a matter of the person’s own prior endorsement. A tool a person distrusts and double-checks every time — precisely the recommended posture toward an AI system whose outputs can be wrong in ways that are hard to detect — does not meet the coupling criterion, on the authors’ own account. That detail matters for everything that follows: whether a given use of an AI tool counts as genuine cognitive extension, or as consulting an outside, unreliable source held at arm’s length, is itself an empirical question about trust and verification behavior, not something settled by the mere fact that the tool is fast and convenient.

Cognitive offloading: what the experiments actually measured

Between the 1998 philosophical argument and today’s AI tools sits a body of experimental psychology that gives the argument teeth. Evan Risko and Sam Gilbert’s 2016 review defines cognitive offloading precisely: “the use of physical action to alter the information processing requirements of a task so as to reduce cognitive demand” [2]. Their review synthesizes decades of small experiments — setting a reminder instead of remembering an appointment, rotating a physical object instead of mentally rotating it, writing a note instead of rehearsing a fact — and finds a consistent pattern: people offload more when the internal task is harder, and people’s self-assessment of when they need to offload is often wrong, meaning the decision to reach for a tool is itself an imperfect metacognitive judgment, not a rational cost-benefit calculation performed correctly every time.

The specific case of externalizing memory to search tools was tested directly by Betsy Sparrow, Jenny Liu, and Daniel Wegner in a set of four experiments published in Science in 2011 [4]. Their finding, precisely stated: when people expect to have future access to information, they show lower recall of the information’s content but equal or better recall of where to find it again. Participants were not simply “getting lazier” — their memory was reorganizing around retrieval location rather than content, a pattern the authors called a transactive memory effect, extending an older idea about how couples and teams distribute what each member remembers so the group as a whole retains more than any one person could.

The AI-specific extension of this literature is recent and should be read with the caution due any single study. Michael Gerlich’s 2025 study in Societies surveyed and interviewed 666 participants and found a statistically significant negative correlation between frequent AI tool use and measured critical-thinking performance, with cognitive offloading identified as the mediating mechanism, and with the effect stronger among younger participants [3]. This is a real, peer-reviewed, published finding — not a rumor. It is also a correlational, self-report-heavy survey design, not a randomized controlled trial: it cannot by itself establish that AI use causes lower critical-thinking scores rather than, say, people who already offload more being drawn to heavier AI use, or age cohorts differing for reasons unrelated to AI at all. The honest reading is: one well-powered study finds an association consistent with the offloading mechanism the broader literature already describes, and it strengthens rather than settles the case.

Put the three findings together and a specific, non-metaphorical mechanism emerges. Extended-mind theory supplies the criterion for when a tool becomes functionally part of a cognitive process. Offloading research supplies the empirical behavior — reaching for that tool preferentially as internal demand rises, sometimes prematurely. And the Google-effects and Gerlich findings supply the consequence specifically for memory and reasoning tasks: content recedes from internal storage while retrieval pathways become the thing actually remembered and practiced. None of this requires believing AI tools are unusually powerful or unusually dangerous. The mechanism is continuous with notebooks, search engines, and calculators; what changes with generative AI tools is the range of tasks that now meet the “reliable, accessible, trusted” bar Clark and Chalmers set — no longer just storing a fact, but drafting an argument, checking a calculation, or producing a first-pass answer to a question a person has not yet tried to answer themselves.

ADVERTISEMENT

Cognitive offloading has an institutional twin: cultural evolution

The individual-level mechanism above has an institutional-level counterpart that is easy to miss if the discussion stays at the level of one person and one tool. Joseph Henrich’s cultural-evolution research program argues that human cognitive success was never primarily a matter of individual brainpower; it was a matter of a species that could reliably transmit, accumulate, and improve practical knowledge across generations faster than any individual could invent it from scratch [5]. Henrich’s evidence base includes historical cases of populations that lost complex technologies — knowledge of tools, food-processing techniques, or watercraft — after their populations shrank or their networks of cultural transmission were disrupted, even though no individual “forgot” anything; the knowledge existed only in the distributed, redundant, cross- checked practice of a large connected group, and it evaporated when that network thinned.

A card-catalog archive room with one drawer pulled open next to a terminal showing a matching lookup, one index card lifted halfway out
Figure 1. A transactive-memory archive: knowing where to look has always counted as a form of knowing.Image prompt and art direction by Brecht Corbeel; generation pending.

This is the same transactive-memory logic Sparrow’s participants displayed, scaled up from a household to a civilization: the unit that “remembers” is not the individual, it is the network, and what any one member needs to hold internally depends on what the rest of the network reliably holds instead. Henrich’s account explains why new transmission technologies — writing, printing, searchable databases, and now AI systems that can answer a question on demand — have historically reorganized institutions rather than simply making individuals within them smarter or lazier. When a technology changes what is cheap to transmit, verify, or reproduce, the institutions built around the old cost structure — how apprenticeship worked, what a credential certified, which errors a review process was designed to catch — become mismatched to the new one, and pressure builds toward institutional redesign, not just individual behavior change.

An eye-tracking booth with a chinrest and monitor mid-calibration, one gaze-marker dot still settling on the screen
Figure 2. An eye-tracking booth mid-calibration, the instrument that turns 'looked at the tool first' into a measurement.Image prompt and art direction by Brecht Corbeel; generation pending.

This is analysis built on a documented framework, not a new finding of its own: Henrich’s book synthesizes archaeological, ethnographic, and experimental evidence for a general claim about cumulative culture; applying that claim specifically to AI-era institutions is this article’s inference, clearly separated from Henrich’s own evidence base, which predates large language models entirely.

Henrich’s framework also supplies a second, less obvious point that keeps the AI case from being treated as unprecedented by default: cumulative culture has always required a mechanism for filtering good information from bad, and that filtering mechanism has repeatedly needed to be rebuilt each time a transmission technology changed who could produce information cheaply. The printing press did not just spread accurate knowledge faster; it also spread inaccurate and fraudulent material faster, and the institutions that emerged in response — peer review, editorial gatekeeping, citation norms, credentialing bodies — were themselves cultural inventions built to solve a verification problem the new technology created. Read this way, the current debate about AI- generated misinformation and unverified AI-assisted output is not a new category of problem; it is a specific instance of a pattern cultural-evolution theory already documents, which is a reason for calibrated concern rather than either panic or dismissal.

Where the mechanism meets institutions: education, work, and oversight

Three institutional domains show this mismatch pressure most clearly, and it is worth being precise about which claims in each are fact, which are vendor or advocacy assertion, and which are analysis.

Education. The fact: task performance on writing and problem-solving assignments can now be produced, in whole or in part, by tools external to the student, at a quality that is often indistinguishable from unaided work without dedicated detection effort — this is a direct consequence of the offloading and extended-mind mechanisms already described, not a separate finding. The analysis: assessment systems built to certify what a student can produce unaided are now testing a condition (guaranteed absence of tool access) that is increasingly hard to guarantee and, outside the exam room, increasingly irrelevant to how the same task will actually be performed later in life. That mismatch is Henrich’s institutional-lag pattern applied to a classroom: the credential’s design assumption has fallen out of step with the transmission technology students actually use.

ADVERTISEMENT
A stack of paper assessment booklets on a classroom table, the top booklet's cover lifted as a tablet slides into the stack's place
Figure 3. A classroom assessment table, mid-transition from a graded booklet to a tool that can draft the answer.Image prompt and art direction by Brecht Corbeel; generation pending.

Work. Tyna Eloundou and coauthors’ 2023 study is a documented, methodologically explicit piece of labor-economics research, not a vendor claim: using the U.S. Department of Labor’s O*NET task database, human annotators and a language model independently rated occupational tasks for exposure to large-language-model capability, and the study reports that roughly 80 percent of the U.S. workforce could have at least 10 percent of their tasks affected, and around 19 percent could have at least 50 percent of their tasks affected [6]. Two qualifications the authors themselves make explicit: “exposure” measures task susceptibility to being performed or assisted by the model, not job elimination, and the estimates depend on complementary tooling and organizational adoption that the paper does not itself measure. The World Economic Forum’s 2025 employer survey adds a complementary, differently sourced data point: employers report AI and big data as the technology skill category expected to matter most through 2030, with 39 percent of core job skills expected to change in that period [8]. This is a survey of employer expectation, not a measurement of realized change, and it should be read as such — it documents what organizations say they anticipate, which is itself useful evidence about institutional pressure even where it turns out directionally wrong about magnitude.

An office approval desk where a paper stamp hovers above a folder next to a review queue glowing softly on a monitor
Figure 4. An approval desk where a stamp still lands on paper beside a queue that never runs out of items to review.Image prompt and art direction by Brecht Corbeel; generation pending.

Oversight and governance. Allan Dafoe’s 2018 research agenda is not a prediction; it is a scoping document, written for the Centre for the Governance of AI, that lays out which institutional questions actually need answers before AI’s societal effects can be well governed — questions about who holds verification capacity, how safety and capability races interact, and how existing political institutions adapt to a general-purpose technology whose pace of capability change outstrips normal regulatory cycles [7]. Its value for this article is definitional: it is the clearest documented statement of what “institutional adaptation to AI” would need to consist of, against which any specific claim of adaptation (or its absence) can be checked, rather than a forecast that adaptation will or will not happen.

Separating fact, claim, analysis, and scenario

It is worth stating explicitly, in one place, how the claims above sort:

  • Documented fact: the parity principle as originally argued [1]; the offloading mechanisms and metacognitive miscalibration Risko and Gilbert catalog [2]; the specific recall-versus-location finding from Sparrow et al. [4]; the correlational finding in Gerlich’s survey [3]; the historical loss-of-technology cases Henrich documents [5]; the O*NET exposure percentages Eloundou and coauthors report [6]; the employer-expectation percentages in the WEF survey [8].
  • Vendor or advocacy assertion: none used directly in this piece; where AI capability claims appear anywhere in this series they should be attributed to the company making them, not stated as fact, and none of the above sources are vendor material — they are peer-reviewed research, an academic press book, and two institutional reports.
  • Analysis (this article’s own inference): that offloading and cultural-transmission theory describe the same mechanism at different scales; that assessment and credentialing systems face a Henrich-style institutional lag; that Dafoe’s governance agenda defines the adaptation bar against which real institutional change should be measured.
  • Scenario, explicitly flagged as such below: any statement about what happens over the next decade.
Two research desks laid out side by side for comparing scenarios, one folder open on each desk, a lamp switching on over only one
Figure 5. A divergence room, two scenario desks laid out for comparison — a discipline of holding more than one future open at once.Image prompt and art direction by Brecht Corbeel; generation pending.

Plural futures, not a single forecast

A series opener earns the right to end with scenario work only if it refuses to collapse the scenarios into one forecast. Three plausible, non-exhaustive equilibria follow from the mechanisms above, each with its own horizon, assumptions, and — critically — a condition that would show it wrong.

Equilibrium one: distributed-competence institutions. Assessment and credentialing redesign around Henrich’s actual insight — that a network’s competence was never purely individual — and begin certifying supervised human–tool collaboration directly, the way open-book, tool-permitted exams already do in some technical fields. Horizon: visible within five to ten years in fields where licensing bodies move first. Assumption: verification costs for collaborative output fall enough that credentialing bodies trust it. Disconfirmation: if credentialing bodies instead double down on tool-free, proctored assessment as the default over the same period, this equilibrium is not occurring.

Equilibrium two: offloading without institutional redesign. Individual and organizational offloading (per Risko and Gilbert, and per Gerlich’s correlational finding) increases faster than institutions adapt, producing a widening gap between what credentials certify and what practitioners can actually do unaided — a scenario, not yet a measured trend. Horizon: five years. Assumption: Gerlich’s association proves to reflect a real causal effect at scale, which the existing study alone cannot establish. Disconfirmation: a well-designed longitudinal or experimental study finding no causal link between AI-tool reliance and reduced unaided task performance would undercut this equilibrium’s premise directly.

Equilibrium three: governance catches the general-purpose-technology pattern early. Dafoe’s research agenda gets substantially answered — verification institutions, adoption-pace monitoring, and updated task-exposure tracking (extending Eloundou’s method past a single snapshot) become routine parts of how labor-market and safety policy are made, ahead of, rather than behind, capability change. Horizon: a decade or more, since it requires sustained institutional investment rather than a single policy. Assumption: political institutions treat AI as the general-purpose technology Dafoe frames it as, not as a sequence of unrelated product launches. Disconfirmation: if governance activity remains reactive to individual model releases rather than building the standing measurement infrastructure Dafoe’s agenda calls for, this equilibrium is not materializing.

These three are not mutually exclusive, and nothing in the evidence surveyed here picks a winner among them. That is the honest state of the documented research today: the individual-level mechanism — extended cognition made measurable through offloading experiments — is well established; the institutional-level mechanism — cultural-evolution pressure toward redesign under a new transmission technology — is well established as a historical pattern; and which specific institutional future results from applying the first to the second, at AI’s current scale and pace, remains open. Later pieces in this series take up each equilibrium’s evidence in turn.