Three research communities currently claim to explain how humans and artificial intelligence are reshaping each other, and they rarely cite each other’s central results. Cultural-evolution theorists model cumulative culture as an information-transmission process running across generations of learners [4]. Cognitive scientists and HCI researchers run controlled experiments on individual memory, attention, and offloading when a tool is available [1] [2] [3]. Futures and governance scholars build explicit, falsifiable scenarios about how institutions might or might not adapt to a general-purpose technology over decades [6] [7]. Each tradition is right about something the others cannot see, and each has a blind spot the others cover. This article compares them directly — on falsifiability, time horizon, and actionability — rather than declaring a winner, and then applies the comparison to five live questions before laying out a small set of plural, non-exclusive futures.
Why three traditions instead of one
The reason no single field owns this question is that “how AI changes humans” decomposes into at least three genuinely different objects of study, each with its own natural unit of analysis, its own timescale, and its own evidentiary standard.
Cultural-evolution theory treats the unit of analysis as the population, not the individual. What spreads, what is retained across generations of learners, and what determines whether a technique’s fidelity degrades or improves as it passes hand to hand — these are population-level dynamics, and they are modeled with the same formal tools used for cumulative culture before any computer existed [4]. The approach is about process, not any particular technology: it asks how a “collective brain” made of many partially-informed individuals can produce and retain innovations no single member could invent alone, and it treats AI as one more channel through which information is transmitted, filtered, and recombined across that population.
Cognitive-offloading and HCI research treats the unit of analysis as the individual mind in a controlled trial. Cognitive offloading is defined narrowly and operationally: using a physical action or an external tool to reduce the internal cognitive demand of a task, whether that is tilting your head to read a rotated image or letting a phone hold an appointment instead of memory [1]. Sparrow, Liu, and Wegner’s 2011 experiments found that when people expect future access to information, they retain weaker memory for the content itself and stronger memory for where to find it — a shift from “what” to “where” in encoding, replicated and extended many times since as a documented empirical effect on recall strategy, not a claim about intelligence or wellbeing in general [2]. This tradition’s currency is the effect size from a specific manipulation in a specific population, and its horizon is the length of a study, from a single session to a few years of survey data.
Futures and long-term governance scholarship treats the unit of analysis as the institution or the civilization, and its currency is neither an effect size nor a transmission-fidelity curve but an explicit, falsifiable scenario with named preconditions. Dafoe’s AI governance research agenda frames the field around three linked questions — the technical landscape a governance regime must anticipate, the political and economic effects that follow from it, and the ideal long-run institutional arrangements that would make those effects tolerable — and treats each as requiring its own research program rather than a single grand forecast [6]. Bostrom’s vulnerable world hypothesis goes further, formalizing “vulnerability” as a testable structural property of a technology space (does some level of technological capability almost certainly devastate an unprepared civilization by default, absent exit from a “semi-anarchic” condition) rather than a prediction about any one event [7]. This tradition’s natural output is a small set of named, mutually exclusive scenarios plus the observable indicators that would distinguish between them — closer to structured argument than to measurement.
Falsifiability: what would each approach have to observe to be wrong
Falsifiability is the sharpest line between the three traditions, and it is worth stating plainly rather than treating all three as equally scientific.
The HCI and cognitive-offloading tradition is the most falsifiable in the classical sense. A specific claim — memory for content declines when future access is expected, mediated by increased reliance on external storage — is stated in a form that a failed replication or a null result on a moderator variable could directly contradict [2] [1]. Gerlich’s 2025 mixed-method study of 666 participants reporting a negative correlation between frequent AI-tool use and measured critical-thinking scores, mediated by cognitive offloading and moderated by age, is falsifiable in the same sense: a differently-sampled replication could fail to find the correlation, or could find it explained entirely by a confound such as prior educational attainment [3]. This is a real methodological strength and a real limitation together: the tradition buys falsifiability by narrowing scope to what a single study or survey wave can measure, which is rarely the multi-decade, population-wide process the other two traditions are actually trying to explain.
Cultural-evolution theory is falsifiable at the level of its formal models — a transmission-chain experiment either does or does not show the fidelity curve the model predicts, and a “collective brain” account of innovation makes testable claims about how population size, connectivity, and social learning strategies should affect the rate and character of innovation [4]. But its application to any specific technology, including generative AI, involves an extrapolation the formal theory itself does not license: knowing that cumulative culture in general depends on network structure and transmission fidelity does not by itself tell you whether large language models increase or decrease net transmission fidelity for any particular domain of practice. The theory is falsifiable; a given application of it to AI is often closer to informed analogy.
Futures and governance scholarship is the least falsifiable in the narrow sense and openly says so: a scenario is not a point prediction, and a governance research agenda’s value lies in identifying which levers and indicators matter, not in getting a single forecast right [6]. Bostrom’s vulnerable-world argument is explicitly a structural hypothesis about a space of possible technologies, evaluated by whether its named preconditions and countermeasures are coherent and its examples are apt, not by a single confirming or disconfirming observation [7]. That is not a defect if the scenarios are built with disconfirmation conditions stated up front — a scenario that specifies “if indicator X has not moved by year Y, downgrade this branch’s probability” is doing real epistemic work even though no single experiment refutes it. Many published scenario exercises fall short of that standard; readers should treat a scenario without a stated indicator and horizon as scenario-flavored narrative rather than a scenario proper.
Time horizon and actionability
The three traditions also differ in what a practitioner can actually do with their output today, which is a separate axis from falsifiability.
HCI and cognitive-offloading findings are the most immediately actionable at small scale: a documented negative correlation between AI-assisted task completion and independent recall or skill-building gives a teacher, a product designer, or an individual user something to act on this term, not this decade — reduce reliance on a tool during the acquisition phase of a skill, or design an interface that requires a retrieval step before delivering an answer [1] [3]. Its actionable claims are also the most local: what holds for a university sample doing short problem sets does not automatically hold for professional adults using AI for entirely different tasks, and per §-level editorial practice here that generalization gap should be stated, not assumed.
Cultural-evolution theory is actionable at the level of institutional design over a medium horizon — years, not days: if innovation is a population-level process depending on connectivity and transmission fidelity between partially informed individuals, then institutions that widen access to a “collective brain” (broader participation in a research community, cheaper access to tools, higher-fidelity documentation) should raise the rate of useful recombination, a design implication that follows fairly directly from the theory [4]. It is far less actionable at the individual level — it says little about what a single person should do differently tomorrow.
Futures and governance scholarship is actionable primarily for institutions and policymakers operating on a horizon of years to decades: it identifies which levers (compute access, verification cost, international coordination mechanisms) are worth building monitoring capacity around now, even before any scenario resolves [6]. Its actionability is conditional and preparatory rather than immediate — build the capacity to notice which branch is happening, rather than act as though one branch is already confirmed.
Applying the comparison to five live questions
Cognitive distribution. The HCI tradition supplies the clearest current evidence: cognitive offloading is a real, measured behavior with real consequences for what gets encoded internally versus externally [1] [2]. Cultural-evolution theory supplies the frame for why this matters beyond any one person — a “collective brain” was already distributing cognition across specialists and artifacts before AI existed [4]. Futures scholarship supplies the open question neither of the other two can answer on its own: whether institutions will build verification layers that keep distributed cognition reliable, or let verification cost quietly rise until errors compound unnoticed [6].
Education. Gerlich’s finding of an age-moderated negative correlation between AI-tool use and measured critical-thinking scores is a fact about that sample, not a settled fact about all learners everywhere [3]. It is one empirical data point that a cultural-evolution account would situate inside a longer argument about which learning practices survive transmission across a cohort of students, and that a governance scenario would situate inside an institutional question about whether assessment design adapts. None of the three alone settles what schools should do; together they at least separate the measured effect from the institutional response from the transmission dynamic.
Identity and agency. This is the domain where all three traditions are weakest and should say so. HCI experiments rarely run long enough to measure identity effects. Cultural-evolution models are agnostic about subjective experience. Futures scenarios can name identity-related branches (a world in which AI mediation becomes constitutive of professional identity versus one in which it stays clearly instrumental) but cannot yet attach indicators to them with much precision. Analysts should mark identity claims in this space as analysis or scenario, not fact, until better instruments exist.
Institutional authority and dependence. Acemoglu and Restrepo’s task-based framework — automation displaces labor from some tasks while creating new tasks in which labor retains a comparative advantage, with the net labor-share effect depending on which force dominates — is a fact pattern from decades of automation history, not a claim specific to AI [5]. It is directly relevant to dependence: if AI mostly displaces tasks without creating comparably many new ones where humans hold an advantage, institutional dependence on AI-mediated judgment rises faster than compensating human capability is created. Whether that happens is an open, falsifiable empirical question, trackable through the same kind of task-level labor data Acemoglu and Restrepo used, not something either the cultural-evolution or futures traditions can settle by argument alone.
Cultural selection and agency at scale. The Stanford AI Index’s 2026 data point that generative AI reached population-level adoption faster than the personal computer or the internet is a fact about diffusion speed [8]. It is a vendor-adjacent claim only in the loose sense that adoption figures are compiled from usage and survey data, not audited transaction records, and should be read as measured diffusion, not as evidence about quality or benefit. The same report’s finding that organizational adoption reached 88 percent while measured consumer surplus roughly tripled year over year describes economic uptake, not whether that uptake improves the tasks people actually value doing [9]. Fast diffusion is a precondition for coevolutionary pressure at cultural-evolution timescales, not proof that any particular selection outcome — toward more capable humans, more dependent humans, or some mixture — has already occurred.
Plural future equilibria, not a single trajectory
Because the three traditions answer different questions, none of them alone justifies picking one future as most likely. What can be built responsibly is a small set of named, non-exclusive equilibria, each with an indicator and an explicit disconfirmation condition — the standard futures scholarship itself sets [6] [7].
Distributed-competence equilibrium. Verification layers and education adapt roughly as fast as offloading spreads; measured critical-thinking effects stay confined to narrow tasks and do not generalize to overall capability. Indicator: replication of offloading-critical-thinking correlations across varied populations shows shrinking, not growing, effect sizes over a five-year horizon. Disconfirmation: effect sizes instead grow across replications and generalize beyond the original task types.
Task-reinstatement equilibrium. New tasks in which humans retain a comparative advantage are created roughly as fast as old ones are automated, consistent with the historical pattern Acemoglu and Restrepo document for prior automation waves [5]. Indicator: labor-share and task-content data show new task categories emerging at a pace comparable to displaced ones. Disconfirmation: new-task creation lags displacement for a sustained multi-year period.
Verification-debt equilibrium. Institutions under-invest in the governance capacity Dafoe’s agenda calls for, dependence rises faster than error-correction capacity, and Bostrom’s vulnerability condition becomes more rather than less applicable over time [6] [7]. Indicator: documented institutional near-misses or audits showing undetected AI-mediated errors accumulate. Disconfirmation: independent verification and audit capacity is observed to scale alongside adoption rather than lag it.
These three are not mutually exclusive branches of a decision tree; different domains (education, labor markets, high-stakes institutional decisions) can sit in different equilibria simultaneously, and cultural-evolution dynamics, individual-level offloading effects, and governance responses will keep interacting rather than resolving into one outcome. That is the central reason to keep the three research traditions distinct rather than collapsing them into a single “human-AI coevolution” narrative: no one of them, alone, tracks the whole system, and no one of them should be read as though it does.
Reading disagreements between the three traditions correctly
When the three traditions appear to disagree, the disagreement is usually about scope, not about facts. A cultural-evolution theorist reading Gerlich’s correlation between AI use and lower measured critical-thinking scores would not dispute the correlation; they would ask whether the population-level transmission process compensates over a longer horizon than any single study can observe — whether the practices that survive generational transmission end up being the ones that preserve verification skill, even if any one cohort’s cross-sectional snapshot looks worse [3] [4]. A futures scholar reading the same correlation would ask a still different question: which governance lever, if pulled now, would change the indicator’s trajectory, and what would count as evidence the lever worked [6]. None of these readings contradicts the underlying data point; they attach it to a different explanatory frame with a different time horizon. Treating that as a factual dispute, rather than a difference in scope, is the most common misreading of this literature, and it is worth naming explicitly so that a single study is not asked to answer a question it was never designed to address.
This also clarifies what “actionability” should mean in context. A school administrator deciding this term’s policy on AI-assisted homework needs the HCI evidence base, imperfect and narrow as it is, because no cultural-evolution model or governance scenario operates on a school-term horizon [1] [3]. A national regulator deciding whether to fund AI-verification infrastructure needs the futures and governance framing, because the relevant horizon and stakes exceed what any lab study can speak to [6] [7]. A curriculum designer thinking about which literacies to teach across a generation needs the cultural-evolution frame, because that is the timescale at which transmission fidelity and collective-brain effects actually operate [4]. Matching the question to the tradition built to answer it, rather than reaching for whichever study is most recent or most alarming, is the practical payoff of keeping the three approaches distinct.
What the comparison rules out
It rules out treating any single well-designed HCI study as settling a generational question, since its horizon is a single wave or trial, not a generation [2] [3]. It rules out treating a cultural-evolution model’s formal elegance as itself evidence about any specific current technology, since the model’s scope is the transmission process in general, not large language models in particular [4]. And it rules out treating a governance scenario as a forecast with a probability attached rather than as a structured argument about levers and indicators [6] [7]. Comparing the approaches on these terms is not a way of picking a winner; it is a way of knowing, for any specific claim about human-AI coevolution encountered elsewhere, which kind of evidence would actually be relevant to checking it.