Why this argument needs falsifiable scenarios, not a verdict

Public debate about generative AI and jobs mostly takes one of two unfalsifiable postures: confident replacement (“this occupation is over”) or confident reassurance (“new jobs always appear, they always have”). Both postures survive contact with any single data point, which is exactly what makes them useless as forecasts. Labor economics has a more disciplined vocabulary for this question, built around tasks rather than jobs, and around displacement versus reinstatement rather than a single verdict of gain or loss [1]. This article uses that vocabulary to build three scenarios for 2035, each with a stated horizon, explicit assumptions, named observable indicators, and — critically — a condition that would prove it wrong. The goal is not to predict a winner. It is to specify, in advance, what evidence would settle the question, so that whichever way labor markets move by 2035, the record shows whether this analysis anticipated it or missed it.

Throughout, distinguish four kinds of statement. A fact is a number from an official statistical agency or a peer-reviewed study, cited to its source. A vendor claim is a company’s own description of what its product does, reported as a claim and not independently verified. An analysis is a judgment this article draws from the facts, clearly marked as interpretation. A scenario is a conditional “if this, then that” structure. A prediction is a scenario with a horizon and a disconfirmation condition attached. Conflating these is the single most common failure in popular writing about AI and labor, and it is worth naming every time the register shifts.

The task framework: what automation actually changes

The task-based framework treats an occupation as a bundle of discrete tasks, each performed by labor, by capital, or by some mix of the two [1]. Automation does not eliminate occupations directly; it changes the relative cost of performing specific tasks with capital instead of labor. Two forces then operate in opposite directions. Displacement occurs when tasks previously performed by workers are reassigned to machines or software, which — holding everything else constant — lowers labor demand and wages for the workers who specialized in those tasks. Reinstatement occurs when the same technological change creates new tasks in which labor retains a comparative advantage — new job titles, new quality margins, new categories of work that did not exist before. Acemoglu and Restrepo’s empirical decomposition attributes a substantial share — by their estimate, roughly half to more than two-thirds — of the rise in U.S. wage inequality since the 1980s to industries where routine-task automation expanded faster than new-task creation kept pace [1]. That is a fact about the past four decades of automation broadly, not a fact about generative AI specifically, and the two should not be blurred.

ADVERTISEMENT

Generative AI’s task profile differs from earlier waves of automation in one structural way worth naming precisely: it targets non-routine cognitive and language tasks — drafting, summarizing, coding, classifying — that earlier automation technologies could not touch at all. Occupational exposure research led by Eloundou and colleagues estimated that around 80 percent of the U.S. workforce could have at least 10 percent of their work tasks affected by large language models, and about 19 percent of workers could see at least half their tasks affected, using a rubric built from both human expert judgment and model-based classification [2]. This is an exposure estimate, not a displacement estimate — it measures which tasks a tool could plausibly touch, not whether firms will restructure jobs around that capability, whether workers will retain those tasks in an augmented form, or whether wages will move at all. Treating an exposure number as a job-loss number is the single most common misreading of this literature, and it recurs constantly in press coverage.

A time-and-motion stopwatch clipped to a clipboard survey form on a worktable, its second hand caught mid-sweep past a marked interval line
Figure 1. Task-timing studies, the empirical ancestor of modern occupational task data, measure how long a task takes before and after a tool changes — the same logic economists now apply to generative-AI assistance.Image prompt and art direction by Brecht Corbeel; generation pending.

What field evidence actually shows so far: augmentation, unevenly distributed

The strongest evidence available for how generative AI changes work inside real firms comes from a randomized field study of customer-support agents, not from surveys or vendor pitches. Brynjolfsson, Li, and Raymond studied the staggered rollout of a generative AI conversational assistant across 5,179 support agents at an enterprise software firm and found productivity gains — measured as issues resolved per hour — averaging 14 percent, but concentrated overwhelmingly among novice and lower-skilled workers, who saw gains as high as 34 percent, while the most experienced agents saw little to no measurable benefit [3]. The tool appeared to disseminate the tacit practices of the firm’s most effective agents to everyone else, compressing the skill distribution rather than replacing any of the workers studied. This is a fact about one firm, one tool, and one occupation over one measured period — it is evidence of a compression mechanism, not proof that the mechanism generalizes to every white-collar occupation, and the paper itself does not claim that.

At the level of a national labor market, the OECD’s 2023 Employment Outlook reviewed available cross-country evidence and reported no signs, as of that review, that AI adoption had produced detectable aggregate employment declines in occupations most exposed to it; instead it found that job quality and wage effects were more visible than headcount effects, with outcomes depending heavily on whether workers had the complementary skills to move into augmented versions of their roles [5]. The IMF’s 2024 staff discussion note on generative AI reached a similar structural conclusion using exposure-weighted analysis across advanced and developing economies: advanced economies face higher exposure because their labor forces are concentrated in cognitive-task-intensive occupations, but exposure with strong complementarity potential — the capacity to be augmented rather than replaced — was associated with wage and employment gains, while exposure with weak complementarity was associated with risk of both wage suppression and displacement [6]. Both are institutional analyses synthesizing available evidence, not fresh field experiments, and both explicitly frame their conclusions as conditional on adoption patterns that had not yet played out at the time of writing.

A stack of ring-bound wage ledger printouts on a worktable, the top binder caught half-open as if being set down, pages fanned mid-fall
Figure 2. Wage and occupational-employment data are the aggregate record against which every automation claim must eventually be checked — vendor demonstrations are not evidence of a labor-market effect on their own.Image prompt and art direction by Brecht Corbeel; generation pending.

Scenario one: documented net job loss in a major occupational category by 2035

Horizon: by year-end 2035. Claim: a specific, named major occupational category (at the three-digit U.S. Standard Occupational Classification level or equivalent, covering at least half a million workers) shows a documented net decline in employment level, not merely growth rate, that a peer-reviewed or official-statistics analysis attributes primarily to generative-AI-driven task automation rather than to a business cycle, trade shock, or an unrelated structural cause.

Assumptions this scenario requires: that at least one occupation exists where a large share of economically valuable tasks can be fully substituted rather than merely augmented (unlike the customer-support case, where tasks were reallocated but not eliminated [3]); that firms in that occupation face weak enough complementary-skill requirements, and strong enough cost pressure, to act on that substitutability at scale; and that reinstatement in adjacent new tasks does not absorb the displaced workers within the same statistical category.

ADVERTISEMENT

Observable indicators: BLS Occupational Employment and Wage Statistics showing an absolute decline (not just below-average growth) in employment level for the named occupation across at least three consecutive annual releases; a widening gap between that occupation’s employment level and the BLS’s own 2023-2033 projection for it, which currently projects continued — if slower — aggregate job growth of 6.7 million net jobs across the whole economy over the decade [7]; and at least one peer-reviewed study specifically isolating generative-AI adoption, rather than other factors, as the dominant cause using a design comparable to the exposure-and-adoption methods already established in this literature [2, 1].

Disconfirmation condition: if, by 2035, every occupation with high LLM-exposure scores instead shows stable or growing employment alongside changed task composition and wage dispersion — the Brynjolfsson-Li-Raymond compression pattern generalized rather than a substitution pattern — this scenario is falsified. A single anecdotal round of layoffs attributed to AI in earnings calls or press releases does not confirm it; vendor and executive statements about automation are claims about intent, not verified statistical outcomes, and should be weighted accordingly.

A boxy digital time-use diary tablet on a worktable with a partly filled activity log, one entry row caught being highlighted for input
Figure 3. New-task creation is measured the same way displacement is: through time-use and activity logs that record what workers actually do, updated as new categories of work appear that did not exist before.Image prompt and art direction by Brecht Corbeel; generation pending.

Scenario two: aggregate new-job creation measurably offsets displaced roles by 2035

Horizon: by year-end 2035, using data through 2034. Claim: at the level of total nonfarm employment (not any single occupation), job creation attributable to new, AI-adjacent or AI-enabled task categories measurably offsets job losses in occupations with documented AI-driven displacement, such that aggregate employment-to-population ratios and real median wages do not show a sustained, economy-wide decline traceable to automation.

Assumptions this scenario requires: that new-task creation — the reinstatement side of the Acemoglu-Restrepo framework — operates on generative AI the way it operated on prior automation waves [1]; that the World Economic Forum’s 2025 employer-survey projection of 170 million new roles created against 92 million displaced globally by 2030, for a net of 78 million, holds directionally through 2035 rather than reversing as displacement compounds [8]; and that new roles are counted rigorously as genuinely new task bundles rather than renamed existing jobs, a distinction the WEO survey data — being employer self-report, not independently audited occupational classification — cannot fully guarantee on its own. That gap between a survey claim and a verified occupational count is itself worth flagging explicitly: the WEF figures are the aggregated expectations of over a thousand employers, which is a different kind of evidence than a national statistical agency’s realized employment count.

Observable indicators: BLS employment-projection revisions between now and 2035 showing the healthcare, technical, and newly defined AI-operations occupational groups continuing to add net jobs at a pace consistent with or exceeding the 2023–2033 projections [7]; the emergence of new SOC or ISCO occupational codes for roles that did not exist in the 2020 classification cycle, formally adopted rather than merely described in employer surveys; and stable or rising labor-force participation rates in exposed occupational groups rather than a rising share of workers exiting the labor force entirely.

Disconfirmation condition: if, by 2035, official employment projections are revised downward in successive cycles specifically because new-task job creation fails to materialize at the scale employer surveys anticipated, and if labor-force participation in high-exposure occupations declines alongside displacement without an offsetting rise in adjacent new categories, this scenario is falsified. A single strong jobs report does not confirm it; the claim is about a sustained decade-long pattern, not one data release.

ADVERTISEMENT
A wall-mounted corkboard job-classification chart with printed task cards pinned in a grid, one card caught being re-pinned into a new row
Figure 4. Occupational taxonomies are living documents, redrawn as tasks move between categories — the same institutional process that will decide whether new AI-adjacent job titles get counted as genuinely new work.Image prompt and art direction by Brecht Corbeel; generation pending.

Scenario three: worker-bargaining institutions measurably adapt to automation-era conditions

Horizon: by year-end 2035. Claim: sectoral bargaining structures, new labor-organizing models, or statutory frameworks measurably incorporate AI-specific provisions — task-change notification, retraining guarantees, algorithmic-management limits, or productivity-gain sharing clauses — at a scale that changes bargaining outcomes for a meaningful share of workers in AI-exposed sectors, rather than remaining isolated pilot agreements.

Assumptions this scenario requires: that the early pattern already visible in Europe — where Eurofound’s review found AI-related clauses beginning to appear in collective agreements, generally initiated at company rather than sectoral level as of its most recent review [9] — scales upward from company-level pilots to genuine sectoral or national bargaining structures; and that institutions with the legal capacity for sector-wide bargaining (more common in parts of Europe than in the United States, where collective-bargaining coverage is far lower) use that capacity specifically for automation-era provisions rather than leaving the question to individual employers.

Observable indicators: a measurable rise in the count and coverage of collective agreements containing explicit AI or algorithmic-management clauses, tracked by Eurofound or an equivalent body, moving from company-level instances to sectoral agreements; adoption of statutory AI-at-work provisions (following the pattern of the EU AI Act’s workplace-relevant obligations) in a growing number of jurisdictions; and, in labor markets with low formal bargaining coverage, documented growth in alternative worker-organizing models — sectoral wage boards, occupational licensing bodies, or platform-worker associations — that take on an equivalent function without traditional union structures.

Disconfirmation condition: if, by 2035, AI-specific bargaining provisions remain confined to a small number of company-level pilot agreements without spreading to sectoral structures, and if jurisdictions with low existing bargaining coverage show no growth in alternative worker-organizing models addressing automation specifically, this scenario is falsified — meaning bargaining institutions did not adapt at the scale the scenario requires, whatever else happened to employment levels. Note that this scenario can be true or false independently of scenarios one and two: strong new-task job creation with weak bargaining adaptation, or net occupational loss accompanied by strong bargaining adaptation, are both coherent outcomes and the data should be read for each scenario on its own terms.

A small tabletop audio recorder on a worktable with its input level indicator caught mid-flicker, beside an open field-observation notebook
Figure 5. Bargaining adaptation is documented the same way task change is: through direct fieldwork inside workplaces, not through survey instruments alone — a slower, more contested source of evidence than vendor data.Image prompt and art direction by Brecht Corbeel; generation pending.

Reading these scenarios together without collapsing them into one verdict

These three scenarios are not mutually exclusive, and treating labor-market change as a single yes/no question about “AI and jobs” obscures exactly the distinctions that make the evidence tractable. It is entirely possible for scenario one to resolve true in one narrow occupation (say, a specific back-office data-entry role) while scenario two resolves true in aggregate, because reinstatement elsewhere absorbs the displaced workers without a full sectoral job loss ever showing up in aggregate statistics — this is close to the historical pattern the Acemoglu-Restrepo framework describes for earlier automation waves before the 1980s inflection point [1]. It is also possible for both employment scenarios to resolve in the reassuring direction while bargaining power still erodes, because augmentation without adaptation of the surrounding institutions — as Autor’s synthesis of decades of technological-change literature emphasizes — is compatible with falling worker bargaining power even when headcounts hold steady [4]. Wage and employment counts are not a substitute for measuring who captures the productivity gains, and the current sectoral-bargaining evidence base is thin enough that this question deserves its own tracking independent of the jobs numbers [9].

Why the three scenarios were chosen instead of others

A skeptical reader might ask why these three claims, out of the many that could be built from the same evidence base. The answer is that each targets a different layer of the system the task framework describes, and together the three layers cover the whole causal chain from technology to outcome. Scenario one targets the occupational layer — whether a specific bundle of tasks becomes cheap enough to automate outright rather than merely faster to perform with assistance. Scenario two targets the aggregate labor market layer — whether reinstatement, at the scale of the whole economy, keeps pace with displacement the way it has, on average, across the automation waves of the last two centuries, even though the pace of past reinstatement is itself a matter of active dispute among economists rather than a settled constant [4]. Scenario three targets the institutional layer — whether the rules that decide who captures productivity gains change at all, independent of what happens to raw employment counts. A forecast that only tracked employment levels would miss the possibility, well documented in the history of prior technological transitions, that headcounts hold steady while bargaining power and wage shares shift substantially [4, 1]. Tracking only bargaining institutions, meanwhile, would miss the more basic question of whether specific occupations disappear at all. The three layers are not redundant with each other, and a forecast that collapses them into one headline number throws away the information that would let a reader later check which mechanism actually operated.

It is also worth stating plainly what this framework deliberately excludes. It does not forecast which specific companies will win or lose from AI adoption — that is a competitive-strategy question, not a labor-economics one, and vendor claims about market position should not be mistaken for labor market evidence. It does not attempt a single point estimate of “jobs lost to AI by 2035,” because no credible methodology currently exists to produce one at that level of precision across every occupation simultaneously; the WEF’s aggregate figures are employer expectations gathered by survey, not a measured outcome, and should be read with that provenance in mind every time they are cited [8]. And it does not treat any single quarter’s jobs report, product launch, or executive statement as confirming or disconfirming any of the three scenarios — each is explicitly a multi-year, multi-indicator claim, precisely so that no single data point can be mistaken for a verdict.

What would change this analysis before 2035

This article has deliberately avoided the two unfalsifiable postures described at the start. The honest position, given the evidence assembled here, is that generative AI’s task profile is genuinely novel — it reaches non-routine cognitive work that earlier automation waves could not touch [2] — while the best available field evidence to date shows augmentation and skill-compression effects rather than direct substitution in the occupations studied so far [3], and institutional adaptation on the bargaining side is still at an early, company-level stage rather than a settled sectoral response [9]. None of that licenses a confident prediction in either direction for 2035. It licenses exactly the three scenarios above, each carrying its own horizon, assumptions, indicators, and — the part omitted from almost every popular treatment of this question — the specific condition under which it would be wrong.