Field notes, not a forecast
A claim that “AI will automate 40 percent of jobs” is not a finding you can trace to a method. A claim that “occupation X carries an AI-exposure score of 0.6 on the Felten-Raj-Seamans index, computed by linking ten AI-application categories to fifty-two O*NET work-context items” is a finding you can trace, argue with, and eventually replace with a better one [5]. This guide is about the second kind of claim. It walks through three things a labor economist or workforce analyst actually does — coding the task content of an occupation, running a field study that can distinguish augmentation from replacement, and tracing how the U.S. Bureau of Labor Statistics turns a survey response into a ten-year occupational projection — and it is explicit about what each method proves, what it assumes, and where it breaks.
Throughout, five registers are kept separate. A fact is a number or procedure documented in a primary source. A vendor claim is a productivity or capability assertion made by a company selling the technology in question, reported as a claim, not adopted as true. Analysis is inference from facts, labeled as inference. A scenario is one branch of a conditional future with its assumptions stated. A prediction carries a horizon, the assumptions it depends on, an observable indicator, and a stated condition that would disconfirm it.
Method one: coding what a job is actually made of
The starting unit of modern labor economics is not the “job” or the “occupation” but the task — a unit of work activity that converts inputs into output [1]. This distinction matters because occupations are lumpy statistical categories (the Standard Occupational Classification groups millions of workers into roughly 800 detailed occupations) while automation, augmentation, and skill acquisition happen at the level of what a worker does inside a shift, not at the level of the job title on a badge.
Where the task data actually comes from
The main task inventory used by U.S. researchers is ONET, the Occupational Information Network maintained for the Department of Labor. It did not begin as a research convenience; it replaced the paper-based Dictionary of Occupational Titles in the late 1990s and is rebuilt on a rolling basis. The mechanics, as documented by the ONET Resource Center and reviewed independently by the National Research Council, are:
- Occupations are surveyed in cycles of roughly 100 at a time, chosen to refresh the full set on a multi-year rotation [7].
- For each occupation in a cycle, a sample of job incumbents — not analysts sitting in an office — receives a structured questionnaire. It lists task statements previously associated with the occupation and asks incumbents to rate each on importance and frequency; incumbents may also add tasks the standardized list omits [7].
- Trained occupational analysts separately rate the same occupation on abilities, skills, and work context, after meeting documented training and experience requirements and completing rater training [7].
- The raw task statements — more than 19,000 of them — are aggregated upward: task statements roll into roughly 2,000 detailed work activities, which roll into 325 intermediate work activities, which roll into 41 generalized work activities [7]. This four-level hierarchy is what lets a researcher compare a specific task (“adjust a patient’s IV drip rate”) with a coarse one (“controlling machines and processes”) without conflating them.
The National Academies’ review of this system is worth citing precisely because it is not a promotional document: it examined O*NET’s sampling frame, response rates, and rating reliability as an independent audit, not a vendor claim, and it flagged specific limitations — small occupations are harder to sample reliably, and self-reported importance ratings from incumbents can diverge from what a time-and-motion observer would record [8].
Turning task ratings into an automation-exposure score — worked example
A task-content analysis does not stop at collecting ONET ratings; it recodes them against a theory of what machines are good at. The canonical version, from David Autor, Frank Levy, and Richard Murnane and extended by Autor and collaborators since, sorts tasks along two axes: routine versus non-routine, and cognitive versus manual [1]. The OECD’s cross-country routine-intensity index operationalizes this by building a composite score from ONET items grouped into five categories — routine cognitive, routine manual, nonroutine analytic, nonroutine interpersonal, and nonroutine manual — and standardizing each occupation’s position within its country’s task distribution [9]. Analysis, not fact: an occupation scoring high on routine cognitive and routine manual tasks (data entry clerk, assembly-line inspector) sits in the historically automation-exposed zone; an occupation scoring high on nonroutine interpersonal and analytic tasks (registered nurse, software architect) does not, by this particular measure — which says nothing about future exposure to a different kind of automating technology than the one the index was built to capture.
That gap is exactly what a newer index tries to close. Felten, Raj, and Seamans built an AI Occupational Exposure (AIOE) score by first asking crowdworkers to rate how related each of ten AI application categories (image recognition, language modeling, translation, and so on) is to each of fifty-two human abilities in the ONET taxonomy, then aggregating those relatedness scores up to the task level and finally to the occupation level using each occupation’s ONET ability profile [5]. The authors are explicit that “exposure” is deliberately agnostic about direction — a highly exposed occupation could see its workers augmented, replaced, or both, depending on factors the exposure score does not capture, such as bargaining power, regulation, or the cost of verifying AI output. That caveat is a fact about the method, not a hedge added for this article, and it is the single most commonly dropped qualifier when the index gets cited in press coverage.
What a task-content analysis can and cannot tell you
It can tell you which tasks inside an occupation are, by a stated definition, routine, and it can rank occupations against each other on that basis. It cannot tell you how quickly a given technology will actually be adopted inside a given firm, what a union contract will permit, or whether the tasks freed from automation get reallocated to new output (more units produced) or to new task content within the same job (the worker starts supervising the machine instead). Those require the second method.
Method two: a workplace field study that can separate augmentation from replacement
The task-content approach is observational and cross-sectional: it describes what an occupation looks like at a point in time. To find out what actually happens when a specific automating tool enters a specific workplace, economists need a design that can isolate the tool’s effect from everything else moving at the same time — seasonal demand, management changes, worker turnover. The most credible published example to date is a study of AI assistance in customer-support call centers.
The design, mechanically
Brynjolfsson, Li, and Raymond studied a generative-AI conversational assistant rolled out to support agents at a Fortune 500 software firm, using administrative data on 5,179 agents [4]. The design leans on three features that a field study needs to support a causal claim rather than a correlation:
- Staggered rollout, not a simultaneous switch. The tool reached different teams at different times for reasons unrelated to individual agent performance (a phased technical deployment), which lets the analysis compare newly treated agents against not-yet-treated agents in the same period, controlling for whatever else was changing that month [4].
- An administrative, not self-reported, outcome measure. The dependent variable is issues resolved per hour, logged automatically by the support system, not a survey question about perceived productivity. Self-reported productivity is exactly the kind of measure that flatters whichever technology the respondent was told to evaluate; an administrative log does not know it is being studied.
- Heterogeneity by worker tenure and prior performance, not just an average effect. Reporting only the average treatment effect would have hidden the study’s central finding.
What it found, and why the breakdown matters more than the headline number
The average effect was a 14 percent increase in issues resolved per hour. But the distribution around that average is the finding that actually bears on deskilling versus augmentation: novice and lower-tenure agents gained roughly 34 percent, while the most experienced, highest-performing agents gained close to nothing [4]. The authors’ interpretation — labeled here as analysis, since it goes beyond the raw regression coefficient — is that the tool worked by propagating the observed conversational patterns of the firm’s best agents to everyone else, functioning as a distribution mechanism for tacit expertise rather than a substitute for judgment. The study also recorded a secondary, non-headline result: measured customer sentiment improved and agent attrition fell, which matters for a labor-economics reading because a tool that primarily displaced workers would not be expected to reduce quits.
This is a single firm, a single tool, and a single occupation. Generalizing from it to “AI augments rather than replaces” as a universal law is precisely the overreach this guide’s editorial rules exist to block — the honest statement is that this is one well-identified data point showing augmentation-dominant effects in one high-turnover, script-supported service occupation, and the mechanism (skill transfer from expert to novice workers via suggested text) is plausible in other occupations with a similar structure — expert-to-novice knowledge gaps, real-time text or code suggestion — and not obviously plausible in occupations without that structure, such as tasks that are purely physical or purely judgment-based with no comparable “expert pattern” to transfer.
Building the same design for a different workplace
A practitioner replicating this approach elsewhere needs, at minimum: a rollout with real timing variation across otherwise-comparable units (not a simultaneous company-wide switch, which destroys the comparison group); an administrative outcome measure logged independently of the study; a worker-level covariate — tenure, prior performance, task mix — rich enough to test for heterogeneous effects; and, ideally, a downstream outcome (attrition, error rates, customer complaints) that can distinguish “output went up because quality fell” from genuine productivity gain. Absent any one of these four, a workplace AI study can report a correlation but cannot support a causal augmentation-versus-replacement claim, however large its sample.
Method three: how an official occupational projection is actually built
Task-content indices and field studies both operate below the level the public actually reads: headline numbers like “software developer employment is projected to grow 17 percent over the next decade” come from a third apparatus entirely, the Bureau of Labor Statistics’ Occupational Employment and Wage Statistics (OEWS) program feeding the BLS Employment Projections program. It is worth tracing because the number is routinely treated as if it were a physics measurement, when it is a modeled estimate with a documented, and non-trivial, margin of construction.
The underlying survey
OEWS is a semiannual survey of a probability sample drawn from roughly 8 million in-scope U.S. establishments, stratified by geography, industry, size, and ownership [6]. Two panels of about 186,000-189,000 establishments are contacted each year, one in May and one in November, with the largest establishments sampled with certainty and the rest sampled with probability proportional to size [6]. Because no single panel is large enough to produce reliable estimates for narrow occupations in smaller geographic areas, published OEWS estimates combine six semiannual panels collected over a rolling three-year window — the May 2022 estimates, for example, blend responses collected from November 2019 through May 2022 [6]. That three-year blending is a fact with a direct practical consequence: a published occupational wage or employment estimate is never a snapshot of “right now”; it lags current conditions by up to three years, smoothed across that window.
Filling in the gaps: modeled estimates
Even a sample of over a million establishments cannot directly observe every establishment in the country. For establishments not sampled or not responding, BLS applies what it documents as its MB3 modeling methodology, predicting the staffing pattern (the occupational mix at that establishment) and the associated wages from establishments that were observed, combined with current employment counts from the Quarterly Census of Employment and Wages program [6]. In plain terms: a meaningful share of the published occupational employment total in any area is not counted, it is modeled from the pattern of similar, observed establishments. That is a documented, defensible statistical technique — not a flaw hidden from the public, since BLS states it directly in its methods overview — but it is a fact worth knowing before treating a narrow occupation-by-metro figure as a hand count.
From the wage survey to a ten-year projection
The Employment Projections program takes the OEWS staffing-pattern matrix — which occupations work in which industries, and in what proportions — and projects it forward using separate, documented models of industry output growth, industry employment, and the staffing ratios within each industry, then layers in occupational separations (retirements and career changes that create job openings even in occupations with flat or declining employment). The result reported to the public — “employment in occupation X is projected to grow/shrink Y percent by year Z” — is the composite of at least three modeled components: an industry-level growth forecast, an assumption that staffing patterns within industries stay roughly stable or change along an extrapolated trend, and the underlying OEWS wage-and-employment base described above. Each of those three components can be wrong in ways that do not show up in the headline figure, which is exactly why a serious reader checks the documentation rather than the number.
What this means for reading any “AI will affect N jobs” claim
Put the three methods together and a discipline falls out. A task-content score (method one) tells you which tasks inside an occupation are, by a stated and falsifiable definition, exposed to a described technology. A field study (method two) tells you what actually happened to output, wages, or retention in one real workplace where a specific tool was deployed with real timing variation. A government projection (method three) tells you a modeled extrapolation of industry-level staffing patterns, built from a three-year-lagged survey, under an assumption that the relationship between industries and occupations does not change faster than the model can track — which is precisely the assumption a fast-moving general-purpose technology is most likely to violate. None of the three methods, alone or combined, can tell you what a not-yet-observed technology will do to a not-yet-restructured occupation. Claims that sound precise about that future (“40 percent of tasks in occupation X will be automated by 2030”) are, without exception, scenario or prediction, dressed in the borrowed authority of a measurement.
Wages, bargaining power, and deskilling: separating what is measured from what is argued
The three methods above are the toolkit; here is what they have actually been used to establish, kept separate from what remains contested.
Fact, from the task literature: middle-wage occupations with high routine task content declined as a share of employment across the United States and much of Europe between 1980 and 2010, a pattern documented across many countries using harmonized task measures, not a single national anomaly [9]. Analysis, not settled fact: the leading explanation for that pattern is routine-biased technological change — the idea that automation displaced routine tasks specifically, pushing employment toward both high-wage analytic occupations and low-wage manual service occupations and hollowing out the middle — though the OECD and other researchers have noted that routine-task measures constructed different ways do not always polarize identically across countries, meaning the mechanism is better supported than the claim that it explains all of the observed polarization everywhere [9].
Fact, from the new-task literature: across eight decades of U.S. Census occupational data and patent records, a majority of contemporary employment sits in job titles and tasks that did not exist in 1940, and the analysis explicitly separates two kinds of innovation — those that automate existing tasks (which do not, on their own, generate new work) from those that augment or extend the output of an occupation (which historically have generated new work) [3]. Analysis: the same research finds the locus of new-task creation shifted from middle-wage production and clerical work in the postwar decades toward high-wage professional work and, secondarily, low-wage services since 1980, and that the demand-eroding effect of automating innovation has grown relative to the demand-creating effect of augmenting innovation over the last four decades [3]. That second finding is a documented trend, not a law of nature: it describes what has happened to the ratio of these two effects in the historical data the authors assembled, not a fixed rate at which new work must always replace displaced work going forward.
Fact, from the displacement/reinstatement framework: Acemoglu and Restrepo’s task-based decomposition of U.S. employment growth attributes the slowdown in employment growth over recent decades to a combination of an accelerating displacement effect (automation removing labor from existing tasks, concentrated in manufacturing), a weaker reinstatement effect (fewer new tasks being created where labor holds a comparative advantage), and slower underlying productivity growth than in earlier decades [2]. Analysis: the authors are explicit that this is a decomposition of what already happened, built on assumptions about how to attribute employment changes to each channel, and that the framework’s predictive power for a new wave of automation depends on parameters — the pace of new-task creation, the extent of labor’s comparative advantage in tasks a new technology does not yet perform well — that must be re-estimated for each technology, not carried over from the last one [2].
On bargaining power and wages specifically: none of the primary sources gathered for this guide directly measure how AI-specific automation exposure has, to date, changed union density, strike activity, or wage-setting power in the way the task-content and new-work literatures measure employment shares. That is a real gap, stated plainly rather than papered over with an inferred number — the honest position is that the mechanism by which automation could weaken worker bargaining power (lower switching costs for the employer, if the automating technology substitutes for scarce skill) is theoretically coherent and consistent with the displacement side of the Acemoglu-Restrepo framework, but a specific, verified magnitude for AI-driven bargaining-power effects was not found in the sources checked for this piece, and should not be asserted here as more than a plausible, untested mechanism.
Transition costs: the part every method here is weakest on
Every method above is built to measure a stock or a rate — how many tasks are exposed, how much output per hour changed, how many jobs a decade will add — and each is comparatively weak on the transition itself: how long a displaced worker searches before re-employment, how much of their prior occupation-specific human capital transfers, and how much household income is lost in the interim. The OEWS and Employment Projections apparatus reports separations (retirements, occupation switches) as an input to projecting job openings, but that is a flow used to model future demand, not a study of what happens to the individual worker who leaves. The field-study method (method two) is built around workers who keep their jobs and use a new tool, which is precisely why it cannot speak to the worker whose task bundle is fully automated rather than augmented — a structural blind spot of the design, not a finding within it. A rigorous treatment of transition costs specifically would require a fourth method — worker-level longitudinal earnings data following individuals through a job loss, of the kind used in the broader displaced-worker literature — which falls outside the scope of the three methods this guide verified in depth and is flagged here as a limitation rather than filled with an unverified number.
Two scenarios, stated as scenarios
Scenario A — augmentation-dominant, horizon 2030. Assumption: the expert-to-novice knowledge transfer mechanism documented in the call-center study generalizes to other occupations with a similar structure (large observable performance gaps between experienced and novice workers, task content amenable to real-time text or code suggestion). Observable indicator: new field studies in occupations like coding, paralegal drafting, or technical support would show the same skewed-toward-novices effect size pattern Brynjolfsson, Li, and Raymond found. Disconfirmation condition: if replication studies in structurally similar occupations instead show flat or negative effects concentrated among novices, or show experienced workers’s output declining because verification burden outweighs the suggested-content gain, the augmentation-dominant reading of this mechanism would be wrong for those occupations specifically.
Scenario B — displacement-accelerating, horizon 2030. Assumption: the ratio Acemoglu and Restrepo document — displacement effects strengthening relative to reinstatement effects — continues on its recent trajectory rather than reversing [2]. Observable indicator: the next round of new-work Census analysis, following the Autor et al. methodology, would show a falling (not merely low) share of total employment sitting in job titles newer than ten years old, concentrated in occupations with high measured AIOE exposure. Disconfirmation condition: if the next available new-work dataset instead shows new-task creation accelerating in high-exposure occupations — new job titles specifically built around supervising, correcting, or directing AI systems, which is exactly the kind of augmentation-linked new work the framework predicts should appear — Scenario B’s trajectory claim would be disconfirmed for that period.
Both scenarios are stated deliberately as conditional and falsifiable, not as forecasts this article endorses. The honest summary of the verified record is narrower than either: one well-identified field study shows augmentation-dominant effects in one service occupation; one long-run historical decomposition shows displacement effects strengthening relative to reinstatement effects over four decades in aggregate U.S. data; and the mechanisms behind both are plausible, documented, and not yet shown to generalize to AI specifically in the way either narrative would need to become a confident prediction.
Summary for practitioners
If you are building a task-content analysis, start from O*NET’s work-activity hierarchy, decide explicitly which routine/non-routine and cognitive/manual categorization scheme you are using, and report which occupations were excluded for small sample size. If you are running a workplace field study, insist on staggered rollout timing, an administrative outcome variable, and a worker covariate rich enough to test heterogeneity before claiming an augmentation or replacement effect. If you are citing an official occupational projection, report the three-year survey lag and note that a share of the establishment-level detail is modeled, not counted. And in every case, keep the five registers apart: what a primary source measured, what a vendor asserted, what you inferred, what a scenario assumes, and what a prediction stakes on a specific, checkable indicator.