Task-based econometrics, workplace ethnography, and official labor statistics each measure automation's effects on work differently — and none of them can replace the other two.

Task-based models decompose an occupation into its component tasks and ask which ones move to capital and which stay with labor. — Image prompt and art direction by Brecht Corbeel; generation pending.
Three distinct research traditions try to answer the same question — what does automation do to labor, skill, and pay — and they answer it in incompatible units. Task-based econometric modeling decomposes occupations into tasks and estimates displacement and reinstatement effects from variation across regions and industries. Workplace ethnography and qualitative field studies follow specific workers through specific automation deployments and surface deskilling, reskilling, and bargaining dynamics that never appear in aggregate data. Official labor statistics programs track employment, wages, and task content at national scale but cannot establish what caused an observed shift. This article compares the three approaches on causal inference power, granularity, and forecasting ability, keeps fact separate from vendor claim, analysis, scenario, and prediction, and declines to name a winner among methods built to answer different questions.
Ask three different labor researchers what automation is doing to a given occupation and you will get three different kinds of answer, built from three different kinds of evidence, and none of them will be wrong. One will hand you a regression coefficient estimated from decades of cross-industry variation in robot adoption. One will hand you a year of field notes from inside a single warehouse. One will hand you a spreadsheet of national wage percentiles updated twice a year. Each is measuring something real. None of them is measuring the same thing.
This is not a failure of the field to converge on a method. It is what happens when a question — “what does automation do to work” — actually bundles together several distinct questions that require different instruments to answer: what causes what, at what level of detail, and how far ahead anyone can see. This article compares the three major approaches that dominate the modern study of labor and automation — task-based econometric modeling, workplace ethnography and qualitative field study, and official labor-statistics trend analysis — on the dimensions that actually separate them, rather than treating them as competitors for the same prize.
Task-based econometric modeling starts from the observation that an occupation is not one thing you can automate or not automate; it is a bundle of tasks, and technology acts on tasks, not titles. Daron Acemoglu and Pascual Restrepo formalize this in a task-allocation framework in which automation moves specific tasks from labor to capital, which mechanically reduces labor’s task content and share of value added, while the same technological change can also create new tasks in which labor holds a comparative advantage, which pushes back the other way [1]. The empirical version of this approach exploits variation — across commuting zones, industries, or countries — in exposure to a specific automation technology, most famously industrial robots, and estimates how employment and wages moved differently in more-exposed versus less-exposed labor markets. David Autor’s synthesis of decades of this tradition traces how automation and trade jointly hollowed out the middle of the U.S. occupational structure, polarizing urban labor markets toward high-skill and low-skill work at the expense of the routine, codifiable tasks in between [2]. A related strand tries to forecast exposure directly: Carl Benedikt Frey and Michael Osborne built a computerizability score for hundreds of occupations by rating the task content of each against what algorithms and robotics could plausibly do, concluding that a large share of U.S. employment sat in occupations at high computerization risk over a one-to-two decade horizon [3]. That estimate is a scenario built on an engineering judgment about technical feasibility, not a measured outcome, and it should be read as exactly that — a conditional forecast, not a fact about what did or will happen.
Workplace ethnography and qualitative field study starts from the opposite end: not a national sample, but one factory floor, one hospital ward, one warehouse, followed closely enough to see how an automation technology actually gets used, resisted, gamed, or reinterpreted by the people working alongside it. Shoshana Zuboff’s extended ethnography of a paper mill, an insurance office, and a bank documented a distinction the aggregate literature had no vocabulary for at the time: information technology can either “automate” a task, stripping the tacit knowledge that used to live in a worker’s hands, or “informate” it, surfacing the underlying process as data that workers and managers can both act on — and which of the two happens is a matter of workplace politics, not the technology’s intrinsic nature [7]. More recent field research on algorithmic management extends this tradition to platform work and app-mediated scheduling. Katherine Kellogg, Melissa Valentine, and Angèle Christin’s review of workplace field studies frames algorithmic systems as a new terrain of contest between managerial control and worker response — restriction, recording, resistance, and reengineering are the four mechanisms they identify running through dozens of individual studies of algorithmically managed workplaces [9]. This is where deskilling, reskilling, and shifts in bargaining power show up first, because they are observable only at the resolution of an actual workplace: a dispatcher who no longer needs to know the roads because the app routes for them; a nurse whose clinical judgment is now cross-checked, and sometimes overridden, by a decision-support flag.
Official labor-statistics trend analysis is the least glamorous and most load-bearing of the three. The U.S. Bureau of Labor Statistics’ Occupational Employment and Wage Statistics program surveys roughly 1.1 million establishments a year to produce employment and wage estimates for around 830 occupations across the entire country, twice a year, at metropolitan-area resolution [5]. The O*NET program, run for the Department of Labor, surveys job incumbents, occupational experts, and trained analysts to produce roughly 500 standardized ratings — task importance, task frequency, required knowledge, work context — for every occupation in its 900-plus occupation taxonomy, refreshed on a rolling quarterly and annual schedule [6]. These programs do not test a hypothesis about automation; they describe the labor market as it stands, continuously, at a scale no research team could replicate on its own budget, and every task-based model and most field studies ultimately borrow their occupational categories, task lists, or wage baselines from exactly this infrastructure.
Task-based econometric models are built for causal inference and that is their central advantage. By finding a source of variation in automation exposure that is plausibly unrelated to other things also affecting a local labor market — the pace at which robot prices fell in a particular industry, say — this approach can estimate a defensible answer to “how much of the change in wages or employment in this labor market is attributable to this specific technology,” net of general economic trends. That is a genuinely hard question and task-based models are close to the only tool that answers it with a number attached to an uncertainty interval. The cost is that the estimate is only as credible as the exclusion restriction behind it, and it describes an average effect across many labor markets, which can average away enormous local variation: a robot that displaces machinists in one plant while data show net job growth in the industry overall is not a contradiction, it is what an average conceals.
Workplace ethnography cannot produce that kind of causal estimate and does not try to. Following one workplace through one deployment cannot tell you what would have happened to wages nationally in a counterfactual world without the technology. What it can do is establish mechanism — the actual causal pathway through which a system changes what a worker does, which task-based models can only infer indirectly from an aggregate outcome. Zuboff’s automate/informate distinction is a mechanism claim that no regression on employment counts could have generated, because it depends on watching how authority and knowledge actually moved inside one organization [7]. Field studies of algorithmic management similarly identify specific tactics — workers gaming a scoring algorithm, managers using a system’s data trail to justify decisions they had already made — that show up as noise, not signal, in a national dataset [9].
Official statistics establish neither causal estimate nor mechanism. Their contribution is descriptive: they can tell you, with confidence no sample survey could match, that median wages in an occupation rose, fell, or held flat over a given period, and what the task content of that occupation was rated at in a given year. They cannot tell you why, because the BLS and O*NET programs are not designed as natural experiments — they are censuses and structured surveys of the labor market as it is, not comparisons against a counterfactual labor market as it would have been [5] [6].

Figure 1. Time-and-motion observation at a single workstation measures how long a task actually takes, at a resolution national statistics never reach. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
Granularity runs in the opposite direction from causal power. Task-based models typically operate at the level of an occupation or a broad task category across an entire economy or a large region; that scale is exactly what makes the statistical variation usable, but it means the unit of analysis is often coarser than any single job actually experienced by any single worker. Acemoglu and Restrepo’s framework works with task shares aggregated across an industry, not the daily task list of one employee [1].
Workplace ethnography sits at the opposite extreme: it can resolve detail no other method reaches, down to a specific shift, a specific interaction between a specific worker and a specific piece of software. That granularity is also its limit — a finding from one warehouse in one company does not automatically generalize to a different warehouse using a different vendor’s system, and the field studies literature is explicit that context (union presence, local labor-market tightness, sector norms) determines which of the four mechanisms Kellogg, Valentine, and Christin identify actually dominates in a given site [9].
Official statistics land in between, and their granularity is a design choice rather than an accident. O*NET’s roughly 500 ratings per occupation are pitched at the occupation level, not the individual worker or firm, because the program’s purpose is comparability across the whole economy, not depth on any one job [6]. The BLS wage program reports down to the metropolitan-area and occupation level but not below it — it will tell you the median wage for registered nurses in a given metro area, not what happened inside any specific hospital [5].

Figure 2. Workplace ethnography follows a worker through an actual deployment, capturing dynamics that never register in an aggregate count. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
None of the three approaches forecasts well in the sense of predicting a specific future employment number with precision, and it is worth separating fact from forecast carefully here, because this is where the literature is most often misread.
Frey and Osborne’s computerization-susceptibility scores are frequently cited as a prediction that a specific share of jobs would disappear; what the study actually produced was a scenario — an engineering-judgment estimate of which task bundles were technically automatable given the capabilities visible at the time of writing, not an estimate of how many of those technically feasible automations would actually be adopted, at what pace, under what cost and regulatory conditions [3]. James Bessen’s demand-based model makes exactly this point formally: whether automation destroys net employment in an industry depends heavily on how elastic demand is for that industry’s output, and the same technical automation potential produced rising employment in some historical industries and falling employment in others, purely as a function of whether cheaper output expanded the market or saturated it [8]. Technical feasibility, in other words, is not adoption, and adoption is not a labor-market outcome — three separate steps that task-based forecasting work often compresses into one headline number.
Ethnography forecasts even less well at the aggregate level by design — a study of one workplace’s adaptation to one system does not extrapolate to a national trajectory — but it is often the first place a genuinely new adaptation pattern becomes visible, well before it shows up as a statistically detectable trend in national data. Official statistics forecast worst of all in the causal sense, because they are backward-looking by construction: OEWS and O*NET describe the labor market that already existed at the time of the survey round, refreshed twice a year and quarterly respectively [5] [6]. Their value for anticipating change is indirect — they supply the task inventories and wage baselines that task-based models then use to build their own scenarios, and a widening or narrowing gap between successive rounds of official data is often the first confirmable signal that a scenario sketched years earlier by econometric or ethnographic work is starting to materialize.

Figure 3. Official wage and employment statistics track national aggregates at scale, at the cost of the causal detail a single field study can supply. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
It is worth being explicit about which claims in this literature are fact, which are vendor or advocate assertion, which are analysis, which are scenario, and which are prediction, because the four get blended constantly in popular coverage of “automation and jobs.”
Fact: the BLS OEWS program surveys about 1.1 million establishments annually and reports employment and wage estimates for about 830 occupations [5]. Fact: O*NET’s roughly 500 ratings per occupation are drawn primarily from job incumbents rating the importance and frequency of tasks they actually perform [6]. Fact: Zuboff’s ethnography covered three specific organizations — a paper mill, an insurance office, a bank — over several years in the late 1970s and early 1980s [7].
Analysis: Acemoglu and Restrepo’s claim that automation’s effect on labor demand depends on the balance between displacement in existing tasks and reinstatement in newly created ones is a theoretical and empirical argument built from data, open to being revised by further evidence, not a brute fact [1]. Autor’s account of labor-market polarization is likewise analysis synthesizing several decades of data, not a single measured quantity [2].
Scenario: Frey and Osborne’s occupational computerization scores are an “if this task bundle is technically automatable, here is which occupations are exposed” scenario, not a forecast of realized job loss [3]. Prediction, properly qualified: any claim in this space about what will happen to a specific occupation’s employment or wages over the next decade needs a stated horizon, the assumptions it depends on (adoption cost, regulatory posture, demand elasticity per Bessen’s framework), the observable indicators that would confirm it is on track, and the condition under which it should be considered wrong [8]. The 2020 MIT Task Force on the Work of the Future’s central conclusion is itself best read this way: it argued that the immediate labor-market effects of automation were being overstated relative to slower-moving, policy-driven trends in unequal wage growth, an analytical claim about relative weight of causes rather than a single verifiable fact, resting on a synthesis of exactly the three literatures compared here [4].

Figure 4. Standardized task inventories such as O*NET turn incumbents' own judgments about their work into ratings comparable across occupations. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
The same underlying phenomenon — a technology changing what a worker needs to know and how much leverage they have over their own terms of work — looks different depending on which instrument measures it. Task-based models pick it up only indirectly, through a shift in the task share or wage premium associated with an occupation’s skill requirements as rated in O*NET or classified in the Standard Occupational Classification system; a falling task share for what the model codes as “routine cognitive” work is consistent with deskilling, but the model itself does not observe any worker losing a skill, only an aggregate task allocation shifting [1] [6].
Field studies observe the mechanism directly and are the primary source for terms like deskilling and reskilling actually being used with evidentiary backing rather than as slogans. Zuboff’s automate/informate distinction is precisely a claim about which of these two things happens to a given worker’s knowledge under a given implementation choice [7], and Kellogg, Valentine, and Christin’s synthesis of algorithmic-management field studies documents specific bargaining responses — workers restricting the data an algorithm can see, recording their own interactions as a counter-record, openly resisting a metric, or reengineering the system’s rules from inside — as the observed forms that changed bargaining power actually takes on the ground [9].
Official statistics see the downstream wage and employment consequence of all of this, aggregated across every workplace where it happened, but cannot distinguish a wage change driven by deskilling from one driven by, say, a shift in industry composition, because the wage series and the task rating series are both point-in-time snapshots, not records of causal pathways [5] [6]. Reading the three together is not redundant: the econometric estimate tells you the size of an aggregate shift and the conditions under which it holds; the field study tells you the mechanism that produces a shift of that kind inside one real workplace; the statistics program tells you whether, and where, the shift is actually showing up at scale over time.

Figure 5. Forecasts about which tasks automate next are scenario work, not measurement, and every method disagrees about how far ahead it can see. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
It is tempting to rank these three approaches by rigor, and the ranking usually favors whichever one the reader is most familiar with — economists reach for the identification strategy, sociologists and organizational researchers reach for the field study’s mechanism, statisticians and policy staff reach for the survey program’s coverage. Each objection to the other two methods is usually correct as far as it goes. Task-based models really do average over workplace-level variation that matters. Field studies really cannot tell you whether their site is typical. Official statistics really cannot establish what caused the number they report. None of that is a defect to be engineered away; it is the tradeoff each method accepted in exchange for what it is actually good at — causal estimation at scale, mechanism at a single site, or description at national coverage and update frequency.
The disagreements that persist among researchers who study labor and automation — how much of recent wage polarization is due to automation versus trade versus institutional change, how far current task-content projections should be trusted, whether a given deployment counts as deskilling or as a shift in what skill even means for that job — are not always resolvable by collecting more of the same kind of data. They often persist because the three literatures are measuring different aspects of the same event and each is, correctly, unable to answer the questions posed to the other two. The 2020 MIT Task Force report explicitly built its policy argument by triangulating econometric, survey, and case-study evidence rather than treating any one stream as sufficient on its own [4], which is the working posture this comparison recommends: read the coefficient, read the fieldwork, read the survey update, and expect each to correct a blind spot in the other two rather than expecting any single one of them to settle the question of what automation is doing to work.
Originally published at https://absolutedigitalpublishers.com/articles/comparing-the-main-approaches-to-labor-automation-and-human-capability.