Skip to content

AI and Employment: What the Data Shows So Far

4 min read · updated August 3, 2026

Almost every headline number about AI and jobs is an exposure estimate being reported as a prediction. The difference is not a technicality: it is the difference between counting tasks a model could plausibly do and observing that somebody lost work.

Four different questions

“Will AI take jobs” is four questions with four evidence bases, and the answers do not follow from one another:

  • Task substitution. Can a model perform this task at acceptable quality? Answerable now, by testing.
  • Adoption. Is it being used in production for that task? Answerable by survey and by firm data, and much lower than capability at any given time.
  • Labour demand. Does the firm employ fewer people, or different people, as a result? Answerable in principle, with an attribution problem.
  • Wages and distribution. Who captures the surplus? The hardest question and the one most policy cares about.

A study answering the first is routinely reported as answering the third. That single substitution accounts for most of the noise.

What each design can establish

Exposure indices

The method: decompose occupations into tasks using an occupational task database, score each task for whether a model could do it or assist with it, aggregate to an occupation score. The output is a ranking of occupations by how much of their task content is in principle model-addressable.

What it establishes: a plausible ordering of where to look. What it does not establish: that anything will happen. The score contains no adoption, no cost, no organisational friction, no regulation, no quality threshold and no demand response. Exposure is also symmetric — a task a model can assist with may make the worker more productive and more employable rather than less, and the index cannot distinguish those cases. When one of these studies is reported as a forecast of job losses, the forecast was added by the reporting.

Firm-level experiments and staggered rollouts

The method: give a tool to some workers and not others, or roll it out in a staggered order across teams, and compare output. This design has been used in customer support, professional consulting tasks and software development, and it is the strongest internal-validity evidence that exists in this area.

What it establishes: a causal effect on a measured output metric, for those workers, on those tasks, over that period. What it does not establish: that the effect persists after novelty and workflow adjustment, that it holds for tasks that were not chosen for the study, that measured output corresponds to value delivered, or that the firm changes its headcount. A recurring and important limitation is that the output measure is usually something countable — tickets closed, lines written, documents produced — and countable is not the same as good.

Labour market panel data

The method: administrative or survey data on employment, hours and wages, compared across more- and less-exposed occupations or regions over time. This is the design that can actually answer the third and fourth questions.

What it does not establish easily: attribution. Adoption is not randomly assigned; firms that adopt differ from firms that do not, and both respond to the same macro conditions. Interest rates, sector cycles and post-pandemic normalisation move the same series. Difference-in-differences designs help and depend on a parallel-trends assumption that is questionable when the treated group is defined by exposure to the very technology under study.

Job-posting text analysis

Fast, high-frequency and increasingly popular: count postings, or count skill mentions within them. The caveats are large — a posting is not a hire, posting behaviour responds to labour market tightness for unrelated reasons, and the language of postings follows fashion. Useful as a leading indicator, weak as evidence of outcomes.

Why the aggregate signal is hard to see

Even if the effect is real and large, several mechanisms delay and disguise it in aggregate statistics. Adoption inside firms lags capability by years because it requires process change, not procurement. Effects appear first in hiring rates rather than in separations, which means they show up as jobs not created — the hardest thing to observe. Displaced workers reallocate, so an employment count nets two opposite flows. Productivity gains in services are notoriously hard to measure because output quality is hard to measure. And firms have strong reasons to attribute layoffs to whatever narrative is currently favourable, which contaminates the most-cited data source of all: company statements.

Established, contested, unknowable

Reasonably established. Controlled studies have found substantial throughput gains on well-specified writing, support and coding tasks. A finding that recurs across several of them is that gains are larger for less experienced or lower-performing workers on the measured task — a compression effect. This is one of the more robust patterns in the literature precisely because it appeared in several independent settings.

Contested. Whether task-level gains translate into firm-level productivity; whether they survive quality review over longer horizons; whether adoption reduces headcount or reallocates it; and whether the compression effect is durable or a novelty of early tooling. Reasonable economists reading the same studies disagree, and the honest description is that the evidence is mixed rather than that one side is ahead.

Not knowable now. Net employment effects over a decade. Every confident statement about this is a model with assumptions, and the assumptions are doing all the work.

Reading a claim

Four questions handle nearly everything: what was the design; what was the outcome measured — capability, adoption, output, or employment; over what horizon; and for whom. A claim that cannot answer those is a projection, which is a legitimate genre as long as it is labelled.

This is a fast-moving evidence base and the state of it will change. Treat the framework as durable and any summary of findings, including this one, as something to re-check against recent work.

AI and Employment: What the Data Shows So Far · Multigrid