Is AI actually taking jobs? How to read Anthropic's labor market study

Tied Inc. 日本語で読む

In March 2026, Anthropic published Labor market impacts of AI: A new measure and early evidence (Maxim Massenkoff and Peter McCrory), adding rare primary data to a debate long dominated by speculation. The headline finding: there is roughly a 3x gap between the share of tasks AI could feasibly perform and the share it is actually performing in real workflows, and no detectable impact on unemployment has materialized yet. The one leading indicator worth watching: hiring of young workers (ages 22–25) into highly exposed occupations has slowed by about 14% since ChatGPT launched.

For investors, M&A professionals, and startup operators alike, the study matters beyond labor economics. Its core methodological move — measuring what AI is doing rather than what it could do — translates directly into how you should stress-test an AI startup’s market-size claims, a portfolio company’s headcount-reduction plan, or a target’s assertion that “AI has transformed our engineering productivity.”

What the study measures: observed exposure

Most prior research on AI’s labor market exposure measured theoretical substitutability. The canonical example is Eloundou et al.’s 2023 study (known as “GPTs are GPTs”), which took O*NET — the U.S. Department of Labor database decomposing roughly 800 occupations into their constituent tasks — and estimated, task by task, whether an LLM could meaningfully accelerate the work.

Anthropic’s contribution is to overlay that theoretical layer with actual usage data. The study combines three sources:

  1. O*NET: the government taxonomy mapping U.S. occupations to tasks
  2. Theoretical feasibility estimates (Eloundou et al., 2023): which tasks an LLM could perform
  3. Anonymized Claude usage data (the Anthropic Economic Index): how often usage corresponding to each task is actually observed

The resulting measure is observed exposure — the share of an occupation’s tasks for which AI is being used in practice. Separating “could” from “does” is the study’s central contribution, and the distinction turns out to be quantitatively large.

The key numbers

Four findings stand out (all figures from the original report):

1. The theory–practice gap is wide. For computer and mathematical occupations, LLMs are estimated to be theoretically capable of 94% of tasks, yet observed usage covers only 33%. Business and finance, management, legal, and office administration roles show the same pattern: high theoretical exposure, much lower observed exposure.

2. Observed exposure concentrates in specific occupations. Computer programmers top the list at 74.5%, followed by customer service representatives at 70.1%, with financial analysts among the leaders. Workers in the most exposed occupations skew older, more female, more educated, and higher-paid — the opposite of the common assumption that AI hits low-wage work first.

3. No detectable unemployment effect so far. The authors matched exposure scores to individual respondents in the Current Population Survey (2019–2025) and ran a difference-in-differences comparison between the top quartile of exposed workers and workers with zero exposure. Since ChatGPT’s launch in November 2022, there has been no systematic increase in unemployment among highly exposed workers.

4. But early-career hiring shows a leading signal. Using the CPS panel structure to track workers aged 22–25 who start new jobs, the authors find hiring into high-exposure occupations has slowed by roughly 14% since ChatGPT launched — consistent with the Stanford “canaries in the coal mine” study (Brynjolfsson et al., 2025), which found a ~13% employment decline for young workers in exposed roles. A regression against Bureau of Labor Statistics occupational projections also finds that each 10-percentage-point increase in observed exposure is associated with a 0.6-point lower projected growth rate through 2034.

How much should you trust these numbers?

The authors are candid about limitations, and external critics have added more. Three deserve attention:

  • The usage data reflects Claude’s user base, not economy-wide AI adoption. Usage through ChatGPT, Copilot, or internal enterprise tools is invisible to this measure.
  • Task-level exposure ignores how tasks fit together. Gans and Goldfarb (2025) note that if jobs follow an O-ring structure — where value requires every task to be completed — automating a subset produces no employment effect. Autor and Thompson (2025) argue the expertise required for the remaining tasks determines whether a job is hollowed out or upgraded.
  • “No effect yet” is not “no effect.” The analysis covers the first three years of adoption; organizational restructuring moves slower than tool adoption.

Read correctly, the study is not proof that AI won’t affect employment. It is a methodological warning: never read employment impact directly off theoretical substitutability.

Four implications for investors and technical due diligence

The three-layer distinction — feasible (could do), observed (is doing), realized (shows up in financials and employment) — is directly usable in deal work.

1. Stress-test AI startup TAM claims across the three layers

A pitch that says “X% of this occupation’s work is automatable, therefore the market is $Y billion” is a feasible-exposure argument. This study quantifies the friction between feasible and observed at roughly 3x — friction composed of workflow integration costs, output verification, organizational readiness, and regulatory constraints.

That gap converts into two diligence questions. First: which friction does this product remove, and how? Reducing adoption friction — embedding into existing workflows, lowering verification costs — is where durable differentiation lives, a point we develop in where AI startups build competitive advantage. Second: is the company’s traction narrated in feasible terms (“we can handle these workflows”) or observed terms (“here are the tasks that actually moved to our product in customer logs”)? The latter reflects a materially more validated thesis.

2. Discount headcount-reduction assumptions in value-creation plans

Business plans increasingly embed assumptions like “AI adoption reduces personnel costs by Z%.” The economy-wide evidence says displacement at statistically visible scale has not happened yet. Plans grounded in theoretical automatability should be recomputed against observed-level adoption (roughly one-third of the theoretical figure) plus explicit integration costs. For post-merger integration planning in particular, build a case where labor savings arrive one to two years later than the management case assumes.

3. Verify “we use AI” claims with observed evidence, not declarations

When a target claims AI has transformed its development productivity, the diligence object is usage traces, not statements: seat counts versus actual active usage, monthly API spend trends, and the share of AI-generated code visible in commit and review history. As with the added evaluation criteria in how AI changes technical due diligence, distinguishing “could” from “does” is what separates rigorous assessment from vibes. This observed-first stance is also how we approach technical DD engagements for investors through TiedPro.

4. Read slowed junior hiring as a talent-pipeline risk

The finding that early-career hiring is slowing in exposed occupations adds a long-horizon question to organizational assessment. Substituting AI for junior hires is cost-efficient this quarter and starves the mid-level talent supply three to five years out. In engineering organization health assessments, we recommend examining age and experience distribution together with AI adoption levels — a barbell-shaped org chart with no juniors is a deferred liability.

FAQ: common misreadings

Q1. Did the study conclude AI won’t cause unemployment?

No. It reports that no unemployment increase was detectable in U.S. data through 2025. The slowdown in young-worker hiring is a leading indicator, and the authors explicitly leave open that impacts may materialize later.

Q2. Is high observed exposure a red flag for businesses that depend on those occupations?

High exposure means AI is being used on those tasks — not that the jobs are disappearing. Whether AI complements or substitutes depends on task composition and the expertise required for what remains. Exposure is a screening indicator; business impact requires occupation-specific workflow analysis.

Q3. How should non-U.S. readers translate these findings?

The data is American. In labor markets with stronger employment protections and thinner external mobility — Japan being a prime example — the same adoption level tends to surface as hiring freezes, internal redeployment, and shrinking graduate intakes rather than unemployment. The observed-exposure framework travels across markets; the form the impact takes depends on labor market institutions.

Summary

The study’s value is pulling the AI-and-jobs debate back from theoretical possibility to observed data. The translation into investment practice is simple: distinguish feasible, observed, and realized, and always identify which layer a business plan or cost-reduction thesis is actually standing on. That single discipline avoids most of the over- and underestimation currently priced into AI-related deals.

References

Tied Inc.

Tied Inc.

Tech-leadership advisory for investors and operating companies. We support technical due diligence, value-up engineering, and strategic technology decisions across the investment lifecycle.

Get in touch →