Enterprise AI case studies need measured baselines before productivity claims become workforce policy
An ILO brief combines 21 purposively selected enterprise interviews with a survey of 1,591 professionals in China. Its self-reported gains are useful hypotheses, but workforce decisions need task baselines, comparison groups and job-quality measures.

What happened
The ILO reports interviews with 21 AI-using or AI-planning enterprises and a survey of 1,591 professionals across Chinese sectors. Some firms reported large productivity changes, but the brief says most lacked systematic impact frameworks and the firm figures were not independently verified.
Why it matters
Purposeful case selection and self-reporting can reveal mechanisms and implementation problems, not population effects. Turning the headline gains into workforce targets would hide selection, concurrent process changes, displaced work and unmeasured job quality.
The ILO research brief combines in-depth interviews with 21 enterprises and a survey of 1,591 professionals in China. The firms span manufacturing, finance, business services, construction, education, media and travel, from an eight-person start-up to a company with 270,000 employees. Every interviewed firm was already using AI or had concrete adoption plans.
The release reports striking examples: one insurance operation increased daily handled issues from 6,000 to 15,000 for 300 service employees; another reduced recruitment cycle time from 30 to 13 days; and a manufacturing facility reported a 30% efficiency increase. The same source explicitly says those figures were self-reported and not independently verified. The 21 firms were purposively selected, so firm findings cannot be generalised statistically to all Chinese enterprises. Most studied firms lacked systematic impact frameworks beyond conventional productivity measures.
Convert a case claim into a testable baseline
A case study can identify where to look. It cannot set a workforce target without a denominator and counterfactual. Before expanding an AI workflow, record at least four weeks of task volume, handling time, error and rework, queue age, escalation, staffing mix and worker time spent on hidden coordination. Freeze definitions so a resolved issue does not silently become a shorter interaction or an automated deflection.
Then compare like with like. Use a phased rollout, matched teams or interrupted time series and record seasonality, demand shifts, hiring changes, process redesign and other automation introduced at the same time. A jump from 6,000 to 15,000 handled issues may reflect new routing, more short contacts or transferred follow-up work. The purpose is not to dismiss the number, but to learn which mechanism produced it and whether it persists.
Measure the distribution of work, not only output
The ILO brief says reported gains concentrate in repetitive and data-intensive tasks while firms also cite resistance, skills gaps, output quality, security, regulation and integration. For each task, map what disappears, what is added and who becomes accountable for checking. Track review time, exception complexity, exposure to difficult customers, schedule control, learning opportunities and income alongside throughput.
Survey results are also perception data. Among professionals, 56% viewed adoption as inevitable, 47% believed AI creates more jobs than it displaces and 39% expected income declines. These answers can guide questions and segmentation; they are not forecasts. Link them to observed task changes, vacancies, wages, training access and exits before using them to justify policy.
The counterargument is that rigorous measurement is expensive and can delay useful deployment. A minimum viable evidence plan can be small: one stable baseline, one comparable group, a predefined outcome set and a weekly review of harms and workarounds. Stop or redesign when output rises but material error, unpaid verification, inequality or turnover worsens. Expand only after the mechanism and trade-offs are reproducible.
Ownership matters as much as method. Assign a named operations owner for each metric and a worker representative or equivalent channel for disputed interpretations. Preserve raw definitions and sampling rules in the decision memo. When management changes a target after seeing results, label it exploratory instead of rewriting the baseline. This discipline makes negative or mixed findings usable and reduces pressure to turn an adoption story into a success story.
The decision for a workforce leader is therefore to treat the ILO cases as hypotheses, not targets. Select one task family, instrument the current process, pilot with a comparable group and publish the limits internally. The AI exposure explorer can help separate task exposure from job outcomes. Productivity becomes decision-grade evidence only when the organization can show what changed, relative to what baseline, for whom and at what cost.