ChatGPT Enterprise usage is growing, but usage is not organisational transformation
OpenAI's administrative data shows how activity spreads across firms, roles and tasks; it does not measure productivity, completed work or role redesign.

What happened
An OpenAI working paper linked ChatGPT Enterprise account activity through March 2026 to job-title, task-classification and US public-company financial data.
Why it matters
The study offers unusually detailed product telemetry, but its measures stop at access and activity. Leaders still need evidence connecting use to work quality, outcomes, routines and accountability.
OpenAI's working paper examines workplace use through 31 March 2026 by joining ChatGPT Enterprise account records to usage, administrative job titles, automated task classifications and financial data for a selected sample of US public companies. It is a large first-party view of one product, not a representative account of how all organisations use AI.
What the data measures
The core panel follows organisations from the week they adopt a paid, centrally administered ChatGPT Enterprise workspace. Active workspaces with no observed activity remain in the data with zero use. The measures are messages sent, weekly active users and output tokens, including ChatGPT and Codex. During the study period, however, the paper says output was overwhelmingly generated by ChatGPT and related non-agentic tools. The study therefore should not be presented as evidence of a general shift to autonomous agents.
For worker-level analysis, the authors select 1,764 organisations with usable industry and job-title information and an active observation 26 weeks after adoption. That sample contains 17,446,551 messages. Job titles are classified with GPT-5-mini into broad functions and seniority groups. Coverage is incomplete, and titles are observed at one point rather than tracked as roles change.
A later task-classification sample contains 973 organisations and 8,696,657 messages. An automated classifier assigns each turn to one of 60 task categories. It was available only from 30 October 2025 and was evaluated on an internal benchmark; the paper does not report external validation metrics. Researchers did not manually review individual customer messages.
What the paper finds
Aggregate output tokens rose approximately sevenfold between June 2025 and March 2026. Output also rose roughly fourfold within the fixed cohort of organisations that had adopted by June 2025, so growth was not only the result of adding customers. This is an adoption-and-intensity result. Tokens remain a volume proxy, not a measure of useful work.
Use appears across functions and seniority levels. Among active users, early-career workers and trainees sent about eight to nine more messages per week than the average active user in the same firm, while managers, directors and executives sent fewer. That comparison does not establish a role-specific adoption rate because the study lacks the denominator of all employees in each role. Nor does it show that high-message users saved time or produced better work.
The classified conversations span documentation and technical writing, technical digital work, communication, research, planning, data analysis, legal work and finance. This maps where users bring requests to ChatGPT; it does not observe the downstream work product, whether the output was accepted, or whether a workflow or role changed.
Do not turn association into impact
In the public-company sample, ChatGPT Enterprise adopters were larger, more valuable and more intensive in R&D and selling, general and administrative investment than non-adopters. The paper explicitly treats these as associations, not causal effects. Its “non-adopter” group can include firms using competing systems, APIs, internal tools or personal ChatGPT accounts. Early adopters are consequently not a neutral comparison group.
The companion OpenAI publication page frames enterprise AI as moving from assistance towards execution, but it combines this working paper with a separate Enterprise Signals report and later product data. That broader framing should not be attributed to this study alone.
The practical decision is to instrument the missing steps. A responsible adoption dashboard should distinguish provisioned access, active participation, task attempted, output accepted, quality, time or cost changed, and accountable human review. Human-in-the-Loop AI and AI Output Verification are therefore relevant controls, but this single first-party paper is not enough to revise either Atlas record or a hiring plan.