How Organizations Use AI: Evidence from ChatGPT

Paper · arXiv 2608.12236 · Published August 12, 2026
AI at Work

Abstract We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of new firm adoption and growing intensity among existing adopters. Second, U.S.-based public company adoption is concentrated among larger, more valuable, and more R&D- and SG&A-intensive firms. Third, active use within adopting firms spans job functions and seniority levels, with especially high usage intensity among early-career workers. Fourth, ChatGPT Enterprise usage encompasses a broad range of knowledge work tasks, including writing, technical work, communication, and information synthesis. In aggregate, these results suggest that firms differ widely in the speed, breadth and purpose of their enterprise AI adoption, and that they are still actively learning how to integrate AI into organizational workflows.

Introduction. Generative AI systems can perform a growing range of economically valuable tasks (Eloundou et al. 2024; Patwardhan et al. 2025), and individuals use consumer-facing generative AI chatbots for many work-related and personal activities (Chatterji et al.

2025; Handa et al. 2025). However, less is known about firm AI adoption: which workers account for observed use, how intensively active users engage with generative AI, and for which tasks. Understanding these patterns is important for interpreting recent findings about the impact of AI on productivity and employment (e.g., Brynjolfsson et al. 2025a; Brynjolfsson et al. 2025b). Most evidence on firm AI adoption comes from worker and firm surveys (McElheran et al. 2024; Bick et al. 2026a; Yotzov et al. 2026; Bonney et al.

2026). Although surveys provide broad coverage and can capture non-use, adoption barriers, and organizational context, self-reported usage is typically less detailed and may suffer from imperfect recall or reporting biases.

This paper studies workplace AI adoption using internal data from ChatGPT Enterprise, OpenAI’s centrally administered workplace product.2 We first examine firm adoption by decomposing enterprise growth into within- and between-firm components and linking ChatGPT Enterprise accounts to public company financial data. We then study within-firm heterogeneity by combining usage data with employee job title information and message-level task classifications, which tells us how usage intensity and task adoption varies across worker groups six months after organizational adoption.

We document four stylized facts about enterprise AI adoption and usage. First, enterprise AI usage is growing rapidly, reflecting both increased use among existing customers and the arrival of new adopters. Aggregate output tokens produced by ChatGPT Enterprise customers grew roughly sevenfold between June 2025 and March 2026, and by nearly fourfold within a consistent cohort of firms that adopted between January 2024 and June 2025. Thus, about half of the growth in token consumption over this period occurred within already-adopting firms. Second, among U.S.-based public companies, ChatGPT Enterprise adopters are larger, more valuable, and more R&D- and SG&A-intensive than non-adopters. This pattern suggests that early enterprise AI adoption is associated with greater prior investment in intangible and organizational capabilities.

Third, usage within adopting firms is broadly distributed across job title classes and seniority levels but with heterogeneous intensity. For example, marketing and communications workers send more messages than executives, and early-career workers send many more messages than more senior employees. Fourth, ChatGPT Enterprise use spans many different tasks across workers and organizations, rather than being concentrated in a single workflow. The most common use cases are writing, communication, and information synthesis, but usage is also common in tasks such as research, planning, data analysis, legal and regulatory work, finance, and many other applications. This breadth is consistent with generative AI functioning as a general purpose technology for knowledge work (Bresnahan and Trajtenberg 1995; Bresnahan 2024; Eloundou et al. 2024).

Taken together, these findings portray enterprise AI adoption as a broad but uneven organizational phenomenon. Adoption is concentrated among firms with greater scale and intangible investment, while use within adopting firms is distributed across many worker groups and knowledge work tasks but varies substantially in intensity. This heterogeneity highlights that the long-run economic value of enterprise AI adoption will depend on whether dispersed individual use develops into complementary organizational capabilities (Bresnahan and Greenstein 1996; Bresnahan et al. 2002; Brynjolfsson et al.

2021).

The remainder of the paper proceeds as follows: we first review the related literature and describe the data and measurement. We then follow the structure introduced above:

Sections 4.1 and 4.2 examine adoption and usage across firms, while Sections 4.3 and 4.4 examine the distribution and task structure of use within adopting organizations. In Section 5, we conclude.

Related work. Our work contributes to four related literatures. First, we add to the literature on AI adoption and diffusion within firms. A central insight from research on general purpose technologies is that initial adoption does not imply effective deployment: realizing value requires experimentation, complementary investment, and organizational change, so use often diffuses gradually within firms (Bresnahan and Trajtenberg 1995; Bresnahan and Greenstein 1996; Bresnahan et al. 2002; Bresnahan 2024; Brynjolfsson et al. 2021; Mansfield 1963; Fuentelsaz et al. 2003). Consistent with this view, recent studies of digital technology adoption and use find that organizations continue to discover applications and improve their use after obtaining access (McElheran et al. 2024; Yotzov et al. 2026; Bick et al.

2026a; Bick et al. 2026b; Brand et al. 2024; Kim et al. 2026; Massenkoff et al. 2026a).3 The most closely related paper to ours is Bonney et al. (2026), which distinguishes firm adoption from the subsequent deployment of AI across business functions and worker tasks. We advance this literature by examining which firm attributes predict enterprise AI adoption and, among adopters at a common point in their adoption cycles, measuring how deployment is distributed across workers and tasks.

Second, our paper contributes to research that measures AI usage with telemetry data. One strand uses data on the content of human–AI interactions to characterize the tasks, occupations, and modes of interaction represented in observed use (Handa et al.

2025; Appel et al. 2026; Massenkoff et al. 2026b; Chatterji et al. 2025; Tomlinson et al.

2025). A second strand uses API and product-activity data to characterize demand across applications and organizational settings and to examine how AI is incorporated into production workflows (Demirer et al. 2025; Fradkin 2025; Appel et al. 2025; Daniotti et al.

2026; Chen and Stratton 2026; Demirer et al. 2026b). Two recent papers are particularly closely related to ours. Counts et al. (2026) use telemetry from Microsoft 365 Copilot to characterize aggregate workplace use and document how its task composition varies across occupations and industries. Johnston et al. (2026) use OpenAI telemetry data to study the shift from conversational to agentic AI, including how Codex adoption, usage intensity, and task composition vary across organizational settings, worker roles, and levels of seniority.

Third, we contribute to research on heterogeneous effects of AI across workers and tasks. This literature distinguishes between the activities for which AI is technically capable and the settings in which those capabilities translate into realized use and benefits.

Method. Our analysis draws on four related but distinct samples: an aggregate enterprise usage sample, a smaller sample with employee job title and firm industry information, a further time-limited subset of the job title and industry sample used for task-classification analysis, and a public company sample linked to the Compustat database from S&P Global Market Intelligence.5 We describe the construction of each sample in the following subsections and summarize their relationships in Figure 1.

For our analysis of ChatGPT Enterprise usage data, we use de-identified data and report results only in aggregate. Message content is classified using automated systems, and job title metadata is mapped to broad job title class, seniority, and people manager categories. No researcher manually reviewed individual enterprise customer messages for this study. For our financial analysis of public companies, we securely link aggregate organizational usage data to public-company financial information from Compustat.

3.1 ChatGPT Enterprise Usage Data Our primary data source is an organization-week panel of ChatGPT Enterprise adoption and usage, constructed from organizations whose ChatGPT Enterprise adoption dates range from January 1, 2024 to March 31, 2026. The data capture adoption of a paid, centrally administered ChatGPT Enterprise workspace, rather than use through personal accounts, the API, or other subscription plans. We observe each organization’s enterprise account identifier, adoption date, and product usage over time.

Organizations enter the panel in the week they adopt ChatGPT Enterprise and they remain in the panel while their workspace is active. Organization-weeks with an active workspace but no observed product activity are retained with zero measured usage. For each organization-week, we measure messages sent, active users, and generated output tokens, including tokens generated through both ChatGPT and Codex. Weeks are indexed relative to each organization’s adoption date. This aggregate ChatGPT Enterprise usage sample is used to measure adoption and product use over time.

3.2 Job Titles, Firm Industries, and Task Classifications For analyses of usage by worker characteristics, we also construct a sample of ChatGPT Enterprise organizations for which we observe both firm industry and high-quality employee job title information. Starting from the ChatGPT Enterprise usage sample described above, we retain organizations that can be assigned to a broad industry category using NAICS classifications. We further require that at least some user activity within the organization can be linked to a non-empty administrative job title.6 Because the corresponding analyses measure usage 6 months (26 weeks) after adoption, we additionally require an observed, active organization-week at that horizon. The resulting worker characteristics sample contains 1,764 organizations and 17,446,551 messages.

Within this sample, we normalize the available job title strings and classify them into broad job title classes, seniority levels, and people manager categories.7 These user-title and firm-industry measures are then linked to ChatGPT Enterprise usage and aggregated by organization, week, and worker category. Appendix C describes the normalization, classification, and validation of our job title measures. Importantly, job title coverage within included organizations is incomplete. Active users without usable job title information remain in organization-level usage totals and denominators but are classified as missing or unclassified in analyses of heterogeneous use by worker type.

We also separately construct a task classification subsample of this dataset. We classify ChatGPT Enterprise messages into a taxonomy of work tasks using a message-level classifier that is available beginning on October 30, 2025. Appendix D provides information about the task taxonomy produced by this classifier. The task classification sample is restricted to organizations that satisfy the worker characteristics sample requirements above and have task classification data available at the week 26 horizon.

Discussion. A growing literature argues that AI, and especially large language models, have the characteristics of a general purpose technology (Goldfarb et al. 2023; Eloundou et al.

2024). For such technologies to affect production, firms must do more than obtain access: they must discover valuable use cases, encourage use across workers, and integrate the technology into existing workflows. This paper studies that process using administrative data from ChatGPT Enterprise. One central message is that access to the same underlying system does not imply uniform use: firms differ in whether and when they adopt, workers differ in how intensively they use the tool, and task use varies across industries, job title classes, and seniority levels.

This interpretation has several implications. First, the earliest U.S.-based public company adopters are not average firms; they are larger, more intangible-intensive, and more highly valued. The relationship between adoption and firm capabilities also points toward the importance of complements. Firms with greater scale and accumulated intangible investments may be better positioned to identify valuable applications, support workers in using the technology, and integrate it into business processes. Second, diffusion may initially reinforce existing firm heterogeneity. If larger and more intangible-intensive firms adopt earlier and are better positioned to integrate the technology into work, generative AI could widen differences in productivity or value creation across firms even when the underlying models are broadly available. Third, adopting firms differ substantially in the breadth of participation, the intensity of use, and the task mix to which the technology is applied. These margins matter because the economic role of generative AI depends not only on whether a firm has access, but also on where the technology enters the organization of work.

Conclusion. Even with these limitations, the patterns documented in this paper point to a central feature of enterprise AI diffusion: adoption is only the beginning of deployment. The rapid adoption of generative AI by firms should therefore not be equated with immediate productivity transformation. General purpose technologies rarely generate immediate, economy-wide gains; their impact unfolds through a slower process of co-invention in which firms discover use cases, invest in complements, and reorganize production so that a new capability becomes reliable in everyday work (Griliches 1957; Mansfield 1961; Hall and Khan 2003; Jovanovic and Rousseau 2005; Bresnahan and Trajtenberg 1995). We are still in the early stages of that process. Firms are not merely deciding whether to use generative AI; they are learning where it belongs in their organizational workflow. The economic effects of generative AI will depend on how that learning and decision-making process unfolds across firms, workers, and tasks.

Limitations. Importantly, these results should be interpreted in light of the scope of the data. The analyses measure usage only within OpenAI’s ChatGPT Enterprise product, not usage of other AI systems, API-based tools, internally built applications, or personal accounts. The worker-level results are based on observed administrative job titles, which are incomplete and do not provide denominators for the full workforce in each role. The task results are based on classified message content and do not measure downstream work products, productivity effects, or changes in organizational routines. Finally, the public company analyses are limited to the selected subset of U.S.-based enterprise organizations that can be linked to financial data. Future work should connect enterprise AI telemetry to measures of output, organizational change, and longer-run firm performance, and should examine how adoption, usage intensity, worker composition, and task mix evolve as generative AI continues to become more widespread.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI systems reliably guide voters without introducing political bias? How does AI adoption reshape collaboration patterns in knowledge work? How do AI-exposed occupations change in employment, wages, and skills? How should humans and AI agents share control and decision-making? Does AI deployment reduce or exacerbate workplace inequality and income instability?