Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Paper · arXiv 2607.01418 · Published July 1, 2026
Domain Specialization in LLMs

Organizations rolling out agentic command line tools like Anthropic’s Claude Code and GitHub’s Copilot CLI need to know who will try them, who will keep using them, and whether the tools produce enough output to justify their cost. At organizational scale, token spend can run into millions of dollars annually, so misreading adoption, retention, or impact can make a rollout expensive without changing engineering velocity. Studying tens of thousands of engineers at Microsoft over its early-2026 rollout, we find that first use spread primarily through social networks, retention was associated more with engineers’ coding activity than with demographics, and adopters merged roughly 24% more pull requests than they would have otherwise. We use merged pull requests as our proxy for output — acknowledging that a merged PR is not the same as the value it delivers — and the lift persists across our four-month window. These results suggest that CLI coding agents are neither uniformly adopted nor mere novelty effects and that organizations should treat visible peer use as central to rollout strategy.

Introduction. Agentic command line tools like Anthropic’s Claude Code, Google’s Gemini CLI, and GitHub’s Copilot CLI are increasing in popularity among software developers. Such tools harness agents that call large language models, where the agents execute semi-autonomous commands on behalf of the user from the command line. In early 2026, the Pragmatic Engineer’s survey indicated that Claude Code was the most popular AI-based developer tool among respondents [26]. At the same time, organizations considering whether to purchase such tools do not yet know which of their engineers are likely to adopt and use them. At the end of 2025, StackOverflow’s survey [35] of more than 49,000 respondents indicated that developers’ trust in AI is falling and that “developers remain willing but reluctant to use AI”. Even if organizations choose to use AI, the tokens needed to execute these tools can be expensive. At the extreme high end, Fortune reports on Meta employees’ usage of AI [14]:

In a 30-day period, total employee usage on the dashboard exceeded 60 trillion tokens, and the highest-ranked individual user averaged 281 billion tokens. Using the least expensive version of Claude Opus 4.6, which costs $5 for every million tokens, that one user alone could have cost Meta more than $1.4 million.

Even at more modest levels of usage, organizations may wonder what return on investment they should expect. Three concerns thus follow any rollout: which engineers will adopt, whether they will keep using the tool, and whether the tool produces enough additional output — which we operationalize as merged pull requests — to justify its cost. To examine them, we analyze our experience in early 2026 with Microsoft offering its engineers two agentic command line tools — Claude Code and Copilot CLI — over a roughly four-month window of usage and pull request (PR) activity. This paper contributes the first field study to use developer-level telemetry to analyze both the adoption of agentic command line tools and their effect on pull-request output. Prior developer AI adoption work typically relies on surveys and interviews, and prior impact work largely infers AI use from public-repository signals (Section 2); our enterprise setting instead observes every engineer who could adopt alongside direct usage. Within this setting we separate initial use from retention, trace merged-PR output against how intensively the tools are used, compare two different tools, and draw on an internal developer survey to help interpret the results. We analyze adoption and impact in two studies. The adoption study (Section 4) asks who adopts: among engineers who could use Copilot CLI, who tries it (RQ1) and, among those who try it, who keeps using it (RQ2). The outcomes study (Section 5) asks what adoption produces: whether using either tool yields more merged PRs (RQ3), whether the specific tool matters (RQ4), and which engineers benefit most (RQ5). The adoption study covers only Copilot CLI, the tool with a well-defined eligible-adopter population at rollout; the outcomes study covers both Claude Code and Copilot CLI.

Related work. The study of technology adoption has a long history in the social sciences, from Rogers’ diffusion of innovations theory [32] to more technologically focused work on information systems [12, 39, 40], which tells us that adoption is shaped by individual, social, and organizational factors. In this paper, we use adoption as an umbrella term for a developer’s uptake of a tool, and decompose it into two phases that this literature treats as distinct [5, 32]: initial use — a developer’s first use — and retention — whether that use is sustained. We study the two separately because, as our results show, the factors that predict initial use are not those that predict retention. In the developer AI space, Reyes-Reina and colleagues’ recent systematic review of AI-tool adoption in software development characterized 25 studies [30]. Yet every empirical study in that review rests on surveys, interviews, or focus groups — or, where behavior is directly observed, on one-shot lab tasks; none draw on observational data of developers’ actual adoption. We fill exactly this gap, enabled by access to (1) longitudinal, individually-identifiable telemetry of AI tool usage and (2) human resources (HR) data about developers. Beyond that review, a few recent studies examine developer AI adoption with observational data. Daniotti and colleagues infer AI-assisted coding from a commit classifier to document AI’s global diffusion [11]; Yang identifies Claude Code adopters among 16,000 scientists via the tool’s default commit co-author trailer [43]; and Robbes and colleagues detect AI agent adoption across GitHub projects from commit trailers, config files, and branch names [31]. But all observe only developers who leave a public-repository signal, so a “non-adopter” is merely someone with no such signal — conflating true non-use with suppressed or invisible use. Lacking a roster of eligible engineers, they cannot cleanly separate adopters from non-adopters. Our enterprise setting supplies that missing denominator: the full population of engineers who could adopt Microsoft’s sanctioned agentic command line tools, paired with logs of who tries AI and who keeps using it.

How AI assistance affects developer productivity has been examined along three axes: what developers report, how they perform on controlled tasks, and what they ship in real-world work. The most common evidence of developer productivity is self-reported. A recent systematic review finds that surveys of developer satisfaction are the dominant way studies measure productivity [21]. Such surveys generally find that developers who accept more AI suggestions report feeling more productive and satisfied [1, 25, 45], and they point to writing and implementing code as where AI helps most [15].

Method. The adoption study examines Copilot CLI uptake in two phases — initial use and retention — among the engineers eligible to adopt it at rollout, the sample we define in Section 4.2. We also draw on a developer survey (Section 4.6) to help interpret the adoption patterns we find. The quantitative pipelines and figures in both studies were implemented in code written with an AI assistant under the authors’ direction and verified by the authors; we give our full AI-use disclosure in the Acknowledgments.

The adoption study answers two questions that, although they sit naturally together, require different data structures and different statistical models. We state each question with the outcome variable that operationalizes it. RQ1. Among engineers who could have used Copilot CLI but had not yet, who tries it first? We define initial use as an engineer’s first opening Copilot CLI during the post-rollout window. For each engineer-week (i, w) in which engineer i had no prior Copilot CLI activity, the initial-use outcome Ai,w is 1 if i recorded any Copilot CLI activity in week w and 0 otherwise; once Ai,w = 1, engineer i contributes no further rows. Because the population still at risk of first use shrinks from week to week, this gives a discrete-time hazard on an engineer-week panel. RQ2. Among engineers who adopted Copilot CLI, who keeps using it? We define retention as sustained early use: an adopter is retained if they recorded Copilot CLI activity on at least 5 of the 14 days beginning with their first use. The 5-of-14 threshold approximates “used Copilot CLI on roughly half of working days during the first two weeks,” designed to separate “tried and stayed” from “tried and abandoned.” We exclude adopters whose first use was within 14 days of the end of our tool-use observation window, since their retention window has not yet fully elapsed. Because retention is defined only for engineers who have already adopted, it is a cross-sectional outcome with one row per adopter.

The adoption study’s sample consists of Microsoft software engineers, every one of whom could have tried Copilot CLI on the rollout date. Engineers who already had Claude Code access, by contrast, faced a different adoption decision, with a sanctioned agentic command line tool alternative already on hand. To isolate the population whose adoption decision is uniform, we start from all Microsoft employees with software engineer job titles and drop three groups:

• members of two large divisions that received broad Claude Code access through a managed program; • individual Claude Code licensees outside those divisions; • retracted users — engineers who once held a Claude Code license but no longer do.

Since Copilot CLI was generally available from GitHub starting February 25, 2026 but was available to Microsoft through a product preview program somewhat before that, we chose January 5, 2026, an earlier date consistent with our outcomes study, as the cutoff. We thus excluded engineers who first used Copilot CLI prior to this cutoff. The pre-period for predictor construction runs October 1, 2025 through January 4, 2026 (13 weeks). The post-period over which adoption and retention are observed runs January 5 through April 29, 2026.

We collected five predictor groups for each engineer. For each group we name the construct, why it might shape initial use or retention, the operationalization, and the source.

Career stage The construct is the scope, qualifications, and general expectations of the engineer’s role. Microsoft uses parallel career-stage ladders — IC for individual contributors and M for people-managers — where higher numbers indicate greater scope and seniority within each ladder, and stages at the same number on the two ladders (e.g., IC4 and M4) are calibrated as peers. Operationalization: career stage from HR records, coded as IC2–IC6 (individual contributors) and M4–M6 (managers), with IC4 as the reference category.

Tenure The construct is years at Microsoft. Operationalization: years since first hire date, bucketed into <1y, 1–2y, 2–5y, 5–15y (the largest bucket, used as reference), and 15+y.

Discussion. Adoption is substantially social. The strongest predictor of who tries Copilot CLI in any given week is whether the engineer’s peers — especially the broader skip-level group — have already tried it. These social signals could reflect peers observing each other’s successful work strategies [42], or, in the case of managers’ use, implicit or explicit directives. To the extent that our results reflect prior work — that learning developer tools from peers is highly effective [23, 24] — this suggests that organizations should enable engineers’ use of agentic command line tools to be visible and socially reinforced. An engineer’s prior IDE Copilot use also predicts trying Copilot CLI; we interpret this as an indicator of familiarity and openness to AI tooling. Retention is substantially behavioral. The predictors associated with adoption are not equally associated with retention, and one of them — prior IDE Copilot use — actively predicts against sticking with Copilot CLI. A plausible interpretation is that engineers who already trust AI tooling in their IDE will try Copilot CLI but have a familiar alternative to fall back on, so they may not build a sustained Copilot CLI habit; engineers for whom Copilot CLI is their first such tool have no such fallback and, if they stay, tend to stay more firmly. Engineer attributes do not matter much to adoption and retention. Career stage produces a gentle gradient among individual contributors and nothing detectable among managers. Tenure produces a small newest-hire bump and otherwise nothing. The five groups together suggest that what an engineer does (peer ties, prior tool use, PR cadence) explains who adopts and retains far better than who an engineer is in the organization. Pull request lifts persist. Our finding that the PR lift does not fade stands in contrast to He and colleagues [16], whose Cursor lift faded at month two and was gone by month three — well inside our roughly four-month window, so a too-short observation period is an unlikely explanation. We read the contrast as genuine, and offer two non-mutually-exclusive reasons: (a) tool generation — they study Cursor, a 2024-2025 IDE-based tool, while we study 2026 agentic command line tools; and (b) unit of analysis — their repo-level average reverts toward the mean as adoption spreads from keen early adopters to marginal users, whereas our within-person design conditions on the same engineer and is immune to this compositional drift. Why does Copilot CLI outpace Claude Code? The larger PR lift for Copilot CLI than Claude Code adopters (Section 5.2.2) is surprising in light of developer sentiment around agentic command line tool. As described in the introduction, public early-2026 comparisons of agentic coding tools generally rate Claude Code as the preferred option for autonomous agentic work; our data shows the opposite ordering on merged-PR throughput. We have two hypotheses: (a) the two tools are used for different task mixes — engineers reach for the two tools for different purposes; and (b) Copilot CLI works better for Microsoft employees — because Microsoft owns GitHub, the maker of Copilot CLI, but is a buyer of Claude Code, organizational forces likely helped align the Copilot CLI harness with the way engineers at Microsoft work.

Conclusion. In this paper, we analyzed how emerging CLI-based agentic coding tools are adopted and used in a large software organization. We found that initial use spread substantially through social channels — an engineer’s peers and managers using the tool — and that adopters went on to merge roughly 24% more pull requests, a gain that held steady across our four-month window. For organizations weighing the token costs raised at the outset, this sustained lift in merged pull requests is direct evidence that agentic command line tools can move a concrete output metric — though whether that output justifies its cost is a question of value, not throughput alone. The pressing open question is now about quality, whether this added throughput yields better software. The field still lacks agreed-upon measures to answer it, and building them should be a priority of the research community. While AI developer tools are advancing at a rapid pace — from chat to code completion to agents on the desktop and in the cloud — understanding adoption and outcomes is critical for today’s software organizations to align expectations with reality.

Limitations. Construct validity. In the adoption study, two thresholds are somewhat arbitrary: retention (5 of 14 days from first use) and the 28-day merge window. We re-ran retention at both a looser 3-of-14 and a stricter 7-of-14 threshold, and the retention model fits at all three. The 28-day merge window trades longer-merging PRs for comparability and freedom from right-censoring; much longer windows would be needed to gauge multi-year adoption. Internal validity. The adoption study is cross-sectional, so it cannot rule out several confounds. For the social effect, we cannot separate peer influence from homophily: an engineer may adopt because a peer did, or simply because similar engineers tend to cluster together [3]. Our fixed effects also miss engineer-specific, time-varying shocks (a personal project, a vacation) that can move initial use within a single week. The outcomes study’s within-person design addresses this for outcomes, but no analogue exists for the initial-use choice. We excluded retracted Claude Code licensees and two broadly-licensed divisions from the adoption sample; re-running the fits on two sample-frame variants left direction and rank-order unchanged — top-bucket reviewer-peer exposure, for example, stayed within a few points of its +54% headline odds lift.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do AI coding tools measurably improve developer productivity and code quality? How does AI adoption reshape collaboration patterns in knowledge work? Can AI research automation sustain progress through accelerating feedback loops? Does AI assistance erode cognitive skills while inflating perceived competence? Does AI deployment reduce or exacerbate workplace inequality and income instability? How do real-world evaluations reveal AI capabilities that benchmarks hide?