How much does AI impact development speed? An enterprise-based randomized controlled trial

Paper · arXiv 2410.12944 · Published October 16, 2024
Domain Specialization in LLMs

Abstract—How much does AI assistance impact developer productivity? To date, the software engineering literature has provided a range of answers, targeting a diversity of outcomes: from perceived productivity to speed on task and developer throughput. Our randomized controlled trial with 96 full-time Google software engineers contributes to this literature by sharing an estimate of the impact of three AI features on the time developers spent on a complex, enterprise-grade task. We found that AI significantly shortened the time developers spent on task. Our best estimate of the size of this effect, controlling for factors known to influence developer time on task, stands at about 21%, although our confidence interval is large. We also found an interesting effect whereby developers who spend more hours on code-related activities per day were faster with AI. Product and future research considerations are discussed. In particular, we invite further research that explores the impact of AI at the ecosystem level and across multiple suites of AI-enhanced tools, since we cannot assume that the effect size obtained in our lab study will necessarily apply more broadly, or that the effect of AI found using internal Google tooling in the summer of 2024 will translate across tools and over time.

Introduction. Seven years after the rise of LLM architecture [1] and two years after the start of the “chatbot revolution” [2], significant investments have been made to AI-enhanced products, including in the software developer space. Since the release of GitHub Copilot [3], numerous developer tools offering code editing and generation support have been built for the general developer community [4], [5], [6], [7]. Researchers and educators have also developed prototype tools to assist novice programmers and students. Furthermore, substantial effort has been invested in building tools for internal use, such as at Meta [8] and Google [9], [10], [11]. However, there is still much to investigate to answer how useful these tools are in helping developers, specifically in improving their productivity. Truly understanding the productivity benefits of AI enhanced coding tools remains a nascent field. While some research has shown improvements in coding speed [12], developer throughput [13], and perceived productivity [14], more work must be done to validate these assertions across, for example, tasks, developer contexts, user groups, and more. To date, very few estimates of the impact of AI-enhanced developer tools on time spent on task in an enterprise context have been published. One much-discussed study was a randomized controlled trial by Peng et al. [12] (n = 95), which found a 56% speed increase for developers using GitHub Copilot—an AI code assistant—, compared to those not using it. Another enterprise-specific estimate comes from a pooled analysis of three field experiments (n = 4, 867) conducted by Cui et al. [13], where developers either had access to Copilot or did not have access to it in their daily activities. The authors found a 26% increase in throughput (measured as an increase in the number of pull requests) for developers using Copilot. Although these studies provide valuable insights and help quantify the speed improvements offered by one AI-enhanced developer tool, gaps remain in estimating the overall impact of different coding assistants and across industry settings. Other shortcomings include missing nuance or understanding of the impact of AI tools by developer- or task-level contexts (e.g., [13]), limited tool and task complexity in experimental settings (e.g., using GitHub Classroom rather than a more naturalistic developer environment [12]), and small sample sizes (e.g., only 32 participants in [15], 24 participants in [16], or 21 participants in [17]). Altogether, the rapidly-changing status of AI-enhanced developer tools, combined with the partial portraits provided by current studies at this very early stage in the empirical study of AI tools in production, requires continued inquiry. In this paper, we complement this recent literature by providing an estimate for the impact of AI in an enterprise context (as per [13]) and applying it to speed on task (as per [12]). To simulate the enterprise context in the study, we designed a task covering multiple aspects of software development, from writing and editing code, to updating build files and to testing, within our proprietary internal infrastructure at Google. We aimed to answer the following questions specifically for our internal developer tools at Google:

• RQ1: What impact does AI have on time spent completing an enterprise-grade development task?

• RQ2: How do developer and task characteristics influence our estimates of the impact of AI assistance on time spent on task?

• RQ3: How do developer and task characteristics interact with the use of AI to accelerate or slow down certain developers and not others?

Providing a robust estimate for the impact of AI-enhanced tools on development speed is critical to the long term adoption and success of these tools across the industry. Continued investment in, and adoption of, these tools is dependent not only on how developers feel about them; it is critical to be able to evaluate their business impact in terms of greater output or time gains for the organization. The second and third research questions unlock important new understandings around how to design and develop AI-enhanced developer tools in a product and user-centric manner. When we understand our users, we can better cater to their needs. In this study, we ran a controlled trial with 96 Google software engineers who were randomly assigned either to use (experimental condition) or not use (control group) three AI-enhanced features for code (AI Code Completion, Smart Paste, and Natural Language to Code; see Section III for more details) to complete an enterprise-grade task. We analyzed data using t-tests and linear regressions on time on task data to evaluate the impact of AI on speed on task. We ascertained the robustness of our estimate using multivariate regressions based on a theoretical framework [18]. Then, we answered our research questions by testing hypotheses we built based on the theoretical framework and the literature.

Related work. In recent years, driven by significant improvements in large language models (LLMs), many developer tools have been built upon or incorporated LLMs. GitHub Copilot [3] is one of the earliest LLM-powered developer tools, suggesting code in real time, based on context. Other examples include Alpha- Code [4] from DeepMind, CodeWhisperer [5] from Amazon, Tabnine [6], and Cursor [7]. These tools primarily offer code completion, editing, and generation, often with additional chat functionality or context integration to improve code quality and the user experience. Significant effort has also been invested in developing AI-enhanced software development tools for internal use, such as at Meta [8] and Google [9], [10], [11], tailored to their developers’ workflows and proprietary codebases. The increasing prevalence of AI-enhanced developer tools has spurred significant research into their benefits and drawbacks across various contexts, including computer science education, open source projects, and industry settings. Recent work has explored the potential of these tools to assist students and novice programmers [19], [20], [21], as well as programmers’ perceptions of [22], [23], [24] and trust in [25], [26], [27] AI tools. A key question surrounding these tools in industry is their potential to enhance developer productivity, leading to improved software quality and reduced development effort and cost. While some studies have investigated the perceived usefulness of AI-enhanced developer tools (e.g., [14], [28], [29]), quantifying actual productivity or speed gains has proven more challenging, due to difficulties in accessing and analyzing unbiased, real-world usage data. Moreover, understanding developer productivity itself requires considering multiple factors beyond simple metrics such as lines of code, and existing research attempting to quantify developer productivity—even outside the context of AI tools—highlights the complexities involved. Despite these challenges, some studies have attempted to measure the impact of AI tools on actual developer productivity. For instance, Ziegler et al. [14] analyzed telemetry data to investigate the productivity benefit of GitHub Copilot, and reported over one-fifth of suggestions were accepted by developers, which correlated with their perceived productivity. Similarly, two studies have documented increased throughput for AI in a controlled environment on information-gathering tasks [15] and in a large, pooled analysis of multiple field experiments, and therefore in enterprise contexts [13].

Method. Software development at Google happens in a monolithic source-code repository (or “monorepo”) called Piper [10]. Google developers have access to multiple integrated development environments (IDEs), including Cider V, a variant of the VS Code IDE by Microsoft [34]. To commit new code, developers patch changelists (or “CLs” for short), the equivalent of “pull requests” in the Git ecosystem, and of “diffs” at Meta. CLs can be “patched” to edit the code. Once the CL is patched, developers can read and edit code, and changes are tracked and documented as the CL is submitted for review (see [10] for more details). The features we included in this study were all in production and thus available to all Googlers in Cider V at study time. These features were as follows:

• AI Code Completion. This feature is a novel, transformer-based hybrid semantic AI code completion that enables single- and multi-line code suggestions as developers type, highlighting the recommendations in light-grey text. Our version [35] resembles that built elsewhere in the industry [36]. See Figure 1 for a visual representation of the feature.

• Smart Paste. This feature uses AI to enable contextaware adjustment to code that is pasted from one area to another in the IDE [37]. It works in ways that are similar to the well-known copy/paste shortcuts, and displays only suggestions that are high-confidence using light-grey font for the recommendation, and crossing out the text that will be removed. The suggestion is accepted using the “tab” key, and ignored otherwise. See Figure 2 for a visual overview of the feature.

• Natural Language to Code. This feature leverages an AI-assistant trained in Python, Java, Go, C++, TypeScript and JavaScript. To activate, developers move their cursor to (or select) the code area that they want to change. A natural-language to code prompting window enables connection to the model, which makes a suggestion for code fixes (see Figure 3 and [11] for more detail). They can then review the suggestion and either choose to apply or reject it.

To answer our research questions, we designed a randomized controlled trial (RCT). RCTs are a type of scientific experiment designed to assess the effectiveness of a treatment condition by exposing some participants to the treatment (the “experimental condition”) and some not (the “control group”). RCTs are often considered an empirical standard for establishing causal links between an intervention and its associated effect, observed empirically, and provide unbiased estimates [38]. Our randomized controlled trial was executed as follows (see Figure 4). In June and July 2024, full-time software engineers from across Google were recruited via email by a team that was independent from the research team, and were accepted into the study if they met the following basic criteria: they had been working at Google for at least one year, they were proficient in C++, submitted code to Piper, used Since RCTs do not eliminate the need to include control variables in analyses [38], we created a theoretical framework [18] that connects our dependent variable (time spent on task) with the intervention as well as developer- and tasklevel factors. We created a theoretical framework describing the predicted relationships among multiple components of the developer experience and their impact on time on task (see Figure 5). Theoretical frameworks are a “logically developed and connected set of concepts and premises ... that a researcher creates to scaffold a study” [18]. Our framework was influenced by a recent systematic review of the different factors that influence AI adoption [40] and the literature on developer productivity [32]. Using these papers, our internal metrics framework, and a review of the literature on attitudes towards AI (which included [41], [42], [43], [44], [45]), we generated a list of concepts that we then either instrumentalized through survey questions or telemetry measures. As noted above, survey questions were then tested through two rounds of “think aloud” cognitive testing [39]. 1) Dependent variable: Time spent on task: Our dependent variable was defined as time spent on task.

Discussion. In this study, we aimed to quantify the impact of using AI coding features on the time developers take to complete an enterprise-grade, standardized task. Our analyses found that developers who used AI features were statistically significantly faster than those who did not. This suggests that the AI-enhanced features included in the study (see Section III) do indeed make developers faster, and thus supports our first hypothesis (t(83.6) = 2.11, p = .038). When controlling for other factors in our theoretical framework, we obtained a similar effect size for AI, but the effect did not meet the p < .05 significance threshold. At this stage, we estimate a roughly 21% increase in development speed attributable to AI, controlling for other important predictors (i.e. the effect attributable to AI in our best-fit model, Model 2). This estimate is significantly smaller than the 56% estimate shared by Peng and colleagues about GitHub Copilot [12], but aligns with the 26% productivity increase attributed to Copilot by Cui and colleagues [13]. The difference between our estimate of developer speed gained through AI and that published in [12] is likely attributable to two main factors. There may be important differences between our suite of AI-enhanced tools and those used by GitHub Copilot. More likely, however, the difference is caused by differences in the underlying populations we recruited from. Indeed, Peng et al. [12] recruited from Upwork (a freelancing platform), while we only sampled full-time Googlers. Their sample is therefore likely much more diverse than ours when it comes to coding experience and expertise. It is no surprise, then, that our estimate is closer to that offered by [13], which focused on enterprise users. When it comes to the impact of the time participants spend on coding tasks daily on their speed on our task, our data suggest that developers who code more hours per day may be faster with AI tools than those who code less. While our interaction effect was not significant (β = −0.29), the effect size was large, negative, and the model itself was significant (p = 0.018). This might be because our current AI tools require that developers spend a lot of time verifying and editing code generated by AI, which may confer an advantage to developers who spend more time working with code. Importantly for the future of AI-enhanced developer tools as products serving very diverse subsets of developers, our data suggest that more senior developers, as defined by their level, may work even faster with our AI tools. This is in contrast with some findings in the literature, which suggest that more junior developers stand to benefit more from AI [48], but aligns with other work by Nam and colleagues, who found that coding experience might amplify the value gained from AI [15]. We believe that this may be because AI is not yet able to close a skill gap when applied to complex tasks such as the one on which we tested our participants. To bridge skill gaps, more research and development work will be needed—along the lines of personalization, perhaps—to make more junior developers even faster with AI. While the exact mechanisms connecting hours spent coding and seniority with increased task velocity with AI tools are still unclear, approaches like the one proposed by Mozannar and colleagues [46] are likely to be a viable path forward in identifying what is truly going on at the intersection of expertise and speed with AI. Indeed, the code generated by our AI features in such a context is still complex; users might therefore still need to understand highly-nuanced, largescale systems to be effective while using these features.

Conclusion. While our team is bullish on the speed gains that will be realized by AI-enhanced coding tools, questions of equity are not settled by our study, and questions about the impact of AI on code quality were not explored. Much more research and development is required to explore and potentially remedy the differentiated impact that AI-enhanced tools might have on people of diverse seniority levels. While our study hinted to the positive impact of AI tools on more senior developers, other research has found contradictory effects [15], [48]. The future might see the targeting of different tools to different types of developers, or learn directly from developers what is optimal for their own specific needs and developmental stages, and provide personalized experiences. Importantly, there is a growing literature suggesting that while coding assistants might be increasing the total number of code contributions [13], [49], AI might be lowering code quality at the ecosystem level [50], might increase code churn [17], and might not have reached a quality bar—either on the model or UI side—that improves task completion rates [16], [51]. Careful research and development work that balances the sometimes-contradictory incentives to move fast and guarantee high-quality code will ensure the long-term success of AI in the developer space, especially if such work focuses on where developers might optimally trust AI and therefore want to delegate more work [22], [52].

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do AI coding tools measurably improve developer productivity and code quality? Does AI-assisted work increase total productivity or just shift time? Does AI assistance help or harm professional skill development? What are the real-world consequences of AI citation hallucinations? Can AI research automation sustain progress through accelerating feedback loops?