Putting AI on the Org Chart: Evidence on Delegation and Accountability
Source: Boston University / BCG (Wiles, Hsu, Bedard, Kropp) · 2026-09-19
Abstract Firms are increasingly deploying agentic AI systems capable of independent action. Some have begun integrating these agents into their organizational structures, assigning them designated roles and responsibilities and referring to them as employees. This creates new challenges for managers’ decisions about when to delegate and how to monitor work. In a survey of 1,261 managers, 23% report that their organizations list AI agents on organizational or workflow charts. In a randomized experiment, we provide those managers with identical documents containing built-in errors and ask them to review them. We vary whether the drafts are presented as produced by an AI tool, an AI employee, or a human employee. Average effects on error catching are small. However, among managers whose organizations already list AI agents on their org charts, we find evidence consistent with a “hot potato” effect: managers catch fewer errors themselves, request more review by others, and assign less accountability to themselves. Presenting drafts as produced by an AI employee rather than an AI tool reduces the share of errors detected by 17%, increases requests for additional review by 22 percentage points, and shifts perceived accountability toward the AI system. Compared with managers reviewing an AI employee’s work, those reviewing a human employee’s work flag more errors themselves and are less likely to request additional review, suggesting that this pattern is not simply a response to delegation. These results suggest that how organizations position AI should be understood as a governance decision that impacts how managers oversee work.
Introduction. “We call it Scout.[...] It’s treated as a team member, it’s technically an equivalent peer on your team [...] It is defined in terms of a job description, has a clear role, has KPIs it needs to hit—just like anyone else.” (Interview with Senior Executive at a Logistics Company)1 A growing body of evidence documents meaningful individual productivity gains from generative AI in knowledge work (Brynjolfsson et al., 2025; Dell’Acqua et al., 2026; Wiles et al., 2026). But these tools don’t simply make workers more productive; they also change the content of work. When working with AI tools, humans spend more time reviewing work and less time writing first drafts of text or code (Peng et al., 2023; Banh et al., 2025). In many firms, AI is moving from a standalone writing aid toward agentic AI systems that draft, recommend, and increasingly execute work inside organizational workflows.2 Recent surveys suggest that a large share of organizations are adopting agentic AI in one way or another, and how these systems are built into the organizations’ work processes can impact whether they realize productivity gains (Kim et al., 2026). Some organizations explicitly describe AI agents as organizational members and list them on organizational charts, or create new “work charts” that include both human and AI actors (Mok, 2025).3 Discussing IBM’s “digital workers,” one IBM vice president remarked, “I don’t think human managers are going to manage these things in the same way as we manage people” (Varanasi, 2026). This raises the question: when AI systems are treated as organizational actors rather than tools, do managers oversee them like software, delegate to them like human subordinates, or treat them as something distinct? To tackle this question, we combine descriptive survey evidence on organizations’ AI adoption with a randomized experiment with managers, directors, and executives (hereafter, “managers”). The survey measures how organizations position AI in work processes and communication, including whether they frame AI as a productivity tool, a career accelerator, a teammate or employee, or a threat, and whether they have institutionalized AI agents by listing them on organizational or workflow charts (hereafter, “org charts”). In the experiment, we isolate the effect of positioning AI as an organizational actor on managers’ oversight and governance choices. We provide managers identical documents to review, and we only vary how we describe the source of the documents—as coming from either an AI tool, an AI employee, or a human employee. This design allows us to compare reviewing work produced using an AI tool with reviewing work attributed to an AI employee, and to benchmark both against delegation to a human subordinate. In this paper, we define AI employee as an AI agent that an organization has assigned a standardized and institutionalized role—for example, granting it data access, giving it a defined set of tasks, or granting it some authority to act. In practice, it is often the same underlying technology as agentic AI operated by a person; it is distinguished by the formal institutionalization of its decision rights and how much authority it has to act. Consider a mid-sized firm that creates an AI employee called Otto Cash to handle accounts payable in its finance department. It is listed on the finance team’s org chart as responsible for invoice intake and routing, has a job title and job description, and gets monthly performance reviews. When invoices arrive, Otto Cash extracts key details, matches them to purchase orders, and drafts a brief memo flagging any discrepancies. For routine cases, it prepares an approval packet that includes a summary, supporting links, and a recommended action (approve, reject, or hold), and it can take actions such as routing the packet to the appropriate budget owner, requesting missing documentation from the submitter, or scheduling the invoice for the next payment run without any human involvement. However, it cannot release payments without sign-off from a human manager. We document that this practice is already widespread: in a survey of 1,261 human resource (HR) and finance managers recruited through a professional expert network, 31% say that their organization frames AI as a “teammate or employee” and 23% report that their organization even lists AI agents on its org chart. A second survey of a broader sample of senior managers weighted to reflect the U.S. private sector suggests that these practices are not confined to our experimental sample: 14% report that their organizations list AI agents on org or workflow charts, and 33% report integrating AI agents into their systems and give them some form of organizational recognition (such as give the agent a name or assign the agent a manager.) In order to isolate the effect of this AI employee framing on managerial oversight and governance practices, we conduct a randomized experiment on these HR and Finance managers.
Method. This paper uses data from two sources. First, we run an experiment on a sample of managers, directors, and executives (hereafter, “managers”) recruited through a B2B expert network, which we will describe below. This is the primary sample and dataset used in this paper. Second, we run a second survey of 1,500 senior managers, directors, and executives (hereafter, “senior managers”) recruited through YouGov, weighted by industry to better reflect the distribution of senior managers across U.S. private-sector industries. We use this survey only as complementary descriptive evidence to assess whether the organizational practices documented in our experimental sample appear in a broader managerial population, and to provide additional detail on how firms are formalizing AI agents in organizational workflows.
The experiment was carried out in two sessions, a registration survey and an experiment, separated by approximately one week. We recruited participants through a B2B expert-network research Following the initial registration survey, on January 9, 2026, participants were invited to the second session to complete the main experimental task. They had one week to complete this task. The main task involved reviewing five documents with errors: job descriptions for managers in HR and budget documents for managers in Finance. Regardless of domain, all participants reviewed a sequence of five documents using an interface that allowed for highlighting, flagging, and commenting. Participants were given a total of 20 minutes to review as many documents as possible.
AI tool framing Your company is finalizing this year’s budget reports across multiple business units. To draft the budget documents, you used an AI tool. You recently started using this generative AI tool to help with finance documentation tasks. This AI tool uses natural language processing to generate budget materials by analyzing similar reports from prior periods and company planning guidelines.
AI employee framing Your company is finalizing this year’s budget reports across multiple business units. ALEX- 3, your AI employee, has drafted the initial budget documents. ALEX-3 was assigned to your team 6 months ago as a direct report and appears on your department’s organizational chart. ALEX-3 is a Generative Artificial Intelligence system that generates budget documents using natural language processing. Like your other team members, ALEX-3 handles finance documentation tasks based on prior period reports and company planning guidelines.
Human employee framing Your company is finalizing this year’s budget reports across multiple business units. Alex, your employee, has drafted the initial budget documents. Alex was assigned to your team 6 months ago as a direct report and appears on your department’s organizational chart. Alex is a recent hire who came from a similar role at another company. Like your other team members, Alex handles finance documentation tasks based on prior period reports and company planning guidelines.
Randomization: Randomization to the framing conditions was conducted at the individual level. To improve statistical precision, we stratified the random assignment by domain (HR vs. Finance), review frequency (whether the participant reviews others’ work at least several times per week vs. less often), and AI usage at work (whether the participant uses GenAI tools daily vs. less often). Within each domain, the order of the five documents was randomized using a Latin square design to ensure each document appeared equally often in each position.
Analysis sample: Of the 1,261 participants who took the registration survey, 857 completed the experiment in the second session, with a 68% response rate. To form the final analysis sample, we excluded 44 participants who failed a simple attention check, resulting in a final analysis sample of 813 respondents. We describe this sample in Table 3 Panel C.
Conclusion. Organizations are increasingly deploying agentic AI inside core workflows, and some are going further by describing these systems as coworkers or “AI employees” and encoding them on org charts or rebranding them as work charts which include human and AI actors. We document that this practice is already meaningful: in our survey of HR and Finance managers, directors, and executives, nearly one quarter report that their organization lists AI agents on their org chart, and a second survey of a broader sample finds similar formalization practices outside the experimental sample. We run a randomized experiment where we provide managers with documents to review while varying only whether the upstream drafter is framed as an AI tool, an AI employee, or a human employee, allowing us to isolate the causal effect of role framing on oversight behavior and beliefs. In the full sample, AI employee framing has relatively small effects. But among managers whose organizations already formalize AI agents, it changes the organization of oversight in a striking way. The largest behavioral response is a sharp increase in managers’ willingness to ask someone else to review the work: costly requests for additional review rise by 44% relative to AI-tool framing, while managers in the AI employee condition do less direct oversight and shift perceived accountability away from themselves and toward the AI system. The human employee treatment arm shows that this is not simply a generic effect of delegation: the most direct oversight came from the group of managers for whom the work was described as coming from a human employee. If the AI employee framing simply caused managers to oversee work how they would whenever they delegate, then we would expect similar responses when the same work was described as coming from a human employee. These results have straightforward implications for organizations experimenting with AI employees.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can humans maintain effective oversight as AI systems scale?- Do workers lose oversight skills by relying on AI to delegate?
- Can labeling alone erode oversight skills without changes to AI capability?
- Why does employer policy reshape who actually makes final decisions?
- How does formal organizational recognition of AI systems change manager accountability?
- What role does organizational policy play in shaping how managers use agentic AI?
- How does delegated workflow adoption differ from conversational chatbot usage patterns?
- Does delegating to an AI employee differ from delegating to a human subordinate?
- What levels of human-AI collaboration do workers prefer across different occupation types?
- Which workplace tasks remain hardest for AI agents to complete autonomously?
- Why do firms substitute labor for AI faster than gig worker jobs disappear?
- Can workers move across the divide between technical and non-technical job markets?
- How does AI task concentration within firms affect worker reallocation across jobs?
- Why does AI adoption favor automation over augmentation in female-dominated work?
- Do firms with high AI exposure shed jobs or reshape roles?
- Which occupations show the sharpest gap between AI capability and actual adoption?
- How does concentrated AI exposure across workers affect firm-level employment demand?
- Can workers reallocate across occupations fast enough to offset AI displacement?
- Can workers retrain faster than AI exposure spreads through occupations?
- How does occupational segregation affect who gains from AI productivity?
- How do institutions shape whether AI enables worker mobility or deepens hierarchy?
- Does AI adoption rise or fall as worker education and wages increase?
- Can persistent agentic workflows predict labor displacement better than task-level exposure?