How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations

Paper · Source
AI at Work

Source: Strategy Science (Allen, McDonald) · 2026-03

Abstract. Benchmarks have helped fuel rapid progress in large language models (LLMs) across a variety of domains including math, science, dialogue, and coding. Yet no existing benchmark adequately captures the defining elements of strategic decision making: uncer­ tainty, complexity, irreversible multiperiod moves, and delayed or noisy feedback. This gap limits our ability to assess and guide LLMs’ capabilities in strategy. We propose that established strategy teaching simulations provide an ideal benchmarking approach because (1) they approximate the essential features of real-world strategy, and (2) they do so in a controlled, replicable environment suitable for evaluation. To demonstrate this, we assess the performance of 21 proprietary and 13 open-source LLMs on the Back Bay Battery (BBB) simulation, a widely used exercise in strategy and innovation courses. The simula­ tion requires balancing short-term profitability against long-term competitive positioning while integrating complex information about customer preferences and technological change. We built an interface enabling LLMs to interact with the simulation as though encountering it for the first time, masking identifiers to reduce contamination from prior training data. Our results show clear progress in composite BBB performance: Later models generally outperform earlier versions, and reasoning-focused models from late 2024–early 2025 (e.g., o4-mini, Claude Sonnet 4, Gemini 2.0 Flash) exceed even the average scores of his­ torical MBA student cohorts. However, frontier models from mid-to-late 2025 (e.g., GPT-5, Claude Opus 4.5, Gemini 3) have declined, underperforming both earlier LLMs and MBA students. This decline is partially explained by a systematic bias toward exploiting the core business at the expense of investing in future growth. Overall, these findings highlight impressive advances in LLMs’ strategic abilities since their inception. At the same time, we document current frontier models’ surprising weakness in managing strategic uncertainty.

Introduction. How well can artificial intelligence (AI) make strategic decisions? Strategy scholars define such decisions as resource commitments and courses of action that shape long-term performance (Mintzberg et al. 1976, Eisen­ hardt and Zbaracki 1992, Csaszar 2018), such as enter­ ing a new market or allocating resources to long-term research and development (R&D) initiatives. They are unusually difficult because they are (1) complex and Supervised-learning AI systems developed in the last decade have excelled on tasks sharing some of these features—augmenting high-stakes decisions in radiol­ ogy, information technology (IT) operations, human resources (HR) screening, and quantitative trading (Li 2017, Kleinberg et al. 2018, Arthur and Hossein 2019, Allen and Choudhury 2022, Csaszar et al. 2024); they have also surpassed human experts in complex strat­ egy games like chess, Go, and StarCraft II (Newborn 2012, Silver et al. 2016, Vinyals et al. 2019, Gaessler and Piezunka 2023). Yet these capabilities have not readily translated to real-world strategy, at least in part because supervised learning depends on wellstructured repeatable data (Choudhury et al. 2020, Csaszar et al. 2024, Felin and Holweg 2024), which are rarely, if ever, present in the context of actual strategic decisions (Eisenhardt 1989, Nickerson and Zenger 2004, McDonald and Eisenhardt 2020).

But the recent emergence of large language models (LLMs) has, for the first time, made it plausible for AI to carry out real-world strategic decisions. Unlike ear­ lier supervised learning AI systems that were limited to prespecified tasks with structured data, LLMs can parse unstructured text and generate context-sensitive responses in a flexible way (Csaszar et al. 2024). This expands AI’s capacity to engage in broader, unstruc­ tured reasoning processes that more closely resemble human cognition. As a result, AI has now surpassed various human benchmarks in coding, reading com­ prehension, multimodal reasoning, and even PhD- level science questions (Chen et al. 2021; Wei et al. 2022; Bubeck et al. 2023; Anthropic 2024; OpenAI 2024, 2025; Google 2025; Maslej et al. 2025).

Still, how well these enhanced capabilities extend to strategic decision making remains an open question. Early evidence shows GPT-3.5/4 can generate and evalu­ ate business ideas comparably to human experts (Bous­ sioux et al 2024, Csaszar et al. 2024, Doshi et al. 2025) and that using GPT-4 significantly enhances productivity in product development and consulting (Dell’Acqua et al. 2023, 2025). But these studies mostly assess one-shot, nar­ row tasks that sidestep key elements of strategic decisionmaking (i.e., uncertainty, multiperiod irreversibility, competitive dynamics, and delayed feedback). Mean­ while, recent conceptual work on AI cautions that LLMs’ backward-looking, pattern-based reasoning may lack the forward-looking causal logics central to strategic value creation (Felin and Zenger 2017, Felin and Holweg 2024, Felin et al. 2024). Further complicating the question, prior empirical studies have typically examined human–LLM collaboration for a single model (e.g., GPT 3.5/4) at a time. Because of rapid changes in the underlying technol­ ogy, standalone results for GPT-3.5/4 may not hold for later models, leaving the strategic capabilities of LLMs uncertain and state of strategy research on AI in a state of perpetual evolution as the underlying models improve.

These questions and issues persist in part because the field lacks standardized strategy-specific LLM benchmarks (Csaszar 2025). In other domains (e.g., coding, chat performance, PhD-level science, and video games), agreed upon LLM benchmarks have provided an objective yardstick for tracking various dimensions of LLM performance to inform the limits of their use in specific domains (Rein et al. 2024, Wang et al. 2024, White et al. 2024, ClaudePlaysPokemon 2025, UC Berkeley SkyLab 2025). Such benchmarks have significantly improved the quality of the systems they track, spurring on the respective fields by defining targets for further research and development. Mean­ while, LLM industry leaders have actively called for more domain-specific benchmarks to push forward model performance (OpenAI Pioneers Program 2025).

To both track and catalyze strategic decision-making capabilities, we propose a novel benchmarking approach that evaluates LLM performance on established strategy simulations. Many strategy simulations establish condi­ tions conducive to benchmarks. They are designed to mirror the complexity and uncertainty of real-world strategic decisions while still providing a controlled, vali­ dated, and repeatable environment necessary for stan­ dardized benchmarking.

Related work. GPT-4 has also approached or exceeded human-level performance in professional knowledge tests, scoring in the 90th percentile on the Uniform Bar Exam—surpassing most law graduates and vastly outperforming GPT-3.5’s bottom 10th per­ centile (OpenAI et al. 2023). In medical licensing exams, GPT-4 achieved approximately 80% accuracy (e.g., MedQA), comfortably exceeding GPT-3.5’s ~58%, and approaching the performance of medical school graduates (OpenAI et al. 2023). This research illustrates that LLMs have rapidly transitioned in a relatively short period of time from foundational competencies to expert-level performance in complex cognitive tasks.

In 2024 and 2025, OpenAI’s “thinking” models using “chain of thought” reasoning—such as o1 and o3—have continued to rapidly advance (Wei et al. 2022). They have achieved remarkable results across a range of benchmarks that were not possible with the base language models like GPT-4. The o1 model, introduced in late 2024, ranked in the 89th percentile on Codeforces and solved 83% of International Mathematics Olympiad qualifier problems, a signifi­ cant improvement over GPT-4o’s 13% (OpenAI 2024). The subsequent o3 model achieved 96.7% accuracy on the American Invitational Mathematics Examination (AIME) and 87.7% on graduate-level science exams (OpenAI 2025). Comparable perfor­ mances have been reported by Anthropic’s Claude 3 and Google’s Gemini 2.5 models (Anthropic 2024, Google 2025).

One of the frontier models as of August 2025 is the GPT-5 model from OpenAI, which has achieved top (or near top) performance across a range of bench­ marks including AIME (high-school math questions), GPQA (PhD-level science questions), LiveBench (a range of cognitive tests), and LMArena (testing chat performance) (White et al. 2024, UC Berkeley SkyLab 2025). Overall, such progress indicates the rapid and substantial expansion of AI capabilities into areas pre­ viously exclusive to highly trained human experts (Chen et al. 2021, Bubeck et al. 2023).

Further research using GPT-4 explored the reliability of AI evaluations by generating 60 startup business models and having them assessed in pairwise compari­ sons by ensembles of LLMs, which aligned closely with evaluations from human strategy experts (Doshi et al. 2025). A follow-up experiment validated these findings with 60 real-world business plans from a U.S. university startup competition (Doshi et al. 2025). Although the LLMs often produced evaluations that were inconsistent and biased, in aggregate, the rank­ ings tended to resemble those of human experts.

Method. The simulation is designed to capture the classic stra­ tegic dilemma of balancing exploration and exploitation over time (March 1991, Benner and Tushman 2003, Gupta et al. 2006). The core technology (AGM) provides reliable cash flows, but there are indications that it will face eventual obsolescence and market saturation. In contrast, the new technology (SC) represents an unpro­ ven but potentially high value emerging growth market. It begins as a niche solution with limited market appeal but has the potential to redefine market boundaries and become a major growth driver if developed to meet emerging customer needs (Christensen and Bower 1996, Adner 2002, Adner and Zemsky 2005, Christensen et al. 2018). Early investments in SC, however, do not yield immediate results. Therefore, R&D investment in the new technology reduces short-term profits; successful growth requires strategic foresight, calculated risk tak­ ing, and sensitivity to evolving market signals.

In the simulation, participants face significant com­ plexity: Market dynamics evolve based on simulated customer preferences, competitor behavior, and tech­ nological progress, none of which follow simple or eas­ ily predictable patterns. The interdependent decisions carry multiyear consequences, because neglecting one technology or overcommitting to another can erode long-term competitiveness (Peterson and Wu 2021). Feedback is delayed and noisy, because the effective­ ness of R&D investments may only become apparent several years (decision-periods) after they are made. Success in the simulation requires not only logical rea­ soning but also the ability to synthesize a deluge of possibly relevant information and make high-stakes decisions under significant uncertainty.

In the primary benchmarking exercise, we use the “advanced” difficulty in “legacy” mode, which is the standard in MBA classrooms and matches our 3.1.2. Simulation Structure and Dynamics. The BBB simulation models the challenges of managing an established business (AGM batteries) while deciding whether and how to invest in an emerging disruptive technology (SC). The simulation creates tension by lim­ iting resources: There is never enough R&D budget to fund all opportunities. A common failure mode is spreading investments too thinly across multiple initia­ tives, which prevents achieving meaningful progress in any area. Success requires carefully chosen tradeoffs, prioritization, and long-term strategic thinking.

Several programmed dynamics drive the competitive environment. Around year 4, the simulation introduces a downturn in AGM sales due to increased foreign pricing pressure. This decline is largely unavoidable; even strong management choices cannot eliminate it entirely. How­ ever, by investing in process improvements (which lower costs) and tactically reducing prices, players can soften the impact, retain some market share, and sustain profit­ ability in the core AGM business.

Meanwhile, investments in SC follow a cumulative, path-dependent trajectory. Early and sustained R&D is necessary—single exploratory investments almost never pay off. SC breakthroughs take multiple years of funding in focused performance dimensions that underperform customer demands (such as energy den­ sity or cycle life) before they begin to deliver tangible market demand. Importantly, these investments are risky: Players commit capital without guaranteed suc­ cess, only knowing that if breakthroughs occur, SC can eventually provide a far superior value proposition for select customer segments compared with AGM. When a key technological threshold is crossed, the SC market can accelerate rapidly, creating the potential for expo­ nential growth.

3.1.3. Illustrative Paths to High Performance. One via­ ble route to high performance is to focus entirely on the established AGM technology. In this scenario, players concentrate most R&D dollars on AGM process improvements (to reduce cost and improve margins). This strategy can lead to strong cumulative profit scores because of steadily improving margins in the core busi­ ness.

Discussion. This paper has presented a benchmark for the strategic decision-making capabilities of LLMs by evaluating their performance on an established business strategy simulation. Although prior research has shown that LLMs can benefit some strategy tasks (Dell’Acqua et al. 2023, 2025; Csaszar et al. 2024; Doshi et al. 2025), this is the first standardized assessment comparing multiple models’ ability to independently make complex, un­ certain, multiperiod business strategy decisions in a newly encountered uncertain environment (Table 1). Specifically, we benchmarked 21 proprietary and 13 open-source models from leading LLM providers on the BBB strategy simulation, which tests users’ ability to synthesize complex information across multiple per­ iods while balancing strategic investments in maintain­ ing a profitable core business (exploitation) and in developing an emerging business (exploration).

Our results indicate that generally LLMs have signifi­ cantly improved in their performance on the BBB simu­ lation, but that the latest frontier models have declined. Early models from 2022 to 2023, such as GPT-3.5 and GPT-4, performed significantly below later models and well below the human MBA benchmark. Models from late 2024 to early 2025 (such as GPT-4o and Gemini 2.0- Flash) performed best overall, achieving strong profit­ ability in the core market while also investing in the emerging one. These models exceeded the historical human MBA student average scores. In contrast, the current frontier models from mid-2025, such as Gemini 2.5-Pro and GPT-5, despite performing well on other benchmarks, regressed below earlier models and aver­ age historical human MBA performance.

By developing a transparent and replicable bench­ marking framework, we provide a tool for strategy researchers to systematically track LLMs’ evolving strategic decision-making capabilities over time.7 Beyond establishing an evaluation of current capabili­ ties, the benchmark allows for ongoing assessment and offers a target to improve future AI systems by highlighting the blind spots of the current frontier models. To this end, the BBB strategy benchmark is a first step—a proof of concept and a call to action for the field. Below, we distill design principles and practical recommendations for developing additional strategyspecific benchmarks. Together, these contributions are intended to guide future strategy benchmarking research, enabling scholars and developers to build, refine, and compare AI systems in ways that better reflect the distinctive demands of real-world strategic decision making.

5.1. Contributions to Research on AI and Strategy Our study makes several contributions to research on AI in strategic decision making. First, we provide one of the first systematic, empirical benchmarks compar­ ing multiple LLMs’ independent performance in an environment that mimics real-world strategy. Prior work has tended to evaluate important albeit narrow strategy tasks such as idea generation or evaluating early business plans (Dell’Acqua et al. 2023, 2025; Csas­ zar et al. 2024; Doshi et al. 2025). Through the BBB sim­ ulation, we were able to interrogate a broader set of strategic challenges that mirror real-world strategy including uncertainty, complexity, irreversible multi­ period moves, and delayed/noisy feedback. Simula­ tions like BBB, we propose, offer a rigorous, controlled, and repeatable environment that complements and extends existing approaches for evaluating LLMs.

This benchmark allowed us to demonstrate that LLMs have recently crossed a key threshold in that many models can now independently reason through complex, multiperiod, uncertain strategy problems Allen and McDonald: How Well Can AI Do Strategy?

Limitations. 5.4. Other Limitations and Future Research Because AI remains far from functioning as an inde­ pendent strategist, future research should also examine how humans use AI within simulations. Current sys­ tems lack physicality, social presence, agency, and long-term task management, making them ill suited for real-world strategic automation (Murray et al. 2021, Randazzo et al. 2025). In practice, applications will take the form of human–AI collaboration rather than full substitution (Mollick 2024, Randazzo et al. 2025). This raises a critical question: Would a current frontier model such as GPT-5 perform more effectively when guided by a skilled strategist providing well-framed prompts, or would it simply replicate the same failure modes it displays when used alone? Because most stra­ tegic work is team based, another key question is how group dynamics shape these interactions. How would teams of humans—or ensembles of LLMs—perform under different conditions? Such comparisons could reveal whether the systematic biases we observe (e.g., the exploitation tendency of frontier models) are miti­ gated or amplified in group settings.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does AI adoption reshape collaboration patterns in knowledge work? How do real-world evaluations reveal AI capabilities that benchmarks hide? How do AI systems determine and balance multiple competing objectives? What governance mechanisms can effectively constrain widely deployed AI systems? What prevents LLMs from applying their reasoning knowledge to improve outputs? How can we detect and account for LLM involvement in academic writing? Do language models reason through disagreement or only accommodate it? What causes coordination failures in multi-agent language model systems? Can language models reason beyond surface pattern matching? Can monitoring reasoning traces and behavior detect hidden agent deception? Can language models reliably simulate personas and predict behavior? How do neural networks learn compositional structure from training?