Do newer frontier LLMs actually make better strategic decisions?
A strategy simulation benchmark tests whether the latest LLMs can balance short-term profit against long-term growth investment—a core challenge in real strategic reasoning.
Allen and McDonald benchmark 21 proprietary and 13 open-source LLMs on the Back Bay Battery (BBB) simulation, a strategy-teaching exercise used in MBA courses that requires balancing an established, cash-generating technology (AGM) against investment in an unproven, slow-to-pay-off emerging technology (SC) over multiple uncertain periods. They report "clear progress in composite BBB performance" through late 2024–early 2025, when reasoning-focused models (o4-mini, Claude Sonnet 4, Gemini 2.0 Flash) "exceed even the average scores of historical MBA student cohorts." But "frontier models from mid-to-late 2025 (e.g., GPT-5, Claude Opus 4.5, Gemini 3) have declined, underperforming both earlier LLMs and MBA students," a decline the authors say is "partially explained by a systematic bias toward exploiting the core business at the expense of investing in future growth."
The paper frames BBB as a deliberate correction to existing LLM benchmarks, which it argues fail to "capture the defining elements of strategic decision making: uncertainty, complexity, irreversible multiperiod moves, and delayed or noisy feedback." Prior studies of LLMs and strategy, the authors note, "mostly assess one-shot, narrow tasks" like generating or evaluating business ideas (citing Dell'Acqua et al., Csaszar et al., Doshi et al.) that "sidestep key elements of strategic decision making." BBB instead runs the model through a multiyear simulation where early SC investment "reduces short-term profits" and only pays off after sustained, cumulative funding — a structure designed to reward foresight over pattern-matching to immediate reward signals, and to reveal whether a model manages that tradeoff rather than just score well on a single pass.
This sets the paper's within-run exploitation bias alongside two neighboring findings about what a high score on a long-horizon task actually indicates. Do frontier AI agents actually conduct novel research or just optimize? finds that even strong-scoring agents converge on composing known techniques rather than genuine novelty; BBB's frontier-model regression is a sharper version of the same dynamic — not just a ceiling on novelty, but models actively retreating to the known, profitable path when faced with an uncertain one. Both results sit with Do automated benchmarks hide what frontier AI systems can really do?'s broader point that standard benchmarks misrepresent capability on long-horizon, real-world-shaped tasks — BBB is offered explicitly as that kind of corrective instrument for strategy specifically, and its result (newer models scoring worse) is itself evidence that other benchmarks these same frontier models lead on were not capturing this capability.
The excerpt does not explain why mid-to-late-2025 frontier models regress — whether from training changes, risk-averse alignment tuning, or something else — nor does it report effect sizes, model-by-model breakdowns, or statistical tests behind "partially explained by." It also rests on one simulation and one MBA-cohort comparison population, so the exploitation bias is established for BBB, not for strategic reasoning generally. At the strength the evidence allows: general benchmark leadership (math, science, coding) does not predict performance on multiperiod strategic tradeoffs, and the newest frontier models should not be assumed better than their predecessors at decisions requiring sustained investment under uncertainty.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does AI adoption reshape collaboration patterns in knowledge work?- How much do profit levels determine whether managers pay for frame-expanding search?
- Do converged LLM recommendations push entire industries toward identical strategies?
- What strategic decisions do humans keep when AI handles forecasting?
- Do time constraints and AI assistance reshape strategic thinking in opposite directions?
- Can industry-specific context overcome LLM tendency toward trendy strategic choices?
- Does option order matter more than reasoning depth in LLM strategic recommendations?
- Can LLMs forecast performance improve with retrieval augmentation on venture tasks?
- Do LLMs generalize venture forecasting skill to other strategic foresight domains?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do frontier AI agents actually conduct novel research or just optimize?
Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.
both find high scores on long-horizon tasks compatible with avoiding genuine novelty or risk rather than demonstrating it
-
Do automated benchmarks hide what frontier AI systems can really do?
Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?
BBB is the strategy-domain instance of this corrective: a long-horizon, real-world-shaped benchmark that standard benchmarks miss
-
How close are frontier AI models to expert work quality?
GDPval benchmarked frontier models on 1,320 expert-built tasks across 44 occupations, using head-to-head expert judgment to measure whether AI is approaching human deliverable quality in knowledge work.
contrasts one-shot deliverable quality, where frontier models approach experts, with BBB's multiperiod setting, where the newest frontier models fall behind both predecessors and MBA students
-
Why do AI-delegated firms stop exploring new business models?
When firms hand all strategy decisions to AI agents, do those agents and their human overseers rationally stop searching for genuinely novel approaches, even when better options exist outside their current awareness?
Extends A: B's model shows AI strategies converge to equilibria that favor exploiting the core business, matching A's regression finding
-
Can AI forecasters beat expert humans at venture evaluation?
Do frontier large language models outperform experienced managers and investors at predicting fundraising success? This matters because venture assessment is a genuinely uncertain, ill-structured judgment task where human expertise is assumed essential.
Qualifies A: B finds frontier LLMs beat managers and MBA investors at forecasting venture success, bounding A's regression to this task
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament
- AI-Augmented Strategic Decision-Making Under Time Constraints: An Experimental Study on Mental Representations and Strategic Foresight
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- When Artificial Intelligence Does Strategy: Learning, Good Times, Lock-in, and Human-Driven Strategic Renewal
- Large Language Models Often Know When They Are Being Evaluated
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations
Original note title
Allen and McDonald's strategy simulation benchmark finds frontier LLMs regress toward exploiting the core business over growth investment