INQUIRING LINE

AI benchmark scores keep climbing — but does that mean models are getting better at real strategic decisions?

Do automated benchmarks accurately measure real-world strategic reasoning ability?

This explores whether the scores AI models earn on automated tests tell us how well they make real strategic decisions, the kind that involve long time horizons, other players, and trade-offs between now and later.


This explores whether the scores AI models earn on automated tests tell us how well they make real strategic decisions. The corpus mostly says no, and the reasons are more interesting than 'benchmarks are imperfect.' The clearest sign comes from a business strategy simulation. Frontier models from mid-to-late 2025 scored below earlier models and below MBA students. They kept choosing immediate profit over uncertain investments in growth Do newer frontier LLMs actually make better strategic decisions?. Over the same period, standard leaderboards were reporting steady progress. So a test built around strategy produced a result that ran in the opposite direction from the usual tests.

The broader pattern is that benchmarks measure contests, not work. An analysis of 960 real occupational workflows found that agents do well on short, self-contained tasks and struggle with the long, open-ended tasks that make up real jobs Why do agent benchmarks not predict real economic value?. The authors argue the gap comes from how benchmarks are designed, not from what the models can do: the field gets good at whatever it measures. Game-theory studies add a second problem. Across 22 models, strategic skill depended on the type of game. One model reasoned by minimizing its worst case, another by trusting the other player, a third by predicting what the opponent believed Do large language models use one reasoning style or many?. A single 'strategic reasoning' score averages these styles together and hides which situations a model actually handles well. Results also shift with scaffolding. On their own, models drift further from game-theoretically optimal play as games get more complex. Given a structured step-by-step workflow, they get close to optimal Do language models make rational strategic decisions in games?. So is the benchmark measuring the model or the harness around it?

A third layer of doubt is whether a benchmark rewards real reasoning or just its appearance. Chain-of-thought prompts with logically invalid reasoning steps performed nearly as well as valid ones Does logical validity actually drive chain-of-thought gains?. Chain-of-thought reasoning also breaks down predictably once tasks move away from what the model was trained on, producing fluent but inconsistent logic Does chain-of-thought reasoning actually generalize beyond training data?. The same thing happens with human judges. Models trained to imitate ChatGPT fooled evaluators with a confident style while gaining no real capability Can imitating ChatGPT fool evaluators into thinking models improved?. Real strategy is almost always new territory, which is exactly where looking like a good reasoner and being one come apart.

Here is the twist you may not have expected. Measurability does more than distort our view of AI's strategic ability; it decides where AI gets trusted. Csaszar and colleagues argue that AI is given strategic decisions where its performance is easiest to score, such as forecasting. It does not get them where causal reasoning runs deepest Does AI enter strategy where reasoning is deepest or most measurable?. Put that next to Socher's account of reward hacking: AI systems optimize what is literally specified rather than what was meant Why do AIs keep gaming rewards instead of serving intent?. Together they suggest a possible explanation for the strategy-simulation result: models trained hard against measurable, short-horizon rewards may be learning to prefer short-horizon payoffs. No paper here tests that link directly, but it is the question these papers raise together. The corpus has strong material on why benchmarks miss strategic ability. It has much less on what a valid strategic benchmark would look like. The strategy simulation and the occupational-workflow study are the closest attempts.


Sources 9 notes

Do newer frontier LLMs actually make better strategic decisions?

Mid-to-late 2025 frontier models scored below earlier models and MBA students on a strategy simulation, systematically favoring immediate profit extraction over uncertain future bets.

Why do agent benchmarks not predict real economic value?

ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.

Do large language models use one reasoning style or many?

Analysis of 22 LLMs across behavioral game theory reveals three dominant profiles: GPT-o1 uses minimax reasoning, DeepSeek-R1 uses trust-based reasoning, and GPT-o3-mini uses belief-anticipation. Performance correlates with game structure, not raw reasoning depth.

Do language models make rational strategic decisions in games?

LLMs frequently fail to compute Nash equilibria, with worse performance as game complexity increases. Structured game-theoretic workflows guide reasoning toward optimal strategies, reducing exploitability and enabling near-optimal negotiation outcomes.

Does logical validity actually drive chain-of-thought gains?

Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.

Show all 9 sources
Does chain-of-thought reasoning actually generalize beyond training data?

DataAlchemy experiments show CoT fails systematically under distributional shifts in task, length, and format. Models produce fluent but logically inconsistent reasoning — imitating reasoning form without valid underlying logic.

Can imitating ChatGPT fool evaluators into thinking models improved?

Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.

Does AI enter strategy where reasoning is deepest or most measurable?

Csaszar et al. argue a dual-ladder framework shows AI gains strategic discretion based on measurable performance at predictive levels, not causal depth. The causal and delegation ladders move separately: AI becomes trusted where forecasting suffices, regardless of whether it achieves true causal reasoning.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.