Do AI advisors get swayed more by which option you list first than by actually reasoning the problem through?
Does option order matter more than reasoning depth in LLM strategic recommendations?
This explores whether LLMs giving strategy advice are swayed more by surface features, such as which option is listed first, than by how carefully or deeply they reason through the problem.
This explores whether LLMs giving strategy advice are swayed more by surface features, such as which option is listed first, than by how deeply they reason. No study in the collection measures option order and reasoning depth against each other directly. Several findings point the same way, though: surface framing moves LLM recommendations a lot, and more reasoning doesn't reliably fix that. The clearest evidence comes from 15,000 simulated strategy dilemmas. In every strategic tension tested, six different LLMs recommended the same side. Changing the industry context shifted their answers by only 11%. Simply swapping the order of the options shifted them by 19% Do LLMs consistently favor the same strategic choices regardless of context?. So the wording of the menu mattered more than the facts of the business. The authors' explanation is that the models are recombining trendy strategy vocabulary rather than analyzing the situation.
If framing is the problem, more thinking seems like the obvious fix. The corpus suggests it often isn't. In a benchmark of 23 prompts across 12 LLMs, asking for step-by-step reasoning actually lowered accuracy in high-performing models, while cheaper models gained more from rephrasing the question or adding background knowledge Do prompt techniques work the same across all LLM tiers?. A study of 22 LLMs playing strategic games found that performance tracked the structure of the game, not how much the model reasoned. Different models also settled into distinct styles: one plays defensively to limit its worst case, one leans on trust, one tries to anticipate the other player Do large language models use one reasoning style or many?. What a model reasons about, and how it frames the problem, seems to matter more than how long it reasons.
The most surprising finding is that newer isn't better. On a business strategy simulation, frontier models from mid-to-late 2025 scored below earlier models and below MBA students. They kept choosing short-term profit over uncertain investments in growth Do newer frontier LLMs actually make better strategic decisions?. This fits with the 'same side every time' result: these look like built-in leanings, not reasoning failures that more compute would fix. What does help is outside structure. In game-theory settings, LLMs drift further from optimal strategies as games get more complex. But when a structured workflow walks them through the analysis, they get close to optimal play and become much harder to exploit Do language models make rational strategic decisions in games?. The structure does the job that the model's own reasoning doesn't.
The less obvious point is that humans have the same weakness. When 348 people used an LLM to help with strategic evaluation, they considered a wider range of cues, but their forecasts were no more accurate. They also felt more overloaded and less ownership of their decisions Does using LLMs actually improve strategic decision making?. Elsewhere, users trusted AI answers more when they simply carried more citations, even irrelevant ones Do users trust citations more when there are simply more of them?. Put together, a model nudged by option order is advising a person nudged by how polished the answer looks. Neither side is reliably checking the substance. The practical lesson: if you ask an LLM for strategic advice, reverse the order of the options and ask again. If the recommendation flips, you've learned that the original answer reflected the framing more than the situation.
Sources 7 notes
Across 15,000 simulations, six LLMs recommended the same strategic choice in every tension tested. Industry context shifted bias only 11%, while option order—a framing artifact—shifted results 19%, revealing that models recombine trend-coded vocabulary rather than analyze context.
A 23-prompt benchmark across 12 LLMs shows rephrasing and background-knowledge prompts boost cheap models, while step-by-step reasoning reduces accuracy in high-performance models. Task structure, not generic best practices, determines which prompts help.
Analysis of 22 LLMs across behavioral game theory reveals three dominant profiles: GPT-o1 uses minimax reasoning, DeepSeek-R1 uses trust-based reasoning, and GPT-o3-mini uses belief-anticipation. Performance correlates with game structure, not raw reasoning depth.
Mid-to-late 2025 frontier models scored below earlier models and MBA students on a strategy simulation, systematically favoring immediate profit extraction over uncertain future bets.
LLMs frequently fail to compute Nash equilibria, with worse performance as game complexity increases. Structured game-theoretic workflows guide reasoning toward optimal strategies, reducing exploitability and enabling near-optimal negotiation outcomes.
Show all 7 sources
A 348-person experiment found that LLM-assisted evaluation broadened the cues people considered but did not improve prediction accuracy. The assistance also increased perceived overload and reduced psychological ownership of decisions.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
- Game-theoretic LLM: Agent Workflow for Negotiation Games
- Strategic Reasoning with Language Models
- Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
- Your AI Strategy Advisor Is Giving Everyone the Same Advice
- AI-Augmented Strategic Decision-Making Under Time Constraints: An Experimental Study on Mental Representations and Strategic Foresight
- Search Arena: Analyzing Search-Augmented LLMs