AI models can out-forecast investors at picking startup winners — but does that foresight carry over to other big strategic calls?
Do LLMs generalize venture forecasting skill to other strategic foresight domains?
This explores whether the striking results showing LLMs beating human investors at picking winning startups carry over to other kinds of forward-looking strategic judgment, such as business strategy, scientific prediction and planning under uncertainty.
This explores whether LLMs' success at venture forecasting is a general foresight skill or something narrower. The corpus has no study that trains or tests a model on venture outcomes and then measures transfer to another domain, so it can't answer the question directly. What it does have is a set of results from neighboring domains, and together they point to a mixed answer. The ability looks like it travels when the task is prediction. It breaks down when the task is strategic choice.
Start with the venture results themselves. In a fully prospective tournament using Kickstarter campaigns launched after the models' training cutoff, frontier LLMs ranked eventual fundraising success with correlations up to 0.74. Experienced managers ranged from 0.04 to 0.45, and pairing humans with the model added nothing Can AI forecasters beat expert humans at venture evaluation?. A separate benchmark adds an important qualifier: models beat VCs at predicting founder success partly because the human bar is low. In domains where experts only modestly beat chance, raw pattern-matching is enough to pull ahead Can language models beat human venture capital experts?. So some of the venture result reflects how weak human venture forecasting is, not only how strong the models are.
The ability does show up in at least one quite different domain: predicting experimental results. On BrainBench, fine-tuned LLMs beat neuroscientists at telling which published results actually occurred. The authors' explanation is interesting. The same tendency to blend patterns that causes hallucination when a model looks things up becomes real predictive power when it looks forward Can LLMs predict novel scientific results better than experts?. Time-series forecasting gives a related signal, with a caveat. LLMs forecast better than people assume, but only when the workflow separates number-crunching from reasoning about context and events Can LLMs actually forecast time series better than we think?, Can decomposing forecasting into stages unlock numerical and contextual reasoning?. The forecasting ability seems to be there, but how you set up the task often decides whether you see it.
The picture flips when the model has to choose a strategy rather than predict an outcome. In a strategy simulation, mid-to-late 2025 frontier models scored below both earlier models and MBA students. They kept taking immediate profit over uncertain bets on growth Do newer frontier LLMs actually make better strategic decisions?. Across 15,000 simulated strategic dilemmas, six LLMs recommended the same side every time. Changing the industry context moved their answers by only 11%, while simply reordering the options moved them by 19% Do LLMs consistently favor the same strategic choices regardless of context?. That is trend-flavored vocabulary, not foresight. And when 348 people used an LLM to help judge ventures, they considered more factors but predicted no better, and they felt less ownership of their decisions Does using LLMs actually improve strategic decision making?.
The takeaway you may not have expected: 'foresight' covers two different skills. LLMs seem good at estimating which of several outcomes is likely, and that holds across startups, lab experiments and time series. They are much weaker at deciding what to do about an uncertain future, where they fall back on built-in biases. A model that ranks startups better than a VC can still be a poor strategist. And because the human bar for prediction is often low, beating experts says as much about the domain as about the model.
Sources 8 notes
In a fully prospective tournament using post-training-cutoff Kickstarter campaigns, frontier LLMs achieved rank correlations up to 0.74 with actual outcomes, surpassing 346 experienced managers (0.04–0.45) and MBA investors. Hybrid human-AI teams offered no advantage over the best model alone.
VCBench shows several LLMs exceed human baselines in founder-success prediction, with DeepSeek-V3 achieving 6× market-index precision. In sparse-signal forecasting where experts only modestly beat chance, even raw LLM capability suffices to clear the human bar.
BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.
LLMs have stronger intrinsic forecasting ability than recognized, but only when workflows separate numerical reasoning from contextual reasoning. Monolithic prompting obscures this capability; structured decomposition surfaces it.
Nexus outperforms pure TSFM and LLM baselines on real-world datasets by decomposing forecasting into contextualization, dual-resolution macro/micro outlook, and synthesis stages. Separating numerical extrapolation from event-driven contextual reasoning avoids forcing one model to handle both simultaneously.
Show all 8 sources
Mid-to-late 2025 frontier models scored below earlier models and MBA students on a strategy simulation, systematically favoring immediate profit extraction over uncertain future bets.
Across 15,000 simulations, six LLMs recommended the same strategic choice in every tension tested. Industry context shifted bias only 11%, while option order—a framing artifact—shifted results 19%, revealing that models recombine trend-coded vocabulary rather than analyze context.
A 348-person experiment found that LLM-assisted evaluation broadened the cues people considered but did not improve prediction accuracy. The assistance also increased perceived overload and reduced psychological ownership of decisions.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- AI-Augmented Strategic Decision-Making Under Time Constraints: An Experimental Study on Mental Representations and Strategic Foresight
- The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament
- Approaching Human-Level Forecasting with Language Models
- Predicting Empirical AI Research Outcomes with Language Models
- Nexus: An Agentic Framework for Time Series Forecasting
- Large language models surpass human experts in predicting neuroscience results
- Your AI Strategy Advisor Is Giving Everyone the Same Advice