SYNTHESIS NOTE
Topics›AI at Work›this note

Can AI forecasters beat expert humans at venture evaluation?

Do frontier large language models outperform experienced managers and investors at predicting fundraising success? This matters because venture assessment is a genuinely uncertain, ill-structured judgment task where human expertise is assumed essential.

Synthesis note · 2026-10-09 · sourced from AI at Work

In a fully prospective tournament using Kickstarter campaigns launched after the training cutoffs of all tested models, frontier LLMs outperformed human forecasters at strategic foresight. Thirty live U.S. technology ventures were evaluated "while fundraising remained in progress and outcomes were unknown," and a battery of frontier and open-weight LLMs completed 870 pairwise comparisons predicting fundraising success. Benchmarked against 346 Prolific-recruited experienced managers and three MBA-trained investors working under monitored conditions, "human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 — correctly ordering nearly four of every five venture pairs." Neither wisdom-of-the-crowd ensembles of humans nor human-AI hybrid teams outperformed the best standalone model.

The paper frames the result as evidence for what it calls "unbounding rationality" (Csaszar 2025) — AI freed from the bounded-rationality constraints (limited information processing, inconsistent attention, computational limits) that decades of research (Simon 1947; Kahneman et al. 1982, 2016; Tetlock 2005) show degrade human judgment under uncertainty. The design is built specifically to rule out training-data leakage and lookup: campaigns post-date every model's training cutoff, a roughly 500-word anonymized summary strips the venture's name, the platform, funds raised to date, and consumer comments, and testing confirmed the anonymized summaries were difficult to identify via search. Outcomes had not yet been determined when forecasts were made and were pre-registered to Zenodo before results were known. The authors treat strategic foresight — judging "wicked," "ill-structured," irreducibly uncertain problems (Churchman 1967; Simon 1973; Knight 1921) — as categorically harder than domains like chess or protein folding, where ground truth is fixed and verifiable.

This sits alongside Can language models beat human venture capital experts?, but makes a stronger claim: VCBench's human benchmarks (tier-1 VC precision of 5.6%) were already near chance, so a modest LLM edge sufficed to win. Here the human range (0.04–0.45) is wider and sometimes substantial, yet the best LLM still clears it by a wide margin (0.74), and the fully prospective design forecloses the leakage objection a historical dataset like VCBench cannot. It also extends Can retrieval-augmented language models forecast like human experts? — where that system only neared the crowd aggregate, this tournament's LLMs surpass individual expert forecasters outright on a genuinely novel strategic-judgment task, without retrieval augmentation. The structure recalls Can machines learn to predict which research ideas will work?: both use pairwise-comparison tournaments to show an LLM beating domain experts, though this venture tournament finds off-the-shelf frontier models already ahead, with no fine-tuning or retrieval needed to win.

The excerpt does not establish why LLMs forecast better — whether from broader information synthesis, freedom from managerial overconfidence or anchoring, or some artifact of how the 500-word summaries were written and scored. The sample is narrow: 30 technology-category campaigns on one platform over one 72-hour window, evaluated by Prolific-recruited managers (not necessarily trained investors) alongside only three MBA investors. Whether the result generalizes to higher-stakes, longer-horizon strategic decisions — M&A, market entry, CEO succession — where outcomes take years to resolve and cannot be reduced to a single crowdfunding total, is untested here. The finding that hybrid human-AI teams and crowd ensembles underperform the best standalone model is reported but not mechanistically explained, which leaves open whether human input is actively harmful to blend in or merely redundant.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do real-world evaluations reveal AI capabilities that benchmarks hide? What prevents LLMs from applying their reasoning knowledge to improve outputs?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 93 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

a fully prospective venture tournament finds LLMs outforecast human forecasters at strategic foresight — pooling humans and AI does not help