Predicting Empirical AI Research Outcomes with Language Models
Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea’s chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks. We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs published after our base model’s cut-off date for testing, and 6,000 pairs for training. We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with. In the NLP domain, our system beats human experts by a large margin (64.4% v.s. 48.9%). On the full test set, our system achieves 77% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation. We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests. Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent. Our system achieves 63.6% accuracy, demonstrating its potential as a reward model for improving idea generation models. Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research.
Introduction. Many promising-looking AI research ideas turn out to be ineffective when executed. Unfortunately, the only way to find out is to actually implement them. The failed ideas can add up to a significant cost in both human labor and computational resources. Better prioritization—implementing the more promising ideas first—can thus significantly improve research efficiency. But doing so requires predicting the outcomes of experiments without actually running them, a seemingly impossible task.
We hypothesize that language models (LMs) can do this task better than human experts. Humans develop such research intuition through experience, but LMs can acquire it more efficiently by consuming countless research papers, analyzing experimental results, and potentially discovering subtle patterns that are difficult for humans to identify.
Concretely, given the description of a pair of research ideas, our goal is to predict which one works better on a set of benchmarks. For example, given two jailbreaking methods, we aim to predict which one achieves higher attack success rates. This is an objective task: we can obtain the groundtruth by actually implementing both ideas. But making a prediction without implementing the ideas is difficult, and any non-trivial accuracy could be valuable as it informs resource allocation, prioritization of experiments, and iterative refinement of ideas. Crucially, this task provides more actionable information than evaluating subjective aspects of ideas such as novelty or excitement [14, 9, 4], which existing work focuses on.
We construct a benchmark to facilitate the study of this task by scraping both ideas and results from existing conference papers. Each example consists of a research goal specified by a set of benchmarks, two competing ideas, and a binary outcome label indicating which idea performs better across these benchmarks. After four rounds of human verification, we obtain 1,585 verified test examples. To avoid data contamination, we ensure that each test example contains at least one idea published after June 1st, 2024, the knowledge cut-off date of our base model GPT-4.1.
We then develop a system that combines a specialized GPT-4.1 fine-tuned on 6,000 historical idea pairs from the training set and a paper retrieval agent. This system achieves a promising 77% accuracy on the test set. Off-the-shelf frontier models (e.g., o3, Claude 3.7 Sonnet)—even when augmented by the same paper retrieval agent—perform no better than random guessing, demonstrating the importance of proper capability elicitation. We then identify a challenging test subset consisting of 45 NLP idea pairs, and recruit 25 NLP experts to establish a human baseline. Specifically, each human prediction is made by an ensemble of 5 experts who collectively spent over 45 minutes. Our system beats this strong expert baseline by a large margin (64.4% v.s. 48.9%).
To test the system’s generalizability, we design stress tests to measure its sensitivity to superficial features like idea recency and complexity. On three human-designed stress tests and hundreds tests proposed by LMs, our system demonstrates robust behavior.
Lastly, we test our system on brand new, unpublished research projects from a recent study on using AIs to generate ideas [14]. This is a set of 35 ideas with experiments implemented by NLP researchers (recruited by the study’s authors). These projects have never been released publicly anywhere on the internet, thus avoiding any possible contamination. Our system achieves an accuracy of 63.6%, indicating its generalizability and its potential to improve existing automated research systems [14, 9] as an idea ranker or a reward model.
Related work. Accelerating Research with AI. With the overarching goal of using AI to build AI, recent work has explored integrating AI into the entire research workflow, including literature review [2], ideation [14, 4], experimental validation [6, 9], paper writing [9], and review [8]. Notably, one fully AIgenerated paper passed the peer-review process of an ICLR workshop [16]. However, discovering such a workshop-level idea is non-trivial: Starting from 256 candidate papers, the authors manually select the top-3 papers since LM judges are often unreliable, while only one paper gets accepted. Importantly, executing 256 AI research ideas into whole papers is expensive and slow because of the significant GPU or API costs. To accelerate AI research, our paper aims to answer the question: “can we predict the outcomes of ideas before actually implementing them?”
Research Evaluation. Prior work mainly focuses on peer-review style evaluation; the goal is to predict scores that align with human reviewers based on fully written papers. However, such evaluations still require implementing ideas. Moreover, peer-review scores are often subjective. Critiques of novelty, excitement, and paper presentation are often independent of the ideas’ actual empirical effectiveness. Consequently, humans can get misled to prefer “fancy” (e.g., mathematically complex) yet ineffective ideas. Instead, this papers focuses on objective empirical effectiveness, for each we can reliably establish ground truth labels.
Research Outcome Prediction. Prior work also attempts to use AI for predicting empirical research outcomes. The most relevant work is [10], which uses AIs to predict neuroscience results. However, their actual experiment setup is entirely different from ours. Specifically, given a published paper abstract, they use LMs to alter descriptions about results while keeping the method and background unchanged. The goal is then to distinguish the original abstract from its altered counterpart. This task is easy, even a fine-tuned Llama-2-7B model achieves an 80%+ accuracy. In contrast, we study a more realistic and challenging scenario: given two independently written idea summaries that exclude any empirical results, predict which is more effective.
Forecasting. Our task is a type of forecasting, i.e., forecasting the outcomes of ideas before running actual experiments to get the outcome. Prior attempts on neural forecasting methods demonstrate the gains from retrieval and scale (e.g., model and data scale) [19, 13, 1, 7]. For example, [5] builds a human-level forecasting system by retrieval augmentation and fine-tuning on historical data.
Method. We define our task as predicting the outcomes of two competing research ideas for a given research goal. We focus on pairwise evaluation, where the binary outcome label is aggregated over multiple benchmarks to avoid the ambiguity and noise when evaluating individual ideas [14, 9, 4]. Each example involves the following components (Figure 2):
• Idea pair, each defined by a detailed description following a standard format.
• Empirical research goal, defined by a set of benchmarks, each with a quantitative metric.
• Binary outcome label, indicating which idea wins on more benchmarks.
3.1 Retrieval When predicting outcomes of new ideas, human researchers often draw inspiration from existing literature. While the exact same ideas do not exist, prior studies can offer indirect insights, transferable knowledge, or analysis of similar sub-components. Therefore, we develop an agentic retrieval module that iteratively performs the following four steps: query generation, paper retrieval, paper summarization, relevance checking and filtering.
Step 1: Query Generation. At iteration t, given a research goal, two research ideas, and previous queries and retrieval results, our system first determines whether sufficient information has been collected, thus allowing for early exit. If further research is required, the system prompts an LM to generate a new query distinct from previous queries. Importantly, since at least one idea in the given comparison pair is entirely novel, the system would not try to directly search for identical idea comparisons. Instead, the query generation employs two strategies: 1) retrieving ideas that offer indirect or transferable insights, and 2) decomposing the novel idea into sub-components and retrieving related literature for these sub-components.
Step 2: Paper Retrieval. We use https://exa.ai as the paper search engine. Unlike conventional keyword search, they support neural search that can directly take natural language queries as inputs, e.g., “effectiveness of confidence calibration and iterative refinement in question answering”. EXA also supports automatic query optimization [12] for better retrieval performance. For each query, we retrieve the top-15 relevant papers from arxiv.org. To prevent information leakage from retrieval, we only retrieve papers published before June 1st, 2024.
Step 3: Paper Summarization. Prior work mainly resues paper abstracts as summaries [14, 9]. However, abstracts are often too brief to cover rich details (e.g., ablation studies or comparison with specific baselines). Therefore, we download each paper’s PDF, and prompt LMs to summarize the paper with respect to the current query. As shown in Table 3, this substantially boosts the accuracy from 38.8% to 53.0%. We find that using abstracts alone biases the model towards favoring old ideas, while summarizing the whole paper alleviates such biases.
Step 4: Relevance Checking and Filtering. The initially retrieved 15 papers per query ensure coverage, but may also introduce irrelevant papers that may negatively mislead LMs’ predictions. To alleviate this, we prompt LMs to judge the relevance of each paper summary (binary classification: relevant or irrelevant), and subsequently filter out those irrelevant ones.
3.2 Fine-tuning Next, we fine-tune LMs to reason over the research goal, two research ideas, and all retrieval results, to make the final prediction. A straightforward setting is fine-tuning LMs with golden outcome labels.
In addition, we also explore fine-tuning LMs to generate chain-of-thought (CoT) reasoning [15] before the final prediction. High-quality CoTs are not available for this task. Thus, following recent practices [17], we use the LM to augment the training data with its own generated CoTs. Specifically, we sample multiple CoTs for each query, filter out CoTs that lead to incorrect predictions, and try several strategies to select one CoT from all candidates (e.g., random or LM-based selection).
Discussion. Future Work. Our fine-tuning method is straightforward but works well. Future work can explore advanced modeling methods, e.g., inference-time simulation of experiments. Our system can also be used as a reward model to improve automated ideation systems.
Conclusion. A crucial skill in empirical research is being able to predict which ideas are more likely to work out. Human experts develop this skill by reading papers and observing experiments, we show that LMs can do this more efficiently; our LM system outperforms human experts at making direct predictions about the outcome of AI research without running experiments. In addition to higher accuracy, our system is robust to biases like recency that humans are prone to. The system also generalizes to completely novel, unpublished ideas, including AI-generated ones another, demonstrating the potential of further accelerations of AI research.
Limitations. Our current system functions as a black-box and is not conducive to human-AI cooperation beyond prioritizing research ideas based on model predictions. And despite our best attempts to verify its robustness, we cannot rule out the possible reliance on spurious features.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI-assisted research sacrifice exploration breadth for productivity gains?- Can AI agents align their ideas with future research directions as well as humans do?
- Is idea quality or execution capacity the actual bottleneck in AI research?
- Does AI-driven scooping narrow which research topics get explored publicly?
- Why do current AI systems struggle with researcher judgment and taste?
- What role should humans play in reviewing and approving AI-generated research?
- How do template requirements limit AI research systems from true autonomy?
- Are paper mills using NHANES data to automate single-factor research?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Can human reviewers detect when papers have been rewritten by AI?
- Do AI reviews depend more on writing style than scientific merit?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Can automated review systems catch deep methodological flaws or only surface issues?
- Could AI improve peer review rigor and catch human-missed errors?
- Could AI feedback work as a substitute for human peer review entirely?
- How can arXiv and journals scale quality control for AI-generated research?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Do AI-generated research reviews score papers higher than human reviewers do?