Can predicting scientists improve discovery forecasts?
Does modeling the researchers behind discoveries—their collaborations, expertise, and positioning—forecast future scientific breakthroughs better than analyzing research content alone? This matters because it tests whether human expertise patterns reveal signals that text cannot.
The excerpt's central claim is that a prediction model for science improves when it models the scientists who make discoveries, not only what their papers say. The authors build a mixed hypergraph of materials, properties and researchers, then sample random walks over it. Their headline result: "incorporating the distribution of human expertise by training unsupervised models on simulated inferences cognitively accessible to experts dramatically improves (up to 400%) AI prediction of future discoveries beyond those focused on research content alone, especially when relevant literature is sparse." In the drug repurposing test, the human-aware model gives "43% higher precision than identical models accounting for research content alone," measured 18 years after the 2001 prediction year.
The mechanism is a simulation of human inference. Each walk step stands for something a scientist could plausibly do: an author-author step for conversation between collaborators, an author-material step for deep familiarity with a material, and a material-material step for what a reader of two literatures could realize. The authors "hypothesize that identifying topics with high human expert density around them provides us with an informative signal regarding near-future discoveries." The excerpt verifies outcomes, not reasoning. Predictions are scored by overlap with discoveries published after the prediction year, and drug predictions are checked against expert-curated associations in the Comparative Toxicogenomics Database. "These models succeed by predicting human predictions and the scientists who will make them" is the authors' explanation. It is backed by a correlation between precision and how often a drug is mentioned in the literature (r=0.74, p<0.001), not by any checked account of a scientist's reasoning.
Against the nearest notes, this is a different kind of graph. The Mechanist note Can AI automate the discovery of how AI models work? grounds an agent in a knowledge graph of studies. This excerpt argues that content-side graphs cannot see the people positioned to make an inference. HypoEvolve How do collaboration rules shape hypothesis quality? also makes collaboration shape hypothesis quality, but its collaborators are explicit agents, where here they are inferred from co-authorship. The Large Discovery Model note Can language models reliably judge their own candidate quality? ranks candidates by a surrogate's value estimate; this paper ranks them by expert density. The sharpest turn runs the other way: tuned to avoid the crowd, the model yields "scientifically promising 'alien' hypotheses unlikely to be imagined or pursued without intervention until the distant future."
The excerpt does not establish that those alien hypotheses are good. It calls them promising and unlikely to arise otherwise, but reports no evaluation of them, and the COVID-19 progesterone case is an illustration, not a test. The materials-prediction section has only its heading in the excerpt, so the only quantified result in the body text is the 43% drug figure, from one prediction year; the 400% headline appears in the abstract and conclusion without its materials case. Co-authorship is the only human tie modeled, and the Limitations section says other ties, such as shared institutions and conference attendance, could be considered. The implication is narrow: in this drug and materials setting, with dense co-authorship and sparse literature, expert distribution adds predictive signal beyond content. Whether that generalizes, and whether the crowd-avoiding hypotheses are worth pursuing, the excerpt leaves open.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do users confuse explanation quality with actual system accuracy? Can AI systems perform peer review as effectively as humans? Does AI-assisted research sacrifice exploration breadth for productivity gains?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How do collaboration rules shape hypothesis quality?
Can we isolate and test how different ways of coordinating multiple agents affect the quality of scientific hypotheses they develop? This matters because collaboration often helps or hurts depending on conditions.
HypoEvolve also makes collaboration a variable shaping hypothesis quality, but its collaborators are explicit agents, not inferred co-authors.
-
Can AI automate the discovery of how AI models work?
Whether mechanistic understanding of AI systems—traditionally manual and slow—can be accelerated through an agentic system grounded in structured knowledge and curated methods. This matters because AI development is outpacing our ability to understand it.
Mechanist grounds discovery in a content-side knowledge graph; this paper argues such graphs miss the people who would make the inference.
-
Can language models reliably judge their own candidate quality?
LLMs fluently generate candidates across complex spaces but may misestimate their value and uncertainty. Understanding this gap matters for steering AI-driven discovery toward real experimental outcomes rather than internal confidence.
both rank discovery candidates by a predicted signal; the LDM uses a surrogate's value estimate, this paper uses expert density.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Accelerating science with human-aware artificial intelligence
- Predicting Empirical AI Research Outcomes with Language Models
- Towards End-to-End Automation of AI Research
- Large language models surpass human experts in predicting neuroscience results
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Artificial Intelligence Tools Expand Scientists' Impact but Contract Science's Focus (Just accepted by Nature, to be online soon)
- Interesting Scientific Idea Generation Using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders
Original note title
human-aware hypergraphs improve discovery prediction up to 400% over research content alone — by predicting the scientists who make them