Accelerating science with human-aware artificial intelligence
Artificial intelligence (AI) models trained on published scientific findings have been used to invent valuable materials and targeted therapies, but they typically ignore the human scientists who continually alter the landscape of discovery. Here we show that incorporating the distribution of human expertise by training unsupervised models on simulated inferences cognitively accessible to experts dramatically improves (up to 400%) AI prediction of future discoveries beyond those focused on research content alone, especially when relevant literature is sparse. These models succeed by predicting human predictions and the scientists who will make them. By tuning human-aware AI to avoid the crowd, we can generate scientifically promising “alien” hypotheses unlikely to be imagined or pursued without intervention until the distant future, which hold promise to punctuate scientific advance beyond questions currently pursued. Accelerating human discovery or probing its blind spots, human-aware AI enables us to move toward and beyond the contemporary scientific frontier.
Research across applied science and engineering, from materials discovery to drug and vaccine development, is hampered by enormous design spaces that overwhelm researchers’ ability to experimentally evaluate all candidate designs1. To face this challenge, researchers have initialized data-driven AI models with published scientific results to create powerful prediction engines. These models have begun to assist human discovery by focusing scientific attention on the subset of discovery candidates most predicted to possess propertiesrelevant to energy,2 human health3, and other economic and societal values. In this way, AI intervenes in the discovery process by proposing efficient, model-based experiments that would require much longer for unassisted human scientists to identify. Such efforts typically ignore the distribution of scientists and inventors4, however, the human prediction engines who continuously alter the landscape of discovery and invention.
Method. Incorporating Human Experts with Hypergraph Proximity We model the distribution of inferences collectively and cognitively accessible to scientists by constructing a hypergraph over research publications. A hypergraph is a generalized graph where an edge connects a set of nodes rather than a node pair. Our research hypergraph is mixed, containing nodes corresponding not only to materials and properties mentioned in title or abstract, but also the researchers who investigate them (Fig. 1c, first step). Following construction of this research hypergraph, we identify cognitively accessible inferences by generating random walk sequences over it. These walks suggest paths of inference available to active human scientists, which trace mixtures of diverse expertise sufficient for contemporary discoveries. If a valuable material property (e.g., ferroelectricity—reversible electric polarization useful in sensors) is investigated by a scientist who, in prior research, worked with lead titanate (PbTiO3, a ferroelectric material), that scientist is more likely to consider whether lead titanate is ferroelectric than a scientist without the research experience. If that scientist later coauthors with another who has previously worked with sodium nitrite (NaNO2, another ferroelectric material), that scientist is more likely to imagine whether sodium nitrite has the property through conversation than a scientist without the personal connection. In this way, the density of random walks over our research hypergraph is proportional to the density of cognitively plausible and conversationally attainable inferences. If two literatures share no scientists, a random walk over our hypergraph will rarely bridge them, just as a scientist will rarely consider connecting a property valued only in one community with a material understood only in a disjoint one (Fig. 1a). We hypothesize that identifying topics with high human expert density around them provides us with an informative signal regarding near-future discoveries. These topics might be located far from one another in terms of the number of steps required to travel between them in the hypergraph, but a random walker—and the collective scientific mind—can easily travel between them if intermediate steps are socially dense, facilitating conversation and collaboration (Fig. 1a).
To generate each random walk sequence, our model (i) initiates the walk with a valued property (e.g., ferroelectricity) as first node in the sequence, (ii) randomly selects an article (hyperedge) having mentioned that property, (iii) randomly selects a material or author from that article as next node (end of first step), then starts the second step by randomly selecting another article with the newly selected material or author, and repeats this Markov process5,14 a pre-specified number of times (see Fig. 1b for an example, and Supplementary Information for more details). Each random walk step can be viewed as a simulation of human actions: an author-author step mimicks networking or conversation between two expert collaborators; an author-material or author-property step represents how an author is deeply familiar with the selected material/property they have studied and published on; finally, a material/property-material/property step captures the potential for the transition to be realized by human scientists through reading a collection of scientific articles. From the collaborative character of physical and biological science, author nodes in our hypergraph far outnumber materials. To compensate for this imbalance, we devise a non-uniform sampling distribution parameterized by α, which roughly determines the fraction of material to author nodes in resulting sequences. Specifically, we define αwhen sampling a node from a paper (e.g., in step (iii) above) such that the probability of selecting a material is αtimes that of selecting an author (See Supplementary Fig. 1).
Discussion. Results on Anticipating Human Discoveries To demonstrate the power of accounting for human experts, we use transition probability and deepwalk metrics to build two alternative discovery predictors. These algorithms assess the relevance of the focal property to each candidate material based on literature published prior to a given prediction year (e.g., 2001) by embedding the human-aware hypergraph. We contrast our predictions with a random baseline and predictions generated from precisely replicated prior work that uses word embeddings based on the textual content of scientific literature without accounting for the distribution of human scientists15. This prior work measured property/material relevance with cosine similarity from a Word2Vec model12 trained over the contents of scientific articles published prior to the prediction year. Our experiments and evaluation framework are identical to the settings of this study in order to facilitate precise replication. Each evaluated algorithm selects the 50 materials with highest similarity to the focal property based on hypergraph or Word2Vec similarity metrics and reports them as discovery predictions. We evaluate prediction quality based on their overlap with materials discovered and published after the prediction year (see Methods for further details; for alternative evaluation metrics and prediction sizes see Extended Data Fig. 2 and Supplementary Fig. 3).
Energy-related Materials Prediction Drug Repurposing Prediction We used the same approach to explore the repurposing of ~4K existing FDA-approved drugs to treat 100 important human diseases. We used the MEDLINE database of biomedical research publications and set the prediction year to 2001 (Extended Data Fig. 1c). Ground-truth discoveries were based on drug-disease associations established by expert curators of the Comparative Toxicogenomics Database (CTD)17, which chronicles the capacity of chemicals to influence human health. Figure 2e reports prediction precisions 18 years after prediction year, revealing how accounting for the distribution of biomedical experts in our unsupervised hypergraph embedding yields predictions with 43% higher precision than identical models accounting for research content alone. We found a strong correlation between our human-aware prediction precision and drug occurrence frequency in literature (r=0.74, p<0.001), implying that our approach works best for diseases whose relevant drugs are frequently mentioned in prior research.
COVID-19 Therapy and Vaccine Prediction The success of these COVID-19 predictions suggests how fast-paced research on COVID therapies and vaccines increased the importance of scientists’ prior research experiences and networks for the therapies and vaccines they would come to imagine, evaluate and champion in clinical trials. Consider the female progesterone as a candidate material. Despite very few direct literature connections between “Coronavirus” and “Progesterone” before the rise of COVID-19, random walks from our method frequently walked the path between the two literatures through pre-COVID papers published in virology, immunology, studies regarding male/female characteristics of diseases, and the female reproductive system (Fig. 3a, Extended Data Table 1). Shortly after the beginning of 2020 and in 2021, two clinical trials were initiated with similar motivation19,20: (1) the lower global death rate of women compared to men from COVID-19, and (2) the anti-inflammatory properties of progesterone that may moderate the immune system’s overreaction to COVID-19 in men19.
Conclusion. We demonstrate the power of incorporating human-awareness into artificial intelligence systems for accelerating future discovery. Our models succeed by directly predicting human discoveries and the human experts who will make them, yielding up to 400% improvement in prediction precision. These findings offer support for the influence of the human experience and social connection inscribed by our research hypergraph in driving scientific advance. This suggests that the search underlying materials and medical advance is dominated by local exploitation of the familiar over novel exploration of the unknown. Moreover, by tuning our algorithm to avoid the crowd, we generate promising hypotheses unlikely to be imagined, pursued or published without machine recommendation for years into the future. By identifying and correcting for collective patterns of human attention, formed by field boundaries and institutionalized education, these models complement the contemporary scientific community. This demonstrates that connectivities in our expert-aware hypergraph are useful not only for predicting and accelerating human discoveries in the near future, but also for inferring disruptive discoveries that could be imagined by scientists only in distant future.
Limitations. Our analysis examined a limited space of scientific relationships–those between a material possessing a valuable energy or therapeutic property. Many other scientifically meaningful relationships lie beyond this syntax, such as identity (i.e., a is a b), composition (i.e., a is a part of b), or any specific physical or logical relationship (e.g., a chemically reacts with b; a genetically up-regulates b). Using a hypergraph formalism, we could extend such relations beyond logical triples that connect a simple concept pair to larger sets of concepts connected by more complex relations. Another limitation involved our singular consideration of co-authorship as the relationship affecting the distribution of expertise. One could consider other relationships, such as scientist collocation within an institution, at a conference they attend, or geographical proximity. Moreover, there are opportunities to technically improve our approach, such as combining content and human-aware information to amplify prediction accuracy, or inferring and exploiting the body of negative knowledge in science where researchers know that certain scientific claims are false11,21.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do users confuse explanation quality with actual system accuracy? Can AI systems perform peer review as effectively as humans? Does AI-assisted research sacrifice exploration breadth for productivity gains?- Why does expert density around topics signal near-future discoveries?
- How do co-authorship patterns alone capture scientist expertise without surveying reasoning?
- How do community summaries and selective traversal differ as graph scaling strategies?
- Can fixed heuristics like PageRank match learned traversal policies on graphs?
- What graph structures better support multi-hop reasoning than pairwise edges?
- How does quasi-local structure in bipartite graphs differ from global graph patterns?
- Can knowledge graph structure help embeddings represent more combinations?
- Can hierarchical entity extraction from books enable both textual and visual reasoning?
- How does map-reduce over communities compare to flat multi-hop retrieval architectures?
- Can inference-time query decomposition replace pre-built knowledge graph structures?
- Why do binary edges lose information when representing multi-entity relations?
- What is the computational cost of constructing and traversing hypergraphs?
- How does hypergraph accumulation differ from single-pass graph retrieval?
- Can hyperedges replace triple-based externalization in reasoning tasks?