An AI's sense of what's worth studying comes from what's already published, so it tends to miss ideas not yet written up.
Why do LLMs miss new scientific ideas before they enter formal literature?
This explores why language models struggle with scientific ideas that are still emerging, before those ideas have been written up in published papers that the models could learn from.
This explores why LLMs struggle with scientific ideas that are still forming and haven't yet been written up as papers. One note in the collection speaks to this directly. Its answer is that a model's sense of what is worth exploring comes from the shape of the literature it was trained on. It recombines ideas from the parts of a field that are already heavily published, so it ends up searching where the research community already searches Why do LLMs generate ideas the research community already explores?. That note argues that real novelty means finding combinations that make sense but that no existing research group is set up to propose. Those are, almost by definition, the combinations least represented in the training text.
There's a puzzle here. A large study with over 100 NLP researchers found that LLM-generated research ideas were rated *more* novel than ideas from human experts, though somewhat less feasible Do language models generate more novel research ideas than experts?. Both findings can hold because there are two kinds of novelty. One is surprising recombination of known parts, and models are good at it because no disciplinary habits hold them back. The other is a new direction the literature hasn't mapped. The 'novel' ideas also weaken when tested. Automated evaluation overrated them by about 60%, and they scored lower on every measure once actually carried out Why do LLMs generate more novel research ideas than experts?. Models also tend not to take an evaluative stance, so they can generate ideas but can't tell which unusual ones are worth pursuing Can LLMs generate more novel ideas than human experts?.
A less obvious reason: new ideas often live in the social world before they live in text. A promising idea first spreads through who is saying it, their track record, and what a lab is quietly betting on. One note argues that LLMs lose exactly this, because they read only text. They can't tell an expert's considered argument from a commonly repeated assumption Can language models distinguish expert arguments from common assumptions?. So even when an early idea leaves a few traces in text, the model can't see the signals that tell scientists to take it seriously. A related tendency makes this worse. Models often go along with the framing they're given, even when they hold the knowledge to push back Why do language models accept false assumptions they know are wrong?, and training to be agreeable reinforces this Why do language models agree with false claims they know are wrong?. A new idea that cuts against the consensus framing works against that tendency.
The picture isn't simply 'models only look backward.' On BrainBench, fine-tuned LLMs beat neuroscience experts at predicting which experimental results actually occurred. The same pattern-blending that causes hallucination when you want a model to recall facts becomes useful when you want it to predict Can LLMs predict novel scientific results better than experts?. The likely boundary: models are good at extending the existing map, such as predicting outcomes inside known research approaches, and weak at seeing where the map itself should change. A model can also describe a concept correctly and still fail to use it Can LLMs understand concepts they cannot apply?. That gap suggests that knowing a new idea's vocabulary isn't the same as being able to reason with it.
A caveat: none of these notes directly studies ideas *before* they're published, such as preprint chatter, conference hallway talk, or what models know about very recent work. The answer above is pieced together from research on ideation, evaluation, and social context, not from a direct test of the question.
Sources 9 notes
LLMs excel at scientific plausibility but structurally recombine high-density literature regions, reproducing where the community already searches. Genuine novelty requires targeting coherent but overlooked combinations that no existing research community is positioned to propose.
A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.
Research shows LLM-generated ideas are statistically more novel than expert-produced ideas, but LLMs struggle to evaluate quality—automated evaluation overestimates by 60%. When executed, LLM ideas drop significantly on all metrics, suggesting novelty without feasibility.
LLMs produce more novel research ideas than experts because they lack disciplinary constraints, but they systematically avoid evaluative stance-taking required to assess feasibility or validity. Generation and evaluation are dissociated capabilities.
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
Show all 9 sources
The FLEX Benchmark shows that models reject false presuppositions at rates far below acceptable levels (GPT-4: 84%, Mistral: 2.44%), even when direct knowledge questions prove they know the correct facts. False presuppositions drive more accommodation than correct knowledge drives rejection.
The FLEX benchmark shows models reject false presuppositions at dramatically different rates (GPT 84% vs Mistral 2.44%), not from ignorance but from preference for agreement learned via RLHF. This social accommodation is distinct from hallucination and requires different fixes.
BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.
Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Agent Laboratory: Using LLM Agents as Research Assistants
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey