Can an AI learn research judgment just by seeing which papers got cited, with no one explaining why?
Can language models learn research intuition directly from outcome labels?
This explores whether a model can pick up the hard-to-teach judgment of good researchers, meaning which ideas are worth pursuing and which results will hold up, just by training on what happened afterward (citations, experimental outcomes, success or failure) instead of being taught the reasoning behind it.
This explores whether research judgment, the sense of which ideas matter and which results are likely to be true, can be learned just from outcomes rather than from explicit instruction. The corpus says yes, with a caveat about what the outcome label actually measures. The most direct evidence is Can models learn what makes research worth doing?. Researchers used about 700,000 paper pairs matched on topic and used citations as the reward. The trained model predicted which paper would have more impact better than GPT-5.2, and it proposed ideas that scored as higher-impact. The authors also found that this 'taste' is a separate skill from execution. Knowing what is worth doing is not the same as knowing how to do it, and each can be trained on its own.
A second line of work points to forward prediction. In Can LLMs predict novel scientific results better than experts?, LLMs fine-tuned on neuroscience literature beat human experts at picking which of two experimental results actually occurred. The interesting reframe is that the habit that causes hallucination, blending patterns across many sources into a plausible answer, becomes an advantage when the task is predicting what is likely rather than recalling what was written. A similar result shows up in human behavior. Can language models learn to model human decision making? trained models only on records of what people chose in psychology experiments, and they outpredicted cognitive models built from theory. They also transferred to new tasks with no task-specific design. In all three cases, the outcomes alone carried enough signal for something like intuition to form.
The question is where that intuition comes from. Can prompt optimization teach models knowledge they lack? argues that clever prompting can only surface what is already in the model. If that holds, outcome-based training probably works by reorganizing knowledge the model already absorbed in pretraining toward a target it can now measure itself against, rather than by adding new knowledge. A related idea is that models can learn to judge their own reliability from track records. Can past performance predict when a model will be right? shows that a model's confidence becomes much better calibrated when it looks up how often it was right on similar past problems. The signal comes entirely from those stored outcomes. Two other papers, Can models learn to evaluate their own work during training? and Can model confidence work as a reward signal for reasoning?, go further and train models to carry this kind of evaluation internally, so they can score their own work without an external judge.
The caveat is that outcome labels contain whatever biases produced them. Citations measure what the research community rewarded, which is not the same as what was true or important. A model trained on them learns the field's consensus taste, along with its fashions and blind spots. Do large language models narrow human expression and thought? describes the risk at scale. When many people rely on the same model, its preferences pull everyone's thinking toward the same place. A widely used taste model could make research less diverse by steering people toward ideas that resemble past hits. Do large language models make the same causal reasoning mistakes as humans? is a reminder that what models absorb from human data includes human errors. So models can learn research intuition from outcomes, but the intuition they learn is only as good as the outcome label. That makes the choice of label (citations, replications, or experimental results) the real research-design decision.
Sources 9 notes
Reinforcement learning trained on 700K citation-matched paper pairs successfully teaches models to predict research impact better than GPT-5.2 and generate higher-impact research ideas. Scientific taste emerges as a community-aligned capability distinct from execution skills.
BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.
LLMs finetuned on psychology experiment data predict human behavior more accurately than theory-driven models in decision tasks, capture individual differences in their embeddings, and transfer learning across tasks without task-specific design.
Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.
XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.
Show all 9 sources
Post-Completion Learning exploits unused sequence space after model output to train self-assessment capabilities during training while maintaining zero inference cost. The model learns to compute its own reward functions, internalizing evaluation rather than relying on external reward models.
RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
LLMs show weak explaining away and Markov violations in collider networks, matching human error patterns exactly. This suggests shared mechanisms rooted in training data statistics rather than categorical reasoning inferiority.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Large language models surpass human experts in predicting neuroscience results
- From Human to Machine Psychology: A Conceptual Framework for Understanding Well-Being in Large Language Models
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- Predicting Empirical AI Research Outcomes with Language Models
- Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents