Can an AI that models what scientists would expect find unusual ideas worth chasing, and tell the good ones from merely strange ones?
Can human-aware models identify scientifically promising alien hypotheses reliably?
This explores whether AI systems that model what human scientists already know and would think of can find good hypotheses that humans would likely miss, and then judge which of those unfamiliar ideas are worth pursuing.
This explores whether AI that models human scientists, meaning what they know, where they look and what they would plausibly propose, can point to promising ideas outside that territory and pick out which ones are actually good. The corpus has no paper that tests this exact setup. It does cover the separate pieces of the problem, and those pieces suggest the hard part is not producing unusual ideas. The hard part is knowing which unusual ideas are worth anything.
Start with what 'human-aware' requires. To steer away from what people would think of, a system first needs a good model of how people think. That piece is surprisingly feasible. Language models fine-tuned on psychology experiment data predict human decisions better than theory-built cognitive models do, and they capture differences between individuals Can language models learn to model human decision making?. GPT-4.5 judges social appropriateness better than any single human, which suggests a model can map a human landscape from the outside Can AI learn social norms better than humans?. There is a catch, though. All the models made the same systematic mistakes on unwritten norms. A model's picture of 'where humans are' has its own blind spots, so the area it flags as unexplored may just be the area it can't see. The same pattern shows up in values: 106 LLMs cluster in a narrow, idealized region of value space while humans scatter widely Do large language models actually reflect human value diversity?. Models may be less alien to each other than they are to us.
Next, judging promise. One result is encouraging. A fine-tuned GPT-4.1 with paper retrieval picked the better of two AI research ideas 77% of the time and beat expert researchers on a subset, while off-the-shelf models did no better than chance Can machines learn to predict which research ideas will work?. Look at how that ability was built: by training on how past ideas actually turned out. That kind of training favors ideas that look like past successes, which is the opposite of what you want when hunting for alien hypotheses. In materials and molecular discovery, LLMs are good at generating valid candidates but poor at estimating their value or their own uncertainty. They became reliable only when paired with statistical surrogate models fitted to real experimental results Can language models reliably judge their own candidate quality?. The lesson carries over: the model can propose the idea, but judging whether it's any good has to come from the world.
There is also some evidence that escaping familiar patterns is possible when there's a fast way to check results. In one bilevel autoresearch setup, an outer loop read the inner loop's code and invented new search mechanisms that broke its habitual patterns, improving performance about 5x on a GPT pretraining task Can an AI system improve its own search methods automatically?. That worked because every new idea could be scored immediately. Most scientific hypotheses don't get that luxury. Without a check, high-confidence pattern-finding can slide into what one note calls resurrected pseudoscience: correlations that look like discoveries Can AI models be truly free from human bias?.
The twist you might not expect: if these systems work, the bottleneck moves to people. A machine built to find ideas humans wouldn't think of also produces ideas humans are badly placed to judge, and it produces them faster than anyone can evaluate. That is the 'epistemic hyperinflation' scenario, where AI generates claims faster than human judgment can check them Can AI generate knowledge faster than humans can evaluate it?. Agentic judges that collect evidence are far more consistent than plain LLM judges, but their errors cascade through memory Can agents evaluate AI outputs more reliably than language models?. So does it work reliably? Not yet, going by this corpus. Generating alien ideas looks tractable. Ranking them by promise works only where real outcomes or experiments can correct the model.
Sources 9 notes
LLMs finetuned on psychology experiment data predict human behavior more accurately than theory-driven models in decision tasks, capture individual differences in their embeddings, and transfer learning across tasks without task-specific design.
GPT-4.5 outperformed every individual human at judging social appropriateness across 555 scenarios, challenging the theory that embodied cultural experience is necessary. However, all AI models share identical systematic errors on unwritten norms.
Analysis of 106 LLMs across 625 scenarios shows they cluster in a concentrated region of value space while human respondents scatter widely. Models are poor surrogates for diverse populations despite exhibiting coherent value systems.
A fine-tuned GPT-4.1 combined with paper retrieval reached 77% accuracy predicting which of two AI ideas performs better, beating 25 expert NLP researchers 64.4% to 48.9% on a 45-pair subset. Off-the-shelf models performed at chance level, suggesting the capability requires both retrieval and fine-tuning on historical outcomes.
LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.
Show all 9 sources
An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.
Research shows that 'theory-free' AI models mask bigotry behind high accuracy metrics while committing fundamental statistical errors. A 95% accurate criminal justice system would wrongly convict thousands, demonstrating that model sophistication does not validate causal inference.
AI produces knowledge faster than human judgment can verify it, collapsing epistemic confidence just as monetary hyperinflation collapses purchasing power. The gap self-reinforces because evaluation tools are themselves AI-generated, trapping the system in acceleration.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- On Epistemic Diversity in Large Language Models
- A Rational Analysis of the Effects of Sycophantic AI
- AI Models Exceed Individual Human Accuracy in Predicting Everyday Social Norms
- Bilevel Autoresearch: Meta-Autoresearching Itself
- Predicting Empirical AI Research Outcomes with Language Models
- Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
- Determinants of LLM-assisted Decision-Making