INQUIRING LINE

Can an AI learn to judge research quality from where papers got published, and does that hold in STEM as well as sociology?

Does institutional trace learning work equally well in STEM fields and social sciences?

This explores whether training AI on the traces institutions leave behind (such as which papers ended up in which journals) to judge research quality works as well in math and the sciences as it does in fields like management and sociology. The direct evidence covers social science only, so this answer is about what we know, what is missing, and where the two kinds of field differ.


This explores whether learning from institutional traces, meaning training a model on where research ended up rather than on written criteria for judging it, carries over across disciplines. The short answer: the collection has direct evidence for social science only. There is no matching STEM test, so the comparison can't be settled from what's here. That gap still tells you something.

The core finding comes from social science. Models fine-tuned on the publication records of eight social science fields learned to predict which research pitches would land in top-tier journals. In management they reached 59.2% accuracy. Human experts agreed with each other only 41.6% of the time, and the models also beat frontier reasoning models Can institutional publication records train better scientific evaluators?. The striking part is what the models learned. They didn't learn a stated rubric. They picked up the field's unwritten judgment, encoded in which work institutions rewarded. That matters most in fields where 'good research' is contested and experts disagree, because the institutional record is one of the few consistent signals available.

This suggests why STEM might be different, not necessarily better or worse. In math and code, there is often a direct check on whether an answer is right, and that is what reinforcement learning with verifiable rewards (RLVR) relies on. But the collection shows those checks have their own blind spots. RLVR makes reasoning look more coherent from step to step without ensuring that the proof as a whole is valid Does RLVR actually improve mathematical reasoning or just coherence?. Some apparent gains on math benchmarks turn out to be memorized test data rather than better reasoning Does RLVR success on math benchmarks reflect genuine reasoning improvement?. So even where a 'correct answer' exists, outcome signals can teach the shape of good work without its substance. The publication-tier approach has the same risk: it learns what got rewarded, which is not quite the same as what was good.

One more connection is worth a click. Research on pretraining finds that reasoning ability draws on broad, reusable procedural knowledge, while factual recall depends on narrow memorization Does procedural knowledge drive reasoning more than factual retrieval?. If institutional judgment is a kind of procedural knowledge, a learned sense of how a field decides what matters, it might transfer across fields more easily than you'd expect. Whether it does is exactly the open question.

The surprise here is that institutional learning may be most useful where the subject is most contested. Social science has weak ground truth and noisy expert agreement, so learning from institutional history beats the experts. In STEM, a model trained on publication tiers would have to compete with direct correctness checks, and nobody in this collection has run that test yet.


Sources 4 notes

Can institutional publication records train better scientific evaluators?

LLMs fine-tuned on eight social science publication records beat both expert majority votes and frontier reasoning models at evaluating research pitches, reaching 59.2% accuracy in management versus 41.6% expert agreement. The models learned field-level evaluation logic from institutional stratification rather than written criteria.

Does RLVR actually improve mathematical reasoning or just coherence?

RLVR post-training measurably reduces logical errors between adjacent reasoning steps, but locally coherent traces can still be globally invalid proofs. The improvement is structural rather than semantic.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Does procedural knowledge drive reasoning more than factual retrieval?

Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.