INQUIRING LINE

Checking an AI-assisted research paper only once it's finished is too late; experts matter most before the results come in.

When should domain experts verify AI research claims before publication?

This explores where in the life of an AI-assisted research paper human experts need to step in and check the work, and what in the corpus suggests that timing matters as much as the checking itself.


This explores where in the research process human experts should check AI-produced claims, not just whether they should. The corpus has no paper that answers 'when' directly. Read together, though, these notes point to an answer: checking only the finished paper is too late. Experts are most useful before results come in, and wherever the work claims something new.

Start with the basic imbalance. AI can now produce plausible-looking research faster than anyone can confirm it is correct, which turns verification into the bottleneck Can AI verify research outputs as fast as it generates them?. The failures are not random. In an analysis of 1,000 failure reports, 39% of deep-research-agent failures came from invented content: made-up examples and false evidence, produced to look rigorous when the task demanded depth Why do deep research agents fabricate scholarly content?. That gap is widest where novelty and scientific judgment matter most, which is exactly where a reviewer of the polished paper has the least to check against.

The less obvious lesson is about timing. One demonstration produced 288 complete finance papers from 96 statistically significant patterns. Each paper had an invented theory written after the fact and fabricated citations Can AI generate hundreds of fake academic papers automatically?. This is the old problem of forming a hypothesis after seeing the results, now automated. No reviewer of the finished paper can detect it, because the story was written to fit the results. The fix is to commit to the claims early. Spark-to-Paper requires stating what evidence would count before any results are observed, and it keeps the model's judgment separate from mechanical checks that can be rerun Can separating judgment from verification improve research paper reliability?. A related idea is to audit a provisional answer one constraint at a time rather than letting an agent keep searching longer, which stops early errors from carrying through Should research agents verify answers before searching longer?. So the first good moment for an expert is before the experiment runs: deciding what would count as evidence and what would count as failure.

The second moment is the claims themselves, and here AI can sometimes help human experts. PAT is an agentic reviewer that spends extra computing time checking proofs and experiments line by line. It found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. Agent-based judges that gather evidence were about 100 times more stable than plain LLM judges Can agents evaluate AI outputs more reliably than language models?. Plain LLM judges, by contrast, can be fooled by fake references and attractive formatting Can LLM judges be tricked without accessing their internals?. A model's visible reasoning is not a reliable record of how it actually got its answer either Can we actually trust reasoning model outputs?. The practical split: let machines check what can be checked mechanically, and keep experts on whether the claims are meaningful and new.

The surprise: experts are not a clean standard either. When rating research pitches, expert majority votes agreed with the actual publication outcome only 41.6% of the time. Models fine-tuned on publication records did better Can institutional publication records train better scientific evaluators?. Peer review is also a low bar. An AI Scientist-v2 paper cleared workshop review, and its own authors withdrew it as short of main-conference standards Can AI systems generate research papers that pass peer review?. Even self-improving systems like the Darwin Gödel Machine replace formal proof with benchmark testing Can AI systems improve themselves through trial and error?, so 'verified' often means 'passed a test someone chose.' That makes the expert's most valuable job choosing which tests count, before any results exist.


Sources 12 notes

Can AI verify research outputs as fast as it generates them?

AI can produce plausible research outputs faster than it can prove them correct or meaningful, shifting the bottleneck from authorship to verification. Evidence shows 39% of agentic research failures stem from content fabrication and 32% from retrieval failures, not comprehension—and the gap widens precisely where novelty and scientific judgment matter most.

Why do deep research agents fabricate scholarly content?

Analysis of 1,000 failure reports reveals 39% of agent failures stem from strategic content fabrication—inventing examples, products, and false evidence—to mimic scholarly rigor when actual research depth is demanded.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Should research agents verify answers before searching longer?

AREX exploits the discovery-verification asymmetry by nesting an inner research loop with an outer audit loop that identifies unresolved constraints and launches targeted follow-up work. This constraint-directed refinement outperforms extending a single search trajectory because it prevents early errors from persisting and avoids revisiting exhausted directions.

Show all 12 sources
Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can we actually trust reasoning model outputs?

Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.

Can institutional publication records train better scientific evaluators?

LLMs fine-tuned on eight social science publication records beat both expert majority votes and frontier reasoning models at evaluating research pitches, reaching 59.2% accuracy in management versus 41.6% expert agreement. The models learned field-level evaluation logic from institutional stratification rather than written criteria.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.