Why do autonomous research systems release code but not verification artifacts?
Autonomous research systems publish their code at high rates, yet rarely share the seeds, traces, or novelty checks that would let reviewers verify their claims. What explains this gap and what would close it?
The survey codes 26 full-text entries from a screen of 125 candidates (35 included works: 24 runnable systems plus two study or position works) on seven audit dimensions. Its main finding is a split between two kinds of release. "Code release is now common (83% of the 24 runnable systems), but the artifacts and checks that let a reviewer verify a result are not." Only 38% release "the seeds or execution traces needed to reproduce a run," and only 38% report any novelty-verification method. The 22-system LLM-era subset shows the same pattern. These are the authors' own codings of public records, and they call the rates "directional audit evidence."
The authors locate the problem in what evaluation measures. "Most reported evaluation reduces to task success or a reviewer score," while trustworthy research would also need evidence on validity, novelty, reproducibility, and selection, which are "rarely measured." Novelty is the sharpest case: 38% of systems report a novelty-checking step, yet "we found no system that reports independent validation that its novelty check is itself reliable." The contrast is drawn against older lineages. The Robot Scientist "logged hypotheses and provenance in machine-readable form, so each claim was auditable by construction," whereas LLM research agents inherited the generation machinery and not the checks: "the capability transferred; the verification did not."
This sits closest to Can AI verify research outputs as fast as it generates them?, which draws the generation-verification gap from a roadmap's qualitative findings. This excerpt counts one stretch of that gap and places it in release practice: publishing code does not make a claim checkable. The authors' answer is disclosure. Their reviewer checklist asks for seeds, traces, selection policy, and novelty-check method, which fits the governance reading in Does more automation actually hide rather than eliminate errors?. They also name independent, self-preference-robust agent review as an open problem, which is where Can inference scaling help reviewers catch errors humans miss? would need to be tested; the excerpt does not assess that work.
What the excerpt does not establish is how exact these rates are. The coding covers only computational AI/ML research, where "code, experiments, benchmarks, and write-ups can be inspected," so the sample says nothing about other fields. Second-coder agreement was 90% on artifact release but only 50–65% on autonomy level, novelty method, and selection disclosure. The reliability check used abstracts only, so it does not validate the full-text reading behind the headline figures, and the excerpt gives no selection-disclosure rate at all. The implication at the strength the evidence allows: in this corpus, code availability and verifiability diverge sharply, and the 83% versus 38% contrast is a fair directional signal. The size of the gap is not settled until the full-text second pass the authors name as the next step.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do restrictions on reviewer LLM use actually shape peer review behavior? What human oversight must AI research systems have? What external process records should verify agent behavior and benchmark claims? Why does AI verification capability persistently exceed generation capability?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can AI verify research outputs as fast as it generates them?
Research suggests AI systems produce plausible findings rapidly but struggle to verify them at the same pace. This creates a bottleneck in verification across all research stages. Understanding this gap matters for assessing when AI assistance is reliable versus risky.
the survey gives a counted instance of the same gap, located in release practice rather than in generation
-
Does more automation actually hide rather than eliminate errors?
As AI systems become more polished, do they mask failures instead of preventing them? This matters because it changes whether we should focus on detecting problems or governing their disclosure.
the survey's answer is reporting requirements, consistent with governance over detection
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
the survey names independent, injection-robust agent review as open; the excerpt does not assess PAT
-
Can separating judgment from verification improve research paper reliability?
Explores whether dividing model-based decisions from deterministic checks and fixing evidence requirements before observing results could bound errors in automated paper generation and make AI-assisted research more trustworthy.
a design separating judgment from checkable operations, the kind of artifact the checklist asks for; not assessed here
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- AI for Auto-Research: Roadmap & User Guide
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
Original note title
autonomous research systems share code more often than the artifacts a reviewer needs to verify claims — 83% release code, 38% release seeds or traces