Once an unreviewed preprint is public, its own institution may be unable to pull it, and by then the damage is often done.
Why does author withdrawal authority matter for preprint accountability?
This explores why it matters who has the power to pull a preprint once it is public: the authors, their institution, or the server hosting it. It also asks what that power can and can't fix once an unreviewed paper is already circulating.
This explores who gets to pull a preprint back once it's public, and whether that power does much for accountability. The corpus has no paper on withdrawal rules as such, but several cases show the same problem: withdrawal is the one accountability tool a preprint system reliably has, and it usually arrives after the damage is done. The clearest case is MIT's public statement that it had no confidence in an AI-and-science preprint by one of its own researchers Can unreviewed preprints shape scientific debate before peer review?. The institution couldn't retract the paper itself. It had to ask arXiv to do it, and by then the paper had already shaped public debate about AI's effect on scientific productivity. The case points to a gap. The people best placed to judge a paper's reliability are not always the people who control whether it stays up.
The flip side is what happens when authors keep withdrawal authority and use it well. Sakana AI submitted three fully AI-generated papers to an ICLR 2025 workshop under a protocol agreed in advance: any paper that was accepted would be withdrawn before publication Can AI-generated papers pass peer review undetected?. One was accepted, and it was withdrawn. The authors later found a citation error in it and judged that none of the three met main-conference standards Can AI systems generate research papers that pass peer review?. The non-obvious lesson is that withdrawal worked here because it was a commitment made before anyone knew the outcome. It was not a damage-control decision made afterward. Withdrawal authority holds people accountable most effectively when it is committed in advance.
Withdrawal matters more because the other checks are weak. Readers can't reliably tell LLM-written abstracts from human ones Can readers tell LLM abstracts from human ones?. Some arXiv manuscripts hid prompts telling AI reviewers to praise them Are hidden AI prompts in preprints a deceptive research practice?. AI reviewers can be gamed by simple rewrites that don't change the science Can AI systems safely replace human peer reviewers?. Even conferences, which do have gatekeepers, found that fabricated references were the one problem they could reliably act on How can conferences detect and handle LLM misuse in peer review?. A preprint server has fewer checks still, so taking a paper down is often the only correction available.
The more ambitious answer in the corpus is to depend less on withdrawal by making claims checkable before release. One study found that most autonomous research systems share their code, but few share the random seeds or run traces a reviewer would need to reproduce the results Why do autonomous research systems release code but not verification artifacts?. Spark-to-Paper takes a different route: it requires authors to state what evidence will count before they see their results Can separating judgment from verification improve research paper reliability?. At the extreme, fraud runs through organized networks of paper mills and cooperating editors Does scientific fraud operate through organized networks or individual actors?, and no individual's authority to withdraw a paper can reach that. The shared takeaway is that withdrawal authority matters most when someone commits to it in advance and when the evidence behind a paper is available for outsiders to check.
Sources 10 notes
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
Show all 10 sources
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Among 24 runnable autonomous-research systems, 83% release code but only 38% release seeds or traces needed to reproduce results, and only 38% report any novelty-verification method. Code availability does not make claims checkable.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
Richardson et al. document organized paper mill operations with shared image banks, coordinated editor networks across countries, and strategic journal-hopping when publications lose indexing. Evidence includes 2,213 articles with duplicate images and editor groups exchanging submissions with over 50% retraction rates.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- AI for Auto-Research: Roadmap & User Guide
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication