Why does arXiv now demand prior peer review for computer-science survey papers, when AI can mass-produce them faster than moderators can check?
What counts as a survey paper versus a research contribution in arXiv?
This explores how arXiv tells a survey or position paper apart from an original research paper, and why that line suddenly matters now that AI can produce papers cheaply.
This explores where arXiv draws the line between a paper that reviews existing work (a survey or position paper) and one that adds new findings, and why that line now carries real weight. The collection doesn't spell out arXiv's formal categories. It does show that the line has become a gatekeeping decision. arXiv's computer science section now accepts survey and position papers only if they have already passed peer review somewhere else. The stated reason is volume: hundreds of submissions a month, most of them what moderators called shallow annotated bibliographies Can peer review gates stop the flood of AI-generated surveys?. So in practice, a survey is the kind of paper that LLMs can now mass-produce, and arXiv's volunteer moderators can't keep up.
That suggests a working test for what makes a survey worth reading: does it add structure, or does it just list papers? A useful example in the collection is itself a survey. It reviews 230 publications on AI in research and peer review. Instead of summarizing them one by one, it arranges them into six linked dynamics (more papers being produced, automated reviewing, manipulation, defenses, evasion, and feedback across the whole system). It also says where the evidence is strong and where it thins out Does AI create a coupled arms race in research production and review?. That is the difference between a map and a bibliography. A good survey's contribution is the framework, not the list of sources.
Research contributions are judged by a different standard: can someone else check the claim? Here the collection points to a gap. Among 24 autonomous AI research systems, 83% release their code, but only 38% release the random seeds, run traces, or novelty checks a reviewer would need to verify the results Why do autonomous research systems release code but not verification artifacts?. One system for generating papers tackles this by making the author state what evidence will count before seeing any results, and by keeping the model's judgment calls separate from steps that can be checked mechanically Can separating judgment from verification improve research paper reliability?. Seen this way, a research paper earns its label when its claims can be tested, not just because it reports experiments.
The less obvious point is that arXiv's rule moves the quality check somewhere else. It doesn't remove it. arXiv hands the job to refereed venues, but those venues are under strain too. AI-generated manuscripts have already cleared workshop review Can AI systems generate research papers that pass peer review?. Some authors hide prompts in their papers telling AI reviewers to be positive Are hidden AI prompts in preprints a deceptive research practice?. And a preprint that was never peer reviewed can shape public debate before anyone questions it, as MIT's request to withdraw an AI-and-science preprint showed Can unreviewed preprints shape scientific debate before peer review?. The survey-versus-research line used to be a filing category. It is turning into a question of who is responsible for checking which kind of claim. There is also no evidence yet that arXiv's new rule has improved quality or reduced submissions.
Sources 7 notes
arXiv CS now mandates documented prior peer review for survey and position papers, citing hundreds of monthly submissions that are mostly shallow annotated bibliographies. The change delegates quality control to refereed venues rather than volunteer moderators, but no evidence yet shows the rule improved quality or reduced submissions.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Among 24 runnable autonomous-research systems, 83% release code but only 38% release seeds or traces needed to reproduce results, and only 38% report any novelty-verification method. Code availability does not make claims checkable.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
Show all 7 sources
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI for Auto-Research: Roadmap & User Guide
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Assuring an accurate research record