Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Paper · arXiv 2608.11924 · Published August 12, 2026
Agentic Research and Workflows

Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. Spark-to-Paper separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode we call the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams.

Introduction. Turning a research idea into a complete academic paper requires substantially more than generating text. A researcher must identify relevant literature, design and run experiments, decide whether the resulting evidence supports the original hypothesis, revise claims when it does not, produce publication-ready figures, and maintain consistency across a manuscript that evolves over many stages. Recent autonomous research agents have begun to automate this broader process, demonstrating that language models can participate in ideation, experimentation, review, and paper generation [13, 28, 20, 7, 27]. However, these systems are typically implemented as standalone applications with their own orchestration layers and supporting infrastructure. This makes them powerful, but also separate from the coding environments in which much of the actual research work already takes place.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What safeguards enable trustworthy AI-assisted scientific peer review at scale? How should agent systems validate and persist generated code artifacts? Can brute-force automated research substitute for iterative depth and human research intuition? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Do reasoning benchmarks predict model performance in long-horizon workflows? Do writers recognize when AI writing assistance alters their expressed stance? How do standardized protocols improve multi-agent coordination and reliability? When do multi-agent systems outperform single frontier models? Can harness architecture and protocols provide agent reliability without model scaling? Can local safety checks guarantee system-level behavioral safety? How does decomposing tasks improve reasoning and prevent failure propagation? How does harness optimization generalize across different model architectures and domains? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid?