ASI-Bench: At the Dawn of Artificial Superintelligence

Paper · arXiv 2608.17271 · Published August 18, 2026
LLM Evaluations and Benchmarks

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today’s AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems’ capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation.

Introduction. A central challenge on the path toward artificial superintelligence (ASI) is whether AI can move beyond mastering existing human knowledge to explore unfamiliar problems, develop new solutions, and turn them into verifiable results. Today’s AI systems derive much of their capability from learning, compressing, and applying the accumulated knowledge of humanity, and have made rapid progress in scientific reasoning, coding, data analysis, and agentic execution [1, 2, 3, 4, 5, 6, 7]. Yet existing evaluations largely test these capabilities either through problems with known answers or through tasks whose methods and procedures are substantially specified by humans. They therefore provide limited evidence about whether AI can autonomously conduct scientific research when both the problem and the path to a solution are open-ended. In this work, we ask a more direct question: how far can current AI systems independently explore and execute project-level scientific research as human methodological guidance is progressively withdrawn?

Discussion / Conclusion. Beyond a single leaderboard, ASI-Bench is intended to serve as shared research infrastructure for the scientific and AI communities. Its B1–B4 structure and domain-level analyses enable controlled comparisons across models and agents, while revealing where progress occurs and where human methodological guidance remains necessary. Yet no fixed benchmark—and no single research team—can fully represent the breadth of difficult, meaningful, and verifiable problems that future AI systems must confront. The continued value of ASI-Bench therefore depends on collective participation from researchers working across disciplines and at the frontier of model development.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can brute-force automated research substitute for iterative depth and human research intuition? When should work require human-AI partnership versus full automation? How should designers communicate what AI systems truly are and can do? How does the generation-verification gap limit what we can measure about AI reasoning? What training dynamics and scale trigger emergence of reasoning capabilities? How do capability benchmark scores systematically misrepresent true model abilities? How do evaluation practices shape which failures stay visible? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? What fundamental constraints limit how effectively agents can improve themselves? Do reasoning benchmarks predict model performance in long-horizon workflows? What safeguards enable trustworthy AI-assisted scientific peer review at scale?