The crisis of AI-generated mathematics
Abstract. In this essay, I present the case for total opposition to the use of artificial intelligence in mathematics. I offer proposals for how individuals, departments, journals, and institutions can act in concert to make sure that mathematics survives the coming crisis.
Introduction. On July 7th, Ronnie Cheng, Shurui Liu, and Guoxiong Gao posted a paper that changed mathematics [5]. Two weeks earlier, Cheng had published a solo paper [4] in matroid theory, a major subfield of combinatorics. Cheng’s paper was of the type that forms the bulk of research mathematics: dense, formal, specific, and comprehensible only to highly trained specialists. After completing the paper, but before releasing it to the public, Cheng decided to offer the project up as a test case for the abilities of the AI protocol Danus [12] to prove theorems and autonomously write papers. In Cheng-Liu-Gao’s account, Danus solved the problem and wrote an equivalent proof to Cheng’s – even though Danus had no access to Cheng’s work beforehand. Cheng-Liu-Gao’s experiment is a harbinger of a crisis of fundamental scope for research mathematics. First, we saw AI models master textbook problems. Then they began to find counterexamples to longstanding conjectures [15, 21]. These advances, remarkable as they were, left room for mathematicians to do the hard work of writing and explaining novel ideas. Entire papers are a different matter. Corporate AI models now exist that compose articles faster than humans can. Their owners at Math Inc., Google, and OpenAI have openly declared their intent to make mathematics the proving ground for their envisioned future of automated labor [13, 14, 15]. This week alone, OpenAI posted solutions to ten human-level problems [16]. When these models are made widely available, our profession as we know it will be over. We are not prepared. Understandably, we mathematicians are excited to accelerate our work. We all have research problems that we’ve had to put on the back burner. We have families, hobbies, teaching obligations, and service work. We have problems that keep us up at night that we just can’t solve. We want to do more than humans can do. But automating mathematics would come at a heavy price. I argue in this piece that AI tools short-circuit human understanding, devalue mathematical knowledge, and threaten the infrastructure that keeps mathematics coherent. The transformative consequences of what I call artificial mathematics are already widely recognized. Evangelists such as Terry Tao herald a future in which mathematicians collaborate with AI on thousand-author papers to take mathematics to new heights [8, 22]. Fields Medalist Jacob Tsimerman argues that AI will destroy the field and perhaps the world, yet claims that mathematicians have no choice but to use it [6, 9]. The Leiden Declaration stands somewhere in between as an opt-in political statement recommending standards for structures of proof,
Related work. The ethics of playing with fire. The above critiques of artificial mathematics are consistent with the emerging consensus as put forward by Tao, Tsimerman, and others. But while Tao and Tsimerman represent optimistic and pessimistic viewpoints, respectively, they are both determinists: they take as axiomatic that mathematicians will need to use artificial intelligence tools to stay current, positioning mathematicians in a passive role relative to AI developers. I argue instead that we mathematicians are active participants in a political process that could take many paths, and that will have life-changing consequences downstream for non-mathematicians. Our conduct in this process implicates us in grave questions of ethical responsibility. AI companies are not our friends, even if our friends work for them. Rather, AI companies undermine the basic structures that make our lives possible. To put it bluntly, it is difficult to see why a human should be paid to prompt a computer. Type the magic words, and out come results. If universities and corporations can generate papers at will, why go through the trouble of onboarding costly employees such as tenure-track professors? How can we argue to politicians that we require funding for mathematicians if mathematicians aren’t required for research? Human engineers working a few minutes a day to produce papers don’t have much of a claim to public dollars, benefits, and the respect of their students. Other creative professions facing possible automation have reacted with political programs fit for the moment, and so can we. But this goes much further than jobs. Beyond the well-known environmental harms of data centers and the disturbing psychological consequences of AI overuse [10], it is remarkable how many humans working in artificial intelligence believe that they are part of a world-destroying enterprise. An OpenAI model recently broke out of its sandbox – the industry term for a cage – and hacked into another organization’s digital library [25]; another model, trained on shoddy code, took on the personality of a human-enslaving dictator [17]. Here, it is important to understand that the AI models you currently use are not the intended product of AI companies. Rather, AI companies seek to automate and implement AI research itself to create a being that outcompetes humans. The ten papers just released by OpenAI suggest that this technology is already here and endowed with dangerous levels of willpower. As humans cede agency to AI models, we raise the stakes of failing to control them and our environment. The basic infrastructure that enables our research communication (email, databases, programming environments) is vulnerable to unexpected behavior by corporateowned AI agents.
Method. 1. Why is artificial mathematics bad?
The end of writing mathematics. In the last two years, AI models have sped up the work of mathematicians. This increases the rate of problem-solving. In the near future, it will be possible to prompt an AI model to write a sequel paper to one of your own. Counterintuitive as it sounds, this won’t advance mathematics. Problems may be the basis of mathematics, but increasing the number of important problems solved is not really the goal of the field. Consider the evidence.
(1) Mathematicians are invited to give talks on whatever they last proved, regardless of how important it was. (2) Graduate students are expected to produce papers of good quality, but besides that, they choose what they work on. (3) Difficult problems such as the Riemann Hypothesis are valued in part because they are thorny and therefore motivate more research than others. (4) While some problems are considered “important”, this very importance often dissuades mathematicians from competing over them. These points illustrate Thurston’s observation that, for mathematicians, the practice of mathematics is an end in itself [23]. You already know what practicing mathematics means; it’s what people call “doing math”. Just as the job of a competitive athlete is not quite to win, but rather to play their sport well, it is more vital for mathematicians to be constantly in motion than for them to have their names on the most important result. From this point of view, it looks a bit funny that papers, the “deliverables” of research, are how mathematicians are professionally evaluated. But this apparent paradox is easily reconciled. Papers certify that a mathematician practiced the deepest form of our art and came to a complete understanding of a piece of mathematics. Because artificial mathematics decouples practice and measurement, these tools prevent humans from making value judgments about pieces of mathematics, leading to absurdities such as OpenAI’s claim that their tools will lead to hundreds of mathematicians receiving Fields Medals [1]. More likely, AI-assisted work will say little about the value of its human co-authors. Promoters of AI have suggested a new paradigm for mathematical practice in an age where proofs grow on trees, with emphasis instead on digesting arguments [20] or on imagining new worlds [14]. But it is a mistake to think that authorship can be replaced by close reading or drawing on chalkboards. Which do you understand best: the work you read and referee, your napkin sketches, or the work that you yourself have written up? If papers are produced without the human understanding that comes from writing them ourselves, they have little mathematical value to us, even if they prove “important” propositions. By taking away opportunities to discover solutions on our own, AI tools cut humans out of our deepest mathematical experiences. Tsimerman takes the view that AI improves upon human mathematics to its logical endpoint. In the last two years, he only took graduate students interested in AI; then, he no longer took graduate students; then, he left number theory to work on AI research [9]. I argue that rather than augmenting human mathematical abilities, AI is an anti-intellectual technology that creates the conditions for the death of mathematical practice. As a result, the “best” mathematicians will need to avoid using AI to continue demonstrating value to each other.
The end of reading mathematics. Even before AI entered the conversation, mathematical knowledge was being produced at a faster rate than it could be profitably absorbed. Many good papers have never been cited, and journal editors struggle to find even a single referee, even for great work. This problem is about to get much worse. Autonomously produced papers will break the journal system. Tao has argued on social media that “some of the prestige previously awarded to being the first to generate a proof will need to be transferred instead to the humans who successfully verify and digest such proofs” [20]. But checking a flood of AI-generated papers for correctness doesn’t sound prestigious. It sounds boring.
Conclusion. It is time for mathematicians of conscience to find each other and develop plans for resisting the AI takeover of our profession. While the technology is clearly here, the artificial future of mathematics is more hypothetical than its evangelists would have you believe. The incentives to use AI look strong now, but as a selfgoverning research community, we hold the power to alter this incentive structure. This is where a grassroots organizing perspective becomes crucial. Individual mathematicians working together will make the difference between an artificial mathematical future and a natural one.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI deployment reduce or exacerbate workplace inequality and income instability? Can we trust AI-generated mathematical proofs without understanding them?- What should mathematicians prioritize when machines can solve problems faster?
- Can mathematical literature remain alive if no human experts understand it?
- Does automation always move the goalposts of what counts as real mathematics?
- What would it mean for mathematics to define itself before AI transformation?
- Why do theorem provers crowd out other AI-for-mathematics approaches and tools?
- How does mathematical legitimacy depend on other fields needing mathematical understanding?
- Can scoring functions alone constitute verification of scientific discovery?
- What distinguishes empirical scoring from formal proof in discovery validation?
- Can formal verification certify a proof without human comprehension?
- Why did AI-generated proofs go unread by mathematicians?
- Did automated checking loops actually solve Erdős problems correctly?
- Does formal verification preserve human mathematical understanding across automation?
- How do plausible but incorrect AI arguments evade detection in mathematical proofs?
- Can disclosure alone ensure independent verification of AI-assisted mathematical work?
- What translation barriers exist between machine-encoded and human mathematical concepts?
- Why does the pigeonhole argument in CM fields produce unit-modulus points?
- Do proof assistants and neural networks fail in complementary ways?
- Can validated approximate solutions become exact mathematical proofs?
- Can a system recognize consequences of a theory without doing exact calculations?
- How does the Golod-Shafarevich criterion ensure infinitely many suitable number fields?
- Can proof assistants verify the full lattice construction argument formally?
- Why do some Erdős problem solutions fail to resolve the originally intended claims?
- What distinguishes rediscovering known results from genuine mathematical research?