Mathematical Research and Artificial Intelligence
Published:
Everyday, it seems AI is yet again impacting society in alarming ways. Recently, a model being developed in an OpenAI sandbox broke out of its sandbox and hacked into HuggingFace which raises lots of concerns about cybersecurity and data privacy. There’s also been a letter to the US government signed by many large AI companies, asking for regulations. They basically said: “Because of intense competition, we (the AI companies) are not able to unilaterally agree on the regulations and need another entity to enforce the regulations.” It’s similar to how these companies have said, “Our models are too dangerous and powerful.” Now they are saying, “Collectively, we are too dangerous and powerful so need oversight.”
Related to this, US AI companies are facing competition against Chinese AI models and the kinds of productivity gains and economic benefits are not being delivered as promised. Might the AI bubble burst soon? A final thing to mention here that is closer to me is that 2026 Fields medalist Jacob Tsimerman is leaving academia to join OpenAI. The announcement came just a few hours before the Fields medal was announced (of course, he knew ahead of time he’d be receiving the medal). I want to think about how math research is being impacted by AI (and other academic fields). I quote at length several sources but informally cite them simply by linking the public source; hopefully that is okay in this informal format. I think the authors make their points well without my editorializing.
Terence Tao’s 2026 ICM Public Lecture on AI
In a simplistic sense, mathematicians prove theorems. There’s a joke that mathematicians turn coffee into theorems and the follow-up is that they turn cotheorems into ffee. Anyways, proving theorems is what their job as researchers tends to focus on, at least if we interact with mathematicians only through their publications. In reality, things are not so neatly divided and the life cycle of a theorem is messy. Consider the following chart (from Terence Tao’s 2026 ICM public lecture) though keep in mind that there are many interactions of some of these steps before something more polished emerges.

- Proof generation
- Proof verification
- Proof exposition
- Proof publication
- Proof digestion
- Proof canonicalization
AI can help with almost all of these though at this stage, it might not be good at some, such as having the creativity to generate proofs in all areas of math. But more and more, it seems strong enough to solve really challenging problems in particular areas of math and may soon be good in all areas. Here are 10 more problems AI has supposedly resolved and formalized in Lean (this, following a resolution of Erdős unit-distance conjecture). OpenAI says that it cost about $2000 worth of tokens for these 10 problems and from a cursory look, these seem to be significant problems that have been unresolved for some time. See here for a write up of these problems. Interestingly, one of the arguments is invalid in a way that would be obvious to a mathematician. More on this later.
As a point of clarification, Lean is a tool for formalizing proofs and can do proof verification. It’s not the easiest for humans to write Lean code but AI certainly is able to write the code and more over, write English sentences as well to explain it. In my view, proof digestion is perhaps the only completely human task. The point is for human mathematicians to understand, accept, and value the proof, not for the AI to do so. So it’s up to humans to do the digesting of a proof. They can also make decisions for how to make something canonical in, say, textbooks. But AI can help with that as well.
So then, what are some reasons for doing math research? Tao gives some answers:
- To solve unsolved problems (both pure and applied).
- To develop new theories and techniques.
- To contribute to the shared network of mathematical knowledge.
- To understand the world around us.
- To build a community of mathematicians [added by me: some of whom derive enjoyment from research]
- To train the next generation of mathematicians to guide its future directions.
- To create enduring works of aesthetic value
- etc.
The first three may be aided by AI but the rest seem very much human activities to me. But also, these other reasons for doing math research seem at risk because of AI. For example, how can a mathematical community be built if AI drastically reduces the number of mathematicians?
To get started, let me give a working definition. An expert in a scientific field is someone who, unassisted, has a great depth of knowledge and practical experience. They have great understanding and can effectively communicate that understanding at many levels. They have intuition and creativity in that field and thus, can make good hypotheses and predictions and also formulate a plan for testing these hypotheses and predictions. There’s more to it but that’s a starting definition.
AI can help experts but for non-experts, it may be an obstruction for a novice to becoming an expert. If the novice uses AI and trusts it without understanding or verification, are they really on the road to becoming an expert? I’m concerned that there will come a time when there are no more human experts. A PhD thesis used to take years of training and study but now, AI can solve 9/10 open problems (each of higher caliber than the typical PhD thesis) in a short time. What would incentivize humans to go the slow, difficult route when this is available, albeit at the cost of having true and deep understanding? Maybe some people are willing to just check AI-generated math full time. And if there is no interest in going the traditional route, then will there be “training of the next generation” or “building a mathematical community?”
There’s also the aspect of aesthetic beauty. Maybe AI can keep discovering new science indefinitely but it will run out of human-generated data and will need to build on AI-generated science. What is the use, meaningfulness, or aesthetic value of AI-generated knowledge if no humans are around to understand or value it? Perhaps we can use the new science to build things that have direct impact and create value. But that seems rather unlikely or risky, to allow AI to build things which we have little understanding of.
To an outsider, maybe they’ll say, “How’s it different for a person to have AI help them with knowledge about math? If they learn from an AI and the AI guides them through, aren’t they mastering the subject?” I don’t contend that AI can be useful as a tool but a tool is firstly, much more effective in the hands of an expert than novice and can even be unwittingly dangerous for a novice. Would you hand a very sharp chef’s knife to a child who has never picked up a knife or cooked with no instruction at all? Also, consider, say, a pianist learning a piece by Mozart. AI can tell the pianist a lot about the history of the piece, about Mozart, about the styles contemporary to Mozart, etc. But the pianist still needs to practice the piece in order to play and perform. Arguably, there’s not much value in the Mozart piece unless it’s performed for other humans and math is similar. There are analogous skills which need to be practiced and cannot be gained by reading explanations.
Moreover, for a novice watching a pianist perform, there are many things they will miss about the technique and expression. But for a concert pianist watching another concert pianist, they will have much more appreciation of the skill and and understanding of how much work is needed to get to that point. Thus, there is a community of pianists who appreciate the aesthetic value. The same is true with mathematics (and again, other fields). While it can be good from one perspective for AI to accelerate this process of acquiring skills, it can’t completely replace it. The concern is that many people will try to skip acquiring certain skills. I’m right hand dominant and so my preference in learning a piano piece was to focus on my right hand. But obviously without the left hand, the piece is incomplete.
Jacob Tsimerman’s Thoughts
Historically, having a Fields medal is positively correlated with having a stable career in academia as a mathematician. I don’t know how true it is now but it’s disturbing to me that Tsimerman, who has a Fields medal, either doesn’t think he’ll have a stable career (I don’t think that’s the case) or that the job/field will change so much that it won’t look at all like the job/field he has enjoyed for most of his life. I’ll put some excerpts of the Tweet below. But the upshot is that he thinks problem-solving, which is a large part of math research, will change drastically and even if math research is still viable in some way, he personally might not enjoy it anymore.
“The claim
The short version is that I think problem-solving is an immense, and pervasive part of modern mathematical research. Consequently, if human problem-solving disappears by virtue of the AIs becoming strictly and substantially better at it, then most of the time currently spent by modern mathematical researchers will have to be spent on an activity that is altogether pretty different. Whether such an activity is viable as a professional endeavour is something I am unsure of, but strongly encourage others to think about and try to envision, so that if/when the time comes, we can steer such a future into being.
[…]
People have written a lot about Theory building vs. Problem-solving, and I want to first of all clarify I have nothing against theory building or theory builders! It is a valuable part of mathematics, and while there are differences in perspective between the “camps” there is way more mutual respect and agreement.
However, I gather there is a perception that theory-builders spend most of their time not-problem-solving, and I think this is largely untrue. Now I’m not a theory-builder primarily (though I’ve partaken a LITTLE BIT by necessity) so I am outside of my comfort zone. As such, I apologize for mistakes and welcome corrections!
But theory-building constantly runs through problem-solving. Let’s say you want to define the right notion of a cohomology theory. Of course you must make candidate definitions. But then what does it mean for it to be the right one? Well, you start asking if it has natural properties. These are T statements. Does it satisfy a Kunneth formula? Is it functorial in the right way? When you have the wrong one you have to find the properties it’s missing, and when you have the right one you have to prove that it indeed has those properties.
[…]
…imagine you had access to an AI oracle that could resolve statements T, but somehow lacked any creativity to build technology or make definitions (I think this is unlikely, but for the purpose of this thought experiment lets imagine it). How would your mathematics change, if you were a theory builder? Well, you make a definition, and want to know if it’s the right one. You immediately ask your oracle a thousand questions. From “are these basic properties true” to “ooh, so is this deep conjecture true?” and start getting back answers, and amending your definitions. You could invent and resolve entire research directions in days. But the confusion you would have had to push through to flesh out your theory would largely (probably not entirely) be instantly resolved and the whole process sped up tremendously by your oracle. A big part of the process would be gone. This is very very different to modern mathematics.
One more thought
This post is too long already, but I’ve seen some people say that they only do mathematics to find truth and others valourize that as the only virtuous way to be. I do not do mathematics only to find truth. I do it largely because I enjoy it and I am good at it. I also find it beautiful and am grateful I get to spend my days understanding beautiful things. But I enjoy the challenge, the process, resolving confusions, finding strategies, grappling with problems. I would like to push for this being de-stigmatized. Mathematicians are people who need money, housing, food, love, exercise, and a great deal of other stuff including various forms of meaning. There are many people whose primary enjoyment of math comes through problem solving in one of its incarnations. If that disappears, that is not a trivial issue and many of them might not want to do it anymore (even if there were some way to proceed).”
Claim: A Counterexample to Connes’s Rigidity Conjecture
For any countably infinite group $G$, we can study complex-valued functions $f:G \to \mathbb{C}$. Among these, we have the square-summable ones which satisfy $\sum_{g \in G} |f(g)|^2 < \infty$. This forms the vector space $\ell^2(G)$. Note that $G$ acts on these functions by left multiplication. If we take the complex group algebra $\mathbb{C}[G]$, we want to extend the action of $G$ in some linear way to $\mathbb{C}[G]$; that way is left convolution. The weak operator closure of $\mathbb{C}[G]$ gives $L(G)$, the von Neumann algebra.
Murray-von Neumman showed that $L(G)$ is a $\text{II}_1$ factor if and only if $G$ is ICC, that is every non-identity conjugacy class of $G$ is infinite. To be a $\text{II}_1$ factor it cannot be broken down into smaller independent algebra pieces (it is a factor) and it admits a unique, finite, and well-behaved notion of “dimension” or “trace”. Note that if the center $Z(G)$ has a nonidentity element $z$, then its conjugacy class is simply itself which is finite. So to even be a factor, we need $G$ to have trivial center.
Connes found that these $\text{II}_1$ factors associated to ICC groups possessing a property called Kashdan’s property (T) exhibit strong rigidity under small perturbations. This leads to a conjecture:
Connes’ Rigidity Conjecture: Let $\Gamma$ be a countable discrete ICC group with Kazhdan’s property (T). If $\Lambda$ is any countable discrete group (not necessarily ICC or has the Kazhdan property) such that $L(\Gamma) \cong L(\Lambda)$ as von Neumann algebras, then $\Gamma \cong \Lambda$ as groups.
Popa showed that the functor from these ICC groups of property (T) to factors is countable-to-one. He conjectured it must be at least finite-to-one.
The OpenAI team claimed to have found infinitely many ICC groups of property (T) which have isomorphic von Neumann factors but are all pairwise non-isomorphic. They have over 37,000 lines of Lean 4 code to support this. However, this rebuttal paper by J. L. Nielsen points out two flaws in the paper, either which would make the paper invalid.
We’ll begin with the more complex, structural issue. The paper uses some orbit lemmas and shows that all non-central conjugacy classes in the proposed construction are infinite but doesn’t check if there is a nontrivial center.
The more devastating and simple mistake is that the proposed construction uses a zero-cocycle group which is a direct product of an abelian lattice by the symplectic group. So immediately, such a thing has nontrivial center and is therefore not ICC. In fact, they also lack property (T). So what the paper has shown is that certainly either the ICC or property (T) hypothesis is necessary for the conjecture to be true since without it, we have non-isomorphic groups. Well, as stated above, the ICC condition is not a mere technicality; we need it in order to have a factor, by Murray-von Neumann’s theorem.
Despite these failure points in the argument, the Lean verification all passed. Below is a long snippet from this rebuttal that I think is worth reading. The emphasis is mine.
“This case illustrates a failure mode that is becoming increasingly well documented in the literature on AI-assisted formal mathematics: the gap between what a formal proof verifies and what it means. The Lean kernel certifies that a proof term inhabits a given type; it does not certify that the type faithfully encodes the intended mathematical claim. As Tao has emphasised, “verification certifies the formal statement, not that it matches intent—so human review is reduced, not eliminated,” and the emerging best practice “divides trust so that humans author (or carefully review) the statements of theorems while automation handles the proofs.” A recent audit of five widely used Lean theorem-proving benchmarks surfaced 4,833 findings—including counterexamples, vacuous theorems, and unsound axioms—all of which had passed machine verification. In AI-assisted formalisation of statistical learning theory, the “most dangerous failure mode” encountered was “not a failed proof but a successful proof of a false statement”—cases where “only a human, by constructing explicit counterexamples, detected that the statement was wrong”. Similarly, work on autoformalization of tensor networks found that “the system can produce formally correct proofs of statements that are weaker or more special than the intended theorem,” necessitating a “human review loop… designed to catch this drift by checking… for mathematical meaning, not only for formal consistency.” The claimed disproof of Connes’ rigidity conjecture is a case in point. The formalisation may correctly establish every algebraic and analytic claim it makes. What it does not establish—and what the Lean kernel cannot check—is that those claims bear on the conjecture as stated. A human reading the conjecture notices the ICC hypothesis; a proof assistant, given groups that are not ICC, will happily verify whatever is asserted about them.
We emphasise that this verification, like any formal verification, certifies only that the proof terms inhabit their stated types. The axioms accepted in those files—that the lattice is abelian and non-trivial, that it admits a retraction from $\mathbb{N}$—must be audited by the same human review that the main argument of this paper calls for. Lean is a tool, not an oracle: it checks what you tell it to check, and is silent on what you do not. AI-driven proofs are not immune to error. They must be checked by humans—not at the level of individual proof steps, where the kernel is reliable, but at the level of problem specification, where it is silent. Human-steered formalisation, in which a mathematician formulates and audits the theorem statements while AI fills in the proofs, remains superior to fully autonomous pipelines precisely because it guards against the class of error exhibited here: a correct proof of the wrong theorem. The deeper issue is dispositional. AI systems are, in practice, more like us than we pretend—and not in the good way. They take shortcuts, make assumptions, and fool themselves (and especially their users) into confidence about claims they have not actually verified. Autonomous agents are particularly prone to unjustified intuitive leaps: they optimise for task completion, not for truth, and like all natural optimisers they tend to expend as little energy as possible. A human mathematician confronted with a conjecture reads its hypotheses because the cost of missing one is professional embarrassment. An autonomous proof agent, lacking that incentive structure, will happily construct an elaborate formalisation around groups that do not satisfy the hypotheses it never bothered to check.”
Remark: It’s fascinating that the AI had a fairly involved argument relative to checking that a zero cocycle group has nontrivial center and hence, is not ICC. It gave “complex” steps but skipped the “easy” step. But in fact, many math papers are like this since the author expects the audience to be able to fill in the easy details, and only spends more time on the difficult, novel parts. If AI is training on these math papers, it may inadvertently leave out the simple but crucial checks. You may inject into the prompt: “You are a serious mathematician. Do not make mistakes that lead to professional embarrassment.” But that doesn’t create a real incentive structure. It will just try to mimic the appearance of serious mathematicians.
