Mathematical Research and Artificial Intelligence II

9 minute read

Published:

Since writing this blog post, the speed at which events are happening with AI and mathematics has been breakneck. The biggest recent announcement was from OpenAI, claiming a solution to the Navier-Stokes equations which exhibits a singularity in finite time. Of course, this problem is well known as a Millennium Problem and there are rumors of other Millennium problems being solved. It seems very likely that they are being chosen in order to obtain prestige/clout. The day before the announcment, Tristan Buckmaster made a statement saying that he had been working on the problem and other similar ones in fluid dynamics and suspects that OpenAI used ideas/prompts from his private ChatGPT sessions and fed them to their non-public, latest model. The history is roughly as follows: Buckmaster had been collaborating with Levent Alpöge (mathematician at Anthropic) for about a year and published research on a related subclass of the problem in August. OpenAI’s head of mathematics, Sebastien Bubeck, allegedly reached out to Buckmaster privately and offered to share credit for the breakthrough on the condition that Alpöge was excluded. Buckmaster believes that the approach he and Alpöge were trying is unique; no one else was trying it at the time. Buckmaster dutifully attributed the origins of their approach to some other mathematicians, unlike OpenAI who did not give credit and is possibly committing plagarism and intellectual property theft. Moreover, when Buckmaster did not agree with the conditions and asked about them using his private user data, things took an ugly turn. Bubeck stopped “playing nice”, threatening to ruin Buckmaster’s career if he went public with these claims.

An additional important point: at first, OpenAI claimed that their model had pretty much been left to its own devices after the first prompt. But later, they admitted they had a whole team of people working on Navier-Stokes specifically and the number of tokens they used was astronomical, enough energy to power a small city. Buckmaster called this a Deep Blue vs Kasparov moment which I find very fitting. There’s the obvious man vs machine metaphor but also that IBM allowed Deep Blue to analyze/train on Kasparov’s chess games but Kasparov had nothing to study on Deep Blue since it wasn’t a human player with any history of games. And of course, there was a whole IBM team working on Deep Blue.

It’s probably a minority view but I did see one mathematician say that this potential plagarism is the least interesting thing about the story to him. He wants to focus on the power of these latest models and that perhaps soon, the most flashy, popular, publicized math problems will be solved by AI and people can just work on their unflashy, unpopular math problems in peace. Assuming this is a genuine opinion and not a troll, there is a misunderstanding of the situation.

  1. If Buckmaster is right and the AI needed human guidance and lots of human idea/prompts, then it is not as powerful as this opinion puts forth.
  2. A scholar should always be concerned with plagiarism.
  3. The Clay Mathematics Institute chose the Millennium Problems for their importance. Obviously, lots of people can have their preference about what kind of math they like but I don’t think many people are able to comment on the importance of a problem if it’s not their field. Yes, there is definitely politics involved with which research programs get attention and funding and this person may be jaded. But that doesn’t make him qualified to judge other fields of math.
  4. I’m not convinced that AI isn’t coming for the unflashy, unpopular math problems. An undergrad with little knowledge could try to tackle some research problem just to get a bit of attention.

Indeed, a Mathathon is to be held in the (academic) fall 2026 where undergrads are given 2 million AI credits to prompt LLMs to solve math problems. An open letter has been written in response to this idea. I will quote it below at length because of it’s importance:

This event is likely to have destructive impacts for the mathematical community. 

AI companies see research mathematics as an advertising opportunity. Over the past few months, they have advanced a campaign of scientific misinformation about the goals of mathematical research. Major AI companies are engaged in an arms race to declare themselves the first to prove prominent open conjectures. New announcements of AI-generated results appear on social media on a weekly basis. These results are often poorly communicated, and have diffuse negative impacts on the careers of human mathematicians who are concurrently proving the same results. 

Mathematicians do not only prove theorems: they also take measures to ensure that the future of a field remains healthy after the extraction of a result. They do so by sharing their results and explaining their methods to others at talks and conferences. They mentor students and postdocs so that the next generation can advance research further, and take care not to scoop other mathematicians at the last mile. In announcing new LLM-generated theorems, AI companies have not followed such research practices. The intent is clearly to extract the prestige of newly proven results rather than add value to the community in a way that safeguards its future. 

These companies appear guided by an unhealthy instinct to claim certain results before competitors at all costs, regardless of the collateral damage to mathematical understanding. After the latest result is dropped, research mathematicians are compelled to step in to properly verify, disseminate, and sometimes discredit entirely the claimed results. This labor goes uncompensated, uncredited, and unacknowledged. This benchmark-focused arms race shortcuts the standard peer review process, which necessitates time and effort, while extracting unpaid labor from research mathematicians. To put it bluntly, AI companies are engaging in research misconduct.

I think the text explains the issues very well. Mathematicians are not simply in the business of churning out theorems. If that were the case, perhaps mathematicians should all quit their academic jobs to try and work at a big AI company, or pay the hefty AI subscription fees. No, the purpose of math is to foster human understanding; there’s a whole sociological factor to how math research/education is done. The arms race between AI companies is, instead, about profit and prestige, much like the intellectual arms race we see from the Cold War era between the US and USSR, played out in the space race, chess, and technology. AI has not been used in a respectful way of the mathematical community, of the work/lives of mathematicians.

So it is clear there is a severe misalignment between AI and math which is the content of another open letter. I’ll also quote this letter at length because I think it’s important.

In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.

Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.

We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.

My only point to add for now is this: so far, it still seems AI needs human guidance for knowing what problems to solve even if it doesn’t need help solving them. If human experts are out of the job and we end up in a time where there are no human experts, how will we direct AI to do innovate? Or if AI does come up with research directions, how would we be able to judge if those are good directions to go?