Mathematical Research and Artificial Intelligence III
Published:
Once again, the pace at which AI is disrupting mathematics is astounding. OpenAI announced over 700 significant mathematics results which can be found in this repo, many of which have Lean verification. Moreover, they have also said that they will make their frontier models available to “eligible institutions” which I read as: “institutions which can pay or will add to OpenAI’s prestige.” I asked a mathematician friend of mine how mathematicians are reacting and he said, “Absolute despair.” Some of the people had every single one of their research problems scooped by this drop (assuming the papers are actually valid and correct). It’s also overwhelming to try to process, from an intellectual viewpoint as well as a career and what-is-the-future-of-math perspective. On a personal note, there are many annoucements regarding conjectures in my field of symplectic geometry, including a purported counterexample to the nearby Lagrangian conjecture.
As I’ve said before in other posts, the mathematics community cannot afford to spend all their time reviewing AI generated results. They are not paid for it and AI often writes in such a way that each sentence seems to vaguely make sense but often lack a unifying coherence and flow that would otherwise, make it an actually meaningful paper.
Making the frontier models available (possibly at a high price) to a select population increases inequity. Sure, in some sense, it’s always been the case that well-funded universities with their own labs and leading experts have the advantage over small liberal arts colleges, for example. And some people even like that, claiming that it separates out the talent. But the supposed promise of AI was to be good for everyone, that everyone can now learn about their interests without high tuition, everyone can try their hand at coding, etc. When it’s given to just certain people, that would be like giving steroids to some athletes and saying that the competition is simply a venue for finding the real talent. I suspect that soon, different AI companies will make their models available to different institutions and math will be even more a battle ground for these AI companies. Mathematicians and institutions become the pawns.
Why choose math as the place for battle and advertising? The AI models, so far, have no physical manifestation and cannot conduct physical lab research or interact with the world in that capacity (as far as I know). Math requires the least amount of physical interaction with the world. Moreover, math can be very quantitative and there are more “objective” ways to measure an attempt at proving a mathematical result. Personally, I do wish these AI models would be applied to something more useful for society like curing diseases, rather than showing the rational Hodge conjecture is true for CM abelian varieties.
Another observation: because of how much time it takes to review cutting-edge research, it may be that a lot of time passes before these papers are verified to be correct or incorrect. In the meantime, since the public is not well-informed about math or its value and also have no qualifications for evaluating the papers, they will likely assume the papers are correct and take it as a signal of a huge leap forward for math and AI. The public does not typically think of math as an art or creative process. If they did, they might have more sympathy as they might for artists (especially ones working in digital mediums) losing their livelihoods to AI generated art. And in fact, the papers don’t even have to be correct. They just need to convince the public of their correctness, that AI can solve the hardest math problems, and thus, get the buy-in of investors for AI companies to profit.
One might say: “But the papers have been Lean verified!” The thing is, we all know just how easy it is for things to be lost in translation, for meaning to be ambiguious, within the same human language, between different human languages. The sentence, “I wrote my paper on marajuana” could mean I was high when I wrote the paper or that the subject of the paper was on cannabis or even, I took a pen and inscribed my paper on cannabis leaves. The sentence “I saw her duck” could mean I saw her bend her knees and head down or her pet waterfowl. Or that I took a hacksaw and committed animal abuse.
Math should be more precise than typical human conversations but it is naive to think we can escape the need for human clarification and further communication and discussion. Mathematicians all know of math papers that are well-written and how hard it is to write well. And they also know of papers that are poorly written and how frustrating it is to read them.
So to say that the papers are Lean verified and hence, imply they are trustworthy, is to underestimate the difficulty and nuance of translation and formalization. There is a paper on Lean formalization whose title is “Navier-Stokes Lost in Translation: why Lean verification of AI autoformalisation does not guarantee correct natural language proofs.” The authors Bastounis, Circelli, and Hansen write in their abstract:
Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI’s announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI = $\infty$). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI = 1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean ‘verifications’. These include OpenAI’s announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
