Wednesday, September 16, 2026

A Marvelous Demonstration of This Proposition Which This Context Window is Too Narrow To Contain

My friend and former colleague Ash Jogalekar tweeted out an interesting result recently, which is now expanded into a preprint. Using an intelligent agent, he had generated a mathematical proof and then applied it to an interesting problem in nucleic acid structure design.  Ash also freely confessed that he lacked the mathematical background to evaluate the proof himself, which means it is utterly hopeless for me to dream of comprehending the proof since he has far more mathematics understanding than I do.  That touched a very sore spot in current mathematical academic discourse: how to deal with AI-generated proofs
The bigger dustup these days in this area is over OpenAI using $15M of GPU time to tackle a previously open question with the Navier-Stokes equations for fluid mechanics.  In addition to OpenAI solving the problem with what some mathematicians view as gauche brute force, there are allegations that OpenAI may have benefitted from academic mathematicians using OpenAI's models, that OpenAI may have basically threatened academics, and various other bad behavior.  

There's also perhaps the most famous mathematician of the moment, Terence Tao, openly questioning whether these AI-driven proofs are eliminating the intermediate steps of typical proof generation, and that those waypoints are often seeds of new productive directions.

I'll freely confess I can't prove much.  Most of the proofs I ever had to construct were in a particularly disliked geometry class with a particularly disagreeable teacher.  I do remember at a very young age trying to generate counter-proofs of the four color map theorem after my eldest brother introduced me to the idea; it's quite possible it was on his mind from being in Science News or Scientific American (we subscribed to both) shortly after the proof was announced.  Later I'd go down the route of many cranks and try to generate perpetual motion machines or ways to beat the speed of light.  Over time, I've succumbed to believing proofs worked by my mathematical betters.  

I suspect this is one phenomenon that already had mathematicians on edge - the ubiquity of crank "proofs" for various topics.  These have long been the bane of those skilled in math or physics.  I had a now-deceased relative (by marriage!!) who had announced he was going to spend his retirement (he had been a food chemist) proving that quantum mechanics was pure bunkum.  As someone who works frequently around data that could not have been generated except for quantum effects, I wasn't exactly cheering him on.

Back to Ash.  What he announced was two fold: that he had used AI to generate a mathematical proof he didn't understand, but also that he had used the method under the proof to generate an interesting result in nucleic acid design.  So he submitted a preprint to ArXiv.  Which kicked a hornet's nest on Twitter.

Now, of course Twitter is not a representative sample.  It isn't even clear all of these posters are professional mathematicians, though some clearly are.  But the vitriol was impressive.  And usually there were pile-ons later on.    I link to a bunch at the end, but you can see some words I usually keep out of this space. as well as "doofus", "vandalism", "quisling", "should die of shame", "ban him", and other invective.  These aren't the complete set - I left some out that use slurs I avoid.


That's a common theme - if you don't understand, you shouldn't preprint.  And it's not just random Twitter outrage artists stating that - here is a section head at arXiv basically deploring this nature of preprint (though not specifically mentioning Ash) and even suggesting that perhaps authors not famous in a field must pass an oral examination on your preprint or it is not going to make arXiv. 

Which to me is a terrible gatekeeping idea and, along with the condemnations, very wrongheaded.  These proofs - especially when as Ash's does they have real world applications - have added to human knowledge.  And what is a better way of soliciting interest from experts than to post a preprint?

Computer-assisted proofs date to at least the four color theorem I mentioned, which was published only after a computer program generated and proved an exhaustive list of sub-cases.  I believe that was controversial at the time; do mathematicians still regard the four color theorem with a look of disgust?

There's also the question of what it means to understand a proof.  Again, I'm not a mathematician so I'm not sure whether this means "can reiterate the logic of the steps" or "can expound at length on the deep theoretical implications of every step of the proof".   Fermat's Last Theorem, which my title references, was proved by Andrew Wiles in 1993 using some very abstruse math - what fraction of professional mathematicians understand it?  As an aside, a computational verification of Wiles' proof was just announced.

I wonder also if this reflects a fundamental difference in attitude between scientists and mathematicians.  Scientists expect everything they know to be contingent and tentative.  There are certainly numerous things I was sincerely taught in undergraduate courses that has not withstood the test of time.  The human genome does not have 50K protein coding genes and alternative splicing isn't a minor quirk restricted to a few genes.  Math has the luxury of being able to divide their world into "that which has been proven" and "that which hasn't yet been proven - but might never be".  We expect that a publication hasn't covered every case and every interesting implication of the biology it explored - that's just never possible!

The contretemps is also reminiscent of the battles in the sequence world between data generators and data analysts.  Those who spend resources to acquire biological samples and generate sequences from them bristle at being scooped by armchair analysts who dig through their data - conversely computationalists chafe at the indeterminate time - often appearing infinite - which data generators claim ownership but do not publish.  Luckily for everyone, the Bermuda principles from the Human Genome Project calling for immediate sequence release, while not universally observed, have been hugely influential.  All those sequences fed into multiple alignments which were key parts of AlphaFold attacking the protein folded problem (yes, I mean folded not folding).  Data abundance - both sequences and structures - here clearly led to a huge step function advance in biological science.

There's also often a debate around "minimum publishable units" - what is the least science that should be packaged as a publication?  Much of that debate resolves around using publication numbers as metrics for scientific success.  Worrying about MPUs leads to some of the ginormous intellectual conglomerations that pass as Cell/Science/Nature papers.  Related to this is the practice of packing supplementary data with what are really independent papers.

Personally I'd like to see frequent release of small semi-contained preprints, but that could be seen as hypocritical since I've spent my life in corporations that publish much too infrequently.  But if the goal is to maximize available public scientific discourse, more sooner will always be the answer.

But that gets back to some of the other concerns from those worrying first about the flood of  scientific preprints, and now the surge in AI-assisted results.  How can any human keep up with this?  And few making that argument are going to be happy with what seems like the obvious answer: the only way to deal with it is with AI-assisted mining of the growing scientific corpus.




No comments:

Post a Comment