This post has also been submitted to Proofs and Prompts.
In the last half year there has been a massive shift in the "Math and AI" discourse, primarily driven by changes in AI capabilities. In almost no time at all the collective opinion went from "AI can't do research level mathematics" to "AI can solve many small or medium sized research problems and some hard or long-standing problems", and for some of us, even to "AI will likely solve all mathematics problems before a human in the near future". It's been wonderful to see the community so quickly take on the challenge of discussing the future of our discipline both online and at conferences.
However, we need to immediately begin updating the incentive structures within mathematics. What would this mean in practice? Organizations should (as soon as possible) determine the types of mathematical activities they wish to encourage in this post-AI world, then clearly communicate these new expectations and reward work that reflects them through jobs, publications, awards, and other forms of recognition. For example:
Note that I am purposefully non-specific about what these activities and contributions should be: each organization should decide for itself what they think will be important moving forward. Discipline-wide consensus will naturally emerge and evolve as more and more organizations implement changes and then update them in light of new developments. While I do have strong opinions on what these changes should be, which I lay out below, I think the situation is too urgent to debate our way to broad consensus before taking action.
This has already begun to a limited degree. Some journals and conferences are cracking down on AI writing or even clarity of writing in general. New submission quotas are also being implemented (see here and here). However, I believe the changes need to be more widespread and drastic.
The easiest signal to point to is the number of monthly arXiv submissions in mathematics. In the most affected fields[1], this has almost doubled since 2025 and the trend doesn't seem to be slowing down; see for example the recent trend in combinatorics submissions.

Even the "less affected"[1] fields like algebraic geometry have seen a very significant change, and the trend in the last month or two seems to be following in combinatorics' footsteps.

If we take arXiv submissions as a proxy for submissions for publication, this points to a publication crisis in the very near future. While combinatorics and metric geometry have so far been the most impacted,[1] we should expect to see similar trends in other fields in the near future as AI capabilities continue to develop.
For example, in queueing and scheduling (my field), it seemed like the general opinion at the 2026 SIGMETRICS conference in early June was that AI could not yet solve interesting problems in our field. At the time, this was an opinion I mostly shared. However, since the releases of Fable 5 and GPT-5.6 Sol I no longer believe this is the case, and I expect to see a wave of AI solved problems in queueing over the next several months. My opinion primarily changed due to:
These two models marked a (significant) shift in AI capabilities for many fields. Fields in which AI still struggles should expect this to change in the near future.
In summary: we have seen a rapid and somewhat concerning change in the number of mathematics papers written per month and this trend seems likely to continue. I don't know whether these numbers are due to "slop papers" or just due to the increased speed at which problems can now be solved. I hope it's the latter but expect it's at least partially the former.
However, in some ways it doesn't matter. If the rate at which papers are written continues to grow, the systems of publication and review will not be able to handle it (not to mention that it may become impossible to keep up with your field in a meaningful way).
Under the status quo, we reward solving and then publishing solutions to significant, interesting, or long standing problems, but only if you are first. If you are purely optimizing for this success metric, you should use AI to solve as many significant problems as possible as quickly as possible. Furthermore, you should not care about the quality of your writing beyond some minimum threshold, since writing well is a significant slowdown and minimally rewarded compared to solving more problems.
The problem is that if anyone does this, then everyone must if they wish to remain competitive (for jobs, etc) and not regularly get scooped. This forces us into an obviously terrible prisoner's dilemma type situation that can only be resolved by changing the incentives. This issue is addressed more completely and eloquently in this short essay by Ruodu Wang, which I highly encourage everyone to read.
Here's one example of the craziness caused by this situation that I recently experienced.
I recently used AI to solve a significant open problem in queueing theory. After getting the initial proof and verifying its correctness, I spent three weeks iteratively refining the proof with my advisor until it felt like we had found the “right” idea: one that made the argument flow naturally and clearly distinguished the key new ideas.
Partway into this process I started getting anxious about getting scooped. At any moment someone might ask AI to solve the same problem and then post (or even worse, tweet) an AI slop writeup of the proof, potentially without even verifying it. I didn't feel like I deserved credit for solving this problem, after all, that was done entirely by AI. But it would feel extremely disappointing to get no credit for the sleepless weeks I had spent verifying, digesting, refining, and then writing up the result. Because of this anxiety, I (a) started sleeping less in an effort to get it done as soon as possible and (b) eventually tweeted a hash of the proof we had at the time.
I would have chalked this up to my own idiosyncrasies (and maybe reading too much math twitter), but just the other day one of my more measured friends admitted to me that he was experiencing something similar. He discovered a proof for a significant result in metric geometry using AI (with no guidance or ideas from him) and had been spending the last few days verifying and rewriting it. He told me he has been having trouble sleeping as he feels like he needs to get this done ASAP or risk someone else happening to ask AI the same question and posting something about it. I suggested tweeting a hash and he is considering it.
Now let me say this: tweeting a hash of your work is obviously a crazy thing to do. But somehow we have reached the point where it seems like the best of several bad options:
We need to change the system that is creating these situations.
The rest of this post is my own opinions on what activities we should incentivize and why. While I hope that my opinions influence some of you, or at least spark discussion, I will be happy to see any carefully considered changes even if they differ greatly from my proposals.
To start, I hope that we can agree that an AI written paper with a result proven entirely by AI and not carefully verified by a human provides close to zero value. I actually believe something even stronger: such a paper provides close to zero value even if it is carefully verified. Verification is an important and necessary step, but there is no real way to check that an author did it and each reviewer basically has to do this again themselves regardless. Furthermore, I expect verifying correctness will soon become a non-issue through a combination of better AI and Lean verification. Thus, while authors definitely should verify results before submitting them, we should view verification less as a value-add and more as a basic requirement.
Okay so AI written, AI proven papers provide little value. Why is "AI written" specifically a problem? Well, it's not inherently a problem, there could be good AI written papers in theory. But currently, the vast majority of AI written papers are poorly written in a way that makes it hard to understand the key ideas of the proof and verify correctness. Put another way, these papers do not meaningfully support human understanding of the ideas and proof. So the real claim is that any paper containing AI proven results, regardless of who or what writes it, only provides value insofar as it assists humans in understanding the important new ideas, gaining intuition, and verifying correctness.
Note that the significance, beauty, and usefulness of the ideas and results remain important, just as they've always been. It is simply that these qualities are no longer enough on their own. Since the ideas can be easily generated by anyone that thinks to ask an AI, the paper must provide value beyond simply recording them. It can do so by providing an exceptionally clear presentation of the ideas.
Clear presentation does not only mean clear exposition, though that is one important element. Many important results were first proved using long or complicated arguments that were then refined over time as mathematicians found the "right" definitions, objects, perspectives, and techniques that made the proofs shorter, simpler, and more natural. Such refinements are important mathematical contributions even when they do not establish new results. Instead, they help humans better understand existing results and why they are true.
My own experience using AI to solve a long-standing open problem suggests that producing this kind of "proof refinement" may become one of the main mathematical activities left to humans (at least for now). Although the AI easily produced an initial proof, the argument was difficult to understand and did not introduce any new objects or definitions that clarified its underlying structure. Finding the right definitions and objects and then reorganizing the proof around them still required substantial work.
I suspect that this is because such proof refinement is closer to theory building than to problem solving, and AI still seems to struggle with theory building. This gives me some hope that there will remain interesting and enjoyable mathematical work for humans to do even when AI can solve the problems themselves. At least in this case, I found it quite engaging and fun to try to find the cleanest and most natural version of the proof that I could, and doing so required non-trivial insight into the problem.
In his recent ICM talk and paper, Terence Tao describes "canonicalization":
This process of canonicalization — in which a result is restated in its natural generality, given its right proof rather than its first proof, connected to its neighbors, and absorbed into the standard toolkit — is the slowest stage of all. It requires broad, deliberative consensus, and it is the stage least amenable to optimization by AI tools. It is also, in my view, the most valuable part of the entire process.
It is impossible to take a result from new to "canonicalized" in one step, regardless of how much effort you put in. As Tao says, this process inherently requires consensus and absorption within the field. What I describe as "proof refinement" is taking the biggest first step you can towards the canonical proof.
Let me now summarize all of the above and state clearly what I believe we should reward. For any AI proven result, authors should provide value by advancing our collective understanding of both the result and the techniques used to prove it. They can do so by:
Also, when a result is proven by an AI, priority should carry no reward. If an initial publication does not provide the value described above, it should not diminish the credit given to a later author who does provide such value.
There is one last very important wrinkle to all this. Everything I have said so far has been with regard to results proven by AI. However, we cannot reward AI proven results and human proven results differently. If we did, it would create an extremely strong incentive to lie about AI usage. Since there is no real way to check if an AI was used to prove a result and we do not want to reward scientific dishonesty (and, in effect, punish honesty), we must therefore evaluate all results published from here on out as if they were proven using AI.[3] Thus, we must apply the new standards (the ones described above) to all papers, not only those with AI attribution.
So finally, the changes I would personally like to see are:
My primary research community is ACM SIGMETRICS, so I want to appeal to its leadership directly: please update the review guidelines as soon as possible to reflect the new realities of AI assisted research and then communicate these changes to the community. We need to change the incentives to move towards a healthy and sustainable research culture. While I would be very happy if these changes reflected some of my suggestions, I would be equally glad to see SIGMETRICS pursue any changes its leadership believes would best serve the community.
For everyone else: if you think this is important, reach out to the leadership in your own research communities and encourage them to make such changes as soon as possible. It will require many such changes before new norms and incentives become established across mathematics.
The figure below shows, for each arXiv mathematics category, the percentage change from the average monthly submission count over a full calendar year to the average monthly count from March through August of the following year. Categories are ordered by their 2026 growth rates and the labels inside the green bars give the average number of papers submitted per month from March through August 2026. Note that combinatorics and metric geometry are at the top, while algebraic geometry is near the bottom.
Mathematical Discourse is an interesting new venue for such talks! ↩︎
Even if we all agree to this, I expect we will still subconsciously reward non-AI results more. One radical proposal for dealing with this is to ban AI attribution (at the very least for review, but maybe in general), forcing everyone to assume every result was probably proven by an AI. ↩︎