↓Skip to main content

AI and the Future of Mathematics: Why Understanding Still Matters

·2641 words·13 mins
AI Mathematics Artificial Intelligence Mathematical Research Machine Learning Math Education AI Research
Table of Contents

AI and the Future of Mathematics: Why Understanding Still Matters

On September 21, OpenAI announced that an internal model whose training began on August 28 had already solved more than 100 challenging mathematical problems spanning a broad range of fields.

The development has triggered familiar questions about the future of mathematical research. If AI systems can increasingly solve problems that once required years of human effort, what happens to mathematicians, mathematical education, academic publishing, and the traditional process of producing proofs?

University of Toronto mathematician Daniel Litt approached the question from a different direction in his essay, “A beginning for mathematics.” The title is particularly notable because Litt had previously given a talk titled “The End of Mathematics” examining the potential impact of AI on mathematical research.

His newer argument does not deny the possibility of extremely capable mathematical AI. Instead, it accepts a more radical premise: AI systems capable of reliably outperforming humans across most mathematical tasks may arrive relatively soon.

The central question then becomes something else:

If machines can increasingly produce mathematical results, what exactly should humans continue to do?

Litt’s answer focuses on a distinction that could become increasingly important in the AI era: producing mathematical results is not necessarily the same thing as understanding mathematics.

🧠 If a Machine Can Prove Theorems, What Are We Actually Pursuing?
#

Litt begins by outlining the rapid progression of mathematical AI.

Only a few years ago, AI systems could make elementary arithmetic mistakes. More recently, systems from OpenAI and DeepMind have demonstrated mathematical abilities comparable to high-level competition performance. The next stage is increasingly autonomous mathematical research, including work on difficult problems that previously required substantial human effort.

This progression forces mathematicians to confront an uncomfortable question: what is the actual objective of mathematics?

There has never been complete agreement.

Some mathematicians see mathematics primarily as a way to solve problems. Others treat it as an intellectual game, an aesthetic discipline, or something closer to poetry. Some pursue the ideal expressed in Hilbert’s famous declaration, “We must know, we shall know.” Others believe mathematics is an exploration of an abstract or Platonic mathematical reality.

There are also mathematicians who place substantial value on teaching, mentorship, and transmitting mathematical culture to the next generation.

Litt reduces his own priorities to two broad objectives:

  1. Produce and understand high-quality mathematics.
  2. Cultivate high-quality mathematicians.

Neither objective has a perfectly fixed definition. The mathematical community itself continuously determines what constitutes important mathematics and what constitutes mathematical expertise.

The problem is that, historically, the same mechanism has been used to pursue both goals: proving theorems.

A mathematical paper typically needs significant theorems. A doctoral dissertation traditionally demonstrates its author’s research ability through original mathematical results. Solving a long-standing open problem can represent one of the highest achievements available to a mathematician.

But Litt argues that this framework may have always measured only part of what mathematics values.

The Monkey and the ZFC Thought Experiment
#

Consider a hypothetical system that receives the axioms of ZFC set theory and mechanically applies inference rules to generate consequences.

Such a system could theoretically produce an enormous number of valid mathematical statements.

The “monkey” in Litt’s thought experiment is deliberately absurd, but the underlying point is important: logical validity alone does not determine mathematical value.

If the only objective were to generate true mathematical statements, automation would be conceptually straightforward. A sufficiently powerful system could enumerate mathematical propositions and attempt to prove them sequentially.

Eventually, it might encounter problems that human mathematicians currently regard as important.

That would not make the system a mathematician in the meaningful human sense.

The interesting question is not simply whether a proposition has been proven. It is why the proposition matters, what it reveals, how it connects to other ideas, and whether the resulting theory changes how humans understand a mathematical domain.

What If the Monkey Becomes Extremely Intelligent?
#

Litt then pushes the thought experiment further.

Imagine that the hypothetical mathematical system can:

  • Understand its own proofs
  • Produce elegant explanations
  • Identify interesting results
  • Solve previously unresolved problems
  • Discover useful generalizations
  • Propose important new questions
  • Develop coherent mathematical theories

At that point, the distinction between “machine that proves theorems” and “mathematician” becomes much harder to maintain.

Yet Litt still argues that humans would have an important role.

The reason is that machine-generated mathematical knowledge and human mathematical understanding are not identical things.

AI may eventually produce answers faster than humans can. That does not mean humans will automatically understand those answers.

This creates a possible mathematical “big bang”: a period in which the amount of new mathematics generated by machines increases dramatically while the human capacity to absorb and understand it remains limited.

The bottleneck may shift from discovering mathematics to understanding mathematics.

πŸ›οΈ What Mathematics Needs to Preserve May Not Be Today’s Institutions
#

Once mathematical production becomes highly automated, another distinction becomes important: the value of mathematics itself is not necessarily identical to the institutional structure surrounding mathematics today.

Litt places particular value on activities such as:

  • Research seminars
  • Informal mathematical conversations
  • Students discussing difficult problems with professors
  • Unexpected connections discovered through collaboration
  • Groups of mathematicians working through questions nobody fully understands

These activities may remain important even if the mechanisms for publishing and distributing mathematical results change substantially.

By contrast, institutions such as journals, peer review, arXiv, and conventional paper authorship are not necessarily immutable.

The Limits of Protecting Existing Institutions
#

The mathematical community may naturally focus on protecting established structures.

Possible concerns include:

  • Preventing AI-generated papers from overwhelming journals
  • Preserving the role of peer review
  • Maintaining the authority of mathematicians as knowledge gatekeepers
  • Determining how AI-generated research should be credited
  • Preventing automated systems from flooding repositories

Litt’s argument is that trying to preserve today’s equilibrium indefinitely may be unrealistic.

If a widely available AI system can produce useful mathematical results at very low marginal cost, the volume of mathematical output could increase dramatically.

The more productive question may therefore be:

Which aspects of mathematical culture are intrinsically valuable, regardless of how mathematical results are produced?

Do We Simply Reward Better Questions?
#

One tempting response would be to change academic evaluation so that mathematicians are rewarded primarily for asking good questions.

But Litt warns against building institutions around AI’s current weaknesses.

AI systems that struggle with mathematical question generation today may become highly capable at generating conjectures and research directions tomorrow.

Academic institutions change relatively slowly compared with AI capabilities.

A standard designed around today’s limitations could therefore become obsolete within months.

Instead, the more durable question is what should remain valuable even if AI eventually becomes capable across almost every conventional mathematical task.

πŸŽ“ What Should a Mathematics PhD Demonstrate in the AI Era?
#

Education may be one of the areas most directly affected by increasingly capable mathematical AI.

Historically, a mathematics PhD dissertation serves at least two purposes.

First, it contributes new mathematical results.

Second, it provides evidence that the candidate has achieved a high level of mathematical expertise.

AI complicates the second function.

A student could potentially use AI to generate substantial portions of a dissertation, including proofs, exposition, and even novel mathematical results. The resulting document could be technically impressive while providing relatively little evidence about how deeply the student understands the material.

This suggests that mathematical progress and individual mathematical capability may need to be evaluated separately.

From Dissertation Production to Mathematical Expertise
#

Litt proposes a different conception of the future mathematics PhD.

Rather than requiring students to demonstrate that they personally produced every research result without AI assistance, the degree could focus more heavily on becoming a genuine expert in an important mathematical problem or area.

The candidate would need to demonstrate that expertise directly.

A future doctoral evaluation could therefore place greater emphasis on:

  • Rigorous oral examinations
  • Continuous questioning by experts
  • Explaining difficult concepts without prepared text
  • Handling unfamiliar examples
  • Applying methods to new situations
  • Defending assumptions and conclusions
  • Explaining why a result matters
  • Demonstrating understanding beyond the published proof

A dissertation could remain part of the process, but the document itself would become less reliable as a proxy for personal mathematical understanding.

The principle is straightforward:

Students can use AI extensively, but they still have to understand the mathematics themselves.

Why Oral Mathematical Communication May Become More Valuable
#

As AI makes polished written mathematical content easier to generate, direct mathematical communication may become more informative.

A generated paper can look convincing.

A sustained technical conversation is harder to fake.

An expert can ask why a particular assumption is necessary, request a proof in a different form, introduce an unfamiliar counterexample, or ask the student to generalize an argument.

The ability to respond coherently under these conditions provides evidence of internalized mathematical understanding that a static document cannot necessarily provide.

In this sense, AI may not eliminate traditional forms of mathematical evaluation. It could make some of them more important.

πŸ” The Underappreciated Skill of Choosing Good Problems
#

Another potential shift concerns one of the least explicitly rewarded activities in mathematics: deciding which questions are worth asking.

AI systems can already work on difficult open problems.

But those problems did not appear from nowhere.

Human mathematicians spent years identifying conjectures, evaluating their significance, developing partial results, and determining that certain questions were worth sustained attention.

By the time a problem becomes a recognized open problem, it often carries a substantial amount of accumulated mathematical judgment.

The research community has already performed an important filtering process.

From Solving Problems to Building Research Programs
#

If AI becomes capable of solving many existing problems, the ability to construct meaningful research programs could become more important.

That might involve:

  • Identifying a collection of related questions
  • Determining which questions are foundational
  • Finding promising directions
  • Connecting apparently unrelated mathematical areas
  • Developing long-term research programs
  • Convincing other researchers that a direction is worth pursuing

This does not necessarily mean humans will permanently retain this ability.

AI will likely improve at generating conjectures, proposing questions, and constructing research plans as well.

The result could be an enormous expansion in the number of candidate problems.

That brings the discussion back to the fundamental bottleneck: attention and understanding.

πŸ“š When AI Produces Too Many PDFs, Understanding Becomes the Bottleneck
#

Litt takes a deliberately permissive position toward AI-generated mathematics.

Trying to prevent people from using AI to produce mathematics may not be practical.

There may be little value in artificially reserving certain problems for humans or attempting to prohibit researchers from using automated systems.

The more important question is what happens when mathematical production becomes extremely cheap.

If AI can generate thousands or millions of mathematical results, the limiting resource will no longer necessarily be the ability to produce new mathematics.

It will be the ability to determine which results matter and understand them deeply.

If a Conjecture Falls in the Forest
#

Litt uses an intuitive formulation of this problem: if a mathematical conjecture is generated somewhere but nobody understands or cares about it, what mathematical value does it ultimately have?

Imagine an AI system producing a beautiful conjecture and proving it automatically.

If no human understands the proof, knows why the conjecture is important, or can connect it to existing mathematics, the result may have limited practical significance within the human mathematical community.

This does not mean machine-generated mathematics is worthless.

Rather, it means that mathematical understanding remains a separate scarce resource.

Even applied results require people who understand the assumptions behind them, know where the conclusions apply, and can identify situations in which the underlying reasoning fails.

Increasing mathematical productivity therefore does not eliminate the cost of understanding.

It makes that cost more visible.

πŸ’° When Mathematics Becomes “Pay-to-Win”
#

Mathematics has traditionally been unusually inexpensive compared with experimental sciences.

A mathematician generally does not need a particle accelerator, biological laboratory, or expensive physical apparatus to make progress. A notebook, computer, or blackboard can be enough.

AI changes that economic model.

Some mathematical problems may increasingly become solvable by purchasing computation.

More compute, more inference, and more repeated searches could increase the probability of obtaining a solution.

This creates a new relationship between mathematical progress and financial resources.

The Loss of Mathematical Detours
#

There is a potential downside.

Historically, a difficult mathematical problem could force researchers to develop auxiliary theories, invent new techniques, and explore unexpected connections.

Sometimes those detours became more valuable than the original solution.

If an AI system simply produces the answer immediately, some of those exploratory paths may disappear.

That could reduce opportunities for humans to encounter unexpected mathematics while struggling with difficult problems.

Litt acknowledges this possibility but does not treat it as a reason to reject AI-generated solutions.

If a fundamental problem can be solved for the cost of an ordinary dinner, obtaining the solution is still valuable.

The important part is what happens afterward.

Mathematicians can ask:

  • Why does the result hold?
  • What does it explain?
  • Which assumptions are essential?
  • Can the result be generalized?
  • What other mathematical structures exhibit the same behavior?
  • What new questions follow from the result?

Some of these questions may themselves be solved by AI.

Others may generate entirely new areas of confusion.

And that confusion can become the starting point for new mathematics.

🧩 The Final Step From Explanation to Understanding
#

Litt ultimately returns to a familiar mathematical scene.

A student encounters a difficult problem, cannot solve it, and goes to a professor’s office.

They may use AI during the discussion. The model may provide an elegant proof or explanation. The professor and student may inspect it together at a blackboard.

But the interaction still contains a step that cannot simply be outsourced.

There is a difference between seeing an explanation and understanding it yourself.

A model can provide the proof.

It can explain the proof.

It can generate examples.

It can answer follow-up questions.

But the student still has to build the internal representation that makes the mathematics meaningful.

That process involves judgment, intuition, abstraction, and the ability to connect one mathematical idea to another.

Even extremely capable AI does not eliminate the need for a human to perform that act of understanding.

🌌 Mathematics May Be Entering a Beginning, Not an End
#

Litt’s argument ultimately accepts some of the most disruptive possibilities associated with mathematical AI.

AI may become better than humans at solving many mathematical problems.

It may produce enormous quantities of new mathematical knowledge.

Many research papers could become easier to generate.

Traditional doctoral training may need substantial revision.

Existing publishing and evaluation systems may become increasingly difficult to maintain in their current form.

None of this necessarily means mathematics itself is ending.

Instead, the scarce resource may shift.

When producing a proof is difficult, mathematical progress is constrained by the ability to produce proofs.

When producing proofs becomes cheap, progress may be constrained by the ability to understand them.

When generating conjectures becomes cheap, choosing meaningful conjectures becomes more important.

When producing papers becomes cheap, identifying valuable mathematics becomes more important.

And when answers become abundant, asking what those answers actually mean becomes increasingly central.

This suggests a different interpretation of the AI transition.

The future of mathematics may not be defined by humans competing with machines to prove theorems faster. It may instead involve humans and machines operating at different layers of mathematical work: machines expanding the space of possible results while humans continue to decide what deserves attention, what deserves understanding, and what new questions should follow.

The title of Litt’s essay therefore captures the broader possibility.

If AI makes mathematical production dramatically cheaper, mathematics does not necessarily become less important.

The field could become larger, stranger, and more difficult to comprehend.

And if the central objective remains not merely to generate mathematical truths but to understand them, there may still be an enormous amount left for humans to discover.

Related

DeepMind’s Four Paths to ASI: Scaling, Agents, and Self-Improvement
·674 words·4 mins
DeepMind AGI ASI Artificial Intelligence Machine Learning AI Safety Recursive Self-Improvement Multi-Agent Systems LLMs AI Research
Claude Optimizes 30+ Scientific Models as ScienceIDE Scales AI
·2632 words·13 mins
AI Scientific Computing Claude ScienceIDE AI Agents GPU Optimization HPC Machine Learning
AI Self-Improvement by 2028? Inside Anthropic’s Bold Prediction
·767 words·4 mins
AI Machine Learning Anthropic Automation Software Engineering AI Benchmarks LLMs Future of AI