Mathematics has developed a scoreboard problem.
On August 1, OpenAI announced that an internal version of its forthcoming Astra model had produced ten new results across high-dimensional geometry, coding theory, group theory, operator algebras, complexity theory, quantum information, lattice problems, and extremal combinatorics. OpenAI says the model generated the mathematical arguments, humans prepared the manuscripts with it, and the model then formalized the arguments as Lean certificates. The token cost required to find the ten solutions, priced at current Sol API rates, was roughly two thousand dollars.
That is an extraordinary technical event. It also arrived in a field that already knows how to turn extraordinary technical events into races.
Then mathematicians began reading the paper.
Scientific American reported complaints that some of the most impressive results depended on recent human work that the initial public framing did not adequately acknowledge. One mathematician accused OpenAI of research misconduct over a sphere-packing argument. In the group-theory case, researchers found a genuine new result assembled from recent ingredients that made the field less stagnant than the original announcement had suggested.
OpenAI defended the work and said it was holding itself to ordinary mathematical standards. Its current release explicitly says that questions about artificial intelligence in mathematics cannot be answered by a technology company alone.
Two days later, New Scientist put Terence Tao into the middle of the same argument.
Tao's intervention is useful because he turns past the obvious question.
Artificial intelligence is getting better at mathematics.
That is suddenly the harder problem.
◆
The weakest response to the Astra results is also the easiest one:
Welcome to mathematics.
Mathematicians inherit definitions, lemmas, methods, failed approaches, standard constructions, tricks, examples, counterexamples, notation, conjectures, and entire theories from other mathematicians. A new proof does not have to arrive from a spotless dimension containing no previous work. Mathematical creativity often consists in seeing that two structures already present in the field can be made to answer one another in a way nobody had completed before.
The group-theory result is a clean example. Astra's construction drew on ideas from earlier papers. That does not by itself make the resulting theorem fake. Andreas Thom, one of the mathematicians whose earlier work enters that history, described the result as creative and elementary. The interesting act may have been the synthesis.
That is still an act.
Field Instruments: The Mathematics already gave mathematics its strongest possible status inside Modal Path Ethics. Mathematics is one of humanity's most powerful instruments for contacting stable relation. Once the field has been cut into variables, units, operations, and admissible relations, mathematical rigor can carry an argument much farther than intuition alone ever could.
A machine that can operate seriously inside that closure has entered a serious field. So give Astra the point.
If the proof is new, correct, and independently checkable, then something new has entered mathematical reachability. The theorem does not become less true because a model found it. A proof does not acquire a human essence through the sweat traditionally required to produce it.
The structure either survives the correction regime or it does not.
This is where some anti-artificial-intelligence rhetoric breaks contact with mathematics itself. If the model genuinely found a proof, insisting that it did not really count because the path looked unlike human creativity places psychological ceremony above the result.
Modal Path Ethics has no reason to protect that ceremony.
It does have a reason to protect the field around the proof.
That is where the argument changes.
◆
This is the part of the controversy that survives every argument about machine creativity.
Scientific American reports that Steven Miller objected to the sphere-packing result because a key argument had appeared in his earlier work with a collaborator. The same report says the non-sofic group construction combined ideas from papers in 2016 and 2019, while OpenAI's first public framing made the area sound more dormant than specialists believed it was. OpenAI has since updated its language and says it plans ordinary revisions to the paper.
The accusation should remain an accusation while the relevant mathematicians and publication processes do their work. The structural point does not depend on deciding the misconduct charge here.
Citation is path memory.
A citation tells the field that this route did not begin at the current paper. It preserves the bridges that made the new move reachable. It tells later mathematicians where an instrument entered, which prior theorem carries load, which problem was already moving, and which neighboring path may still contain useful structure.
Credit matters. Careers, jobs, grants, reputation, and intellectual ownership all run through attribution.
There is also a deeper epistemic function.
A field that receives results without their path becomes easier to misunderstand.
The Invisible Board made this point through Go on the same day this article was written. A visible position is historical. The stones in front of the player are compressed evidence of earlier moves. The position cannot be understood completely as a fresh inventory detached from the route that produced it.
Mathematics has the same problem at a different scale.
A proof sits inside a literature. The definitions have ancestors. The lemmas have routes. The obstruction somebody finally bypasses may be visible only because ten earlier people spent years discovering where the wall actually was.
Delete the path and the theorem remains true.
But the mathematical field becomes thinner around it.
This is why attribution is larger than etiquette. Proper attribution keeps the discovery connected to the structure that made it possible.
The machine may discover the theorem and still fail to discover where the theorem came from.
That failure becomes more dangerous as generation accelerates.
◆
There is a good case for competition in mathematics.
Mathematical history contains prizes, priority races, rival schools, Olympiads, public challenges, departmental competition, grant competition, journal competition, and the quieter competition inside a room where two people have different ideas about how a proof should work and each would enjoy being correct.
Competition can apply pressure.
Pressure can expose weak structure.
The Invisible Board makes the same argument for games: the opponent is part of the instrument. A private theory can protect itself. Another player has an active interest in locating exactly where your reading of the board fails. The shared rules let disagreement become consequence.
Mathematics has its own adversarial correction machinery.
A conjecture is public enough to attack. A proof can be checked line by line. Another mathematician can search for the hidden assumption, the missing case, the false equivalence, the easier route, the stronger theorem, or the older paper everyone somehow forgot.
Priority can help this field. Prizes can help it. Difficult benchmarks can help it. Even the deeply silly human desire to be the person who finally solved the thing can keep somebody inside a problem long enough for reality to answer.
Competition belongs inside inquiry as one of its instruments.
The instrument works while winning remains coupled to the larger field.
The competition earns its place because the contest produces more mathematical contact than the scoreboard alone contains.
Then artificial intelligence changes the coupling.
◆
Terence Tao's July 24 public lecture at the International Congress of Mathematicians begins by asking the audience to provisionally grant a strong assumption:
He does not spend the rest of the talk celebrating or mourning that possibility.
He asks what the mathematical community is actually optimizing.
His list is wider than problem solving:
Historically, these goals often moved together. Someone who solved an important problem usually had to understand enough mathematics to explain something, teach something, build something, or leave a technique other people could use. The solve-count could stand in as a rough proxy because other goods frequently arrived with it.
Tao's warning is Goodhart's law.
A system can increase the number of solved problems without increasing mathematical understanding at the same rate. It can generate correct proofs faster than the community can read them. It can optimize a class of benchmark-friendly problems while theory-building, pedagogy, taste, exposition, and the selection of genuinely important questions lag behind.
This is the degenerate metagame arriving in mathematical research.
A degenerate meta appears when an incentive system finds an exploit and begins compressing the field around it. Everyone can be playing correctly. The score can keep improving. The activity can still lose the structure that made the score worth pursuing.
Artificial intelligence does not create this problem.
Publication counts were already gameable. Citation counts were already gameable. Priority already rewarded speed. Universities already converted research into rankings. Grant systems already learned to demand legible output. Mathematics already possessed competitions whose local incentives could distort the wider inquiry.
Artificial intelligence changes the gradient.
Suddenly one of the easiest outputs to measure may become one of the cheapest outputs to generate.
That is the fracture.
The competition was tolerable while winning remained coupled to inquiry.
The new systems make it possible to win faster than the field can understand what winning produced.
◆
Tao's most useful contribution is a pipeline.
A mathematical result does not enter the field at the moment a candidate proof first appears.
It develops.
Tao argues that artificial intelligence accelerates the early stages much more aggressively than the later ones.
Proof generation moves first. Formal verification follows. Exposition can be assisted. Community acceptance remains slow. Canonicalization is slower still.
In his ICM lecture, Tao calls canonicalization the stage least amenable to artificial-intelligence optimization and the most valuable part of the whole process.
That sentence should end the fantasy that a theorem count is the same thing as mathematical progress.
A proof can be true before anybody understands its significance.
A proof can be verified before the literature around it has been reconstructed.
A proof can be beautifully typeset before anyone has identified the one idea worth carrying forward.
A proof can exist while the field still does not know what it has.
A proof can arrive before the mathematics arrives.
This is the new bottleneck.
Tao calls the emerging condition proof abundance and proof indigestion. The field has been organized for a world where producing a serious proof was scarce enough that verification, explanation, and digestion could usually gather around it. If generation becomes cheap, those downstream capacities become the scarce resource.
Modal Path Ethics can name the exported burden more generally:
A system produces a result quickly.
Then somebody else has to determine:
The producer can externalize that work onto the mathematical community.
At small scale, this is ordinary scholarship. Everyone depends on reviewers, readers, editors, seminar audiences, librarians, teachers, and future researchers.
At machine scale, the burden can change phase.
One lab can generate hundreds of pages at a tempo that requires dozens of specialists to inspect. Formal certificates lower one part of the burden while leaving several others untouched. The proof assistant can tell you that the encoded derivation closes under its rules. It cannot, by that fact alone, tell you whether the result was framed honestly, whether the right theorem was encoded, whether the literature was represented well, whether the result deserves attention, or what it teaches the field.
The Leiden Declaration on Artificial Intelligence and Mathematics is basically a constitutional attempt to protect those downstream functions before they disappear beneath output volume. It emphasizes attribution, independent verification, transparent arguments, evaluation of depth and significance, community understanding, and human responsibility for the result and its citations.
The declaration is defending the field around proof generation against being compressed into proof generation alone.
◆
OpenAI did something unusually responsible with the Astra work: it formalized each argument in Lean.
Formal verification is exactly the kind of instrument a proof-abundant world will need. If machines can generate candidate mathematics faster than humans can check every line, machine-checkable proofs can preserve a hard boundary against a flood of plausible nonsense.
The boundary is valuable precisely because it is narrow.
This is a direct continuation of the Mathematics Problem. Formal rigor begins after the cut. A formal proof can be immaculate while the larger selection remains confused.
Lean's jurisdiction ends there.
Field Instruments: The Scientific Method makes the parallel point for science. A method becomes trustworthy partly through correction paths: replication, criticism, open records, adversarial review, better instruments, and public repair. The discipline is powerful because it can answer its own earlier claims.
Proof assistants belong to that family. They strengthen correction.
The mistake arrives when the presence of a very strong correction instrument lets the producer skip the rest of the inquiry.
The computer can check the bridge.
The community still has to decide where the bridge goes.
◆
Tao now describes this as mathematics' first major foundational crisis in more than a century.
The early twentieth-century crisis concerned foundations in the familiar technical sense: sets, infinity, axioms, paradoxes, proof, consistency.
The new crisis concerns values and practices.
Tao's current public summary, updated August 6, sharpens this into an unusually urgent warning. He says mathematicians have only months to begin reimagining the profession, calls for organizing and activism rather than passive acceptance, and says the field has temporarily lost the narrative about what a mathematician is.
That language is revealing.
A technology company does not need to pass a law defining mathematics.
It can define the profession indirectly through demonstrated capability.
If every major public announcement says:
If the benchmark says:
If the press cycle says:
The field begins inheriting an external constitution through repetition.
This is an instrument-jurisdiction problem.
Technology companies have standing to say what their systems can do. They have standing to publish results, show evidence, build tools, and participate in mathematics once their systems genuinely participate in mathematics.
They do not acquire sovereignty over the definition of mathematics by becoming exceptionally good at one part of it, however.
The theorem generator does not get to define the theorem field.
The scoring instrument does not get to define the purpose of play.
The proof benchmark does not get to define mathematical life.
Tao's intervention is valuable here because he is not trying to defend a pre-AI profession as sacred territory. His positive vision includes human-machine collaboration, large-scale distributed mathematics, automation of routine work, population studies across huge families of problems, and new infrastructure built around proof abundance.
He expects the instrument to enter. He is arguing over the constitution it enters.
That is the correct fight.
◆
There is a funny timing problem here.
The Invisible Board was published on the same day Modal Path Ethics encountered this controversy.
That article moved from Go into mathematics through exactly the distinction now under pressure.
The visible symbols do not exhaust the active relation.
A proof is public correction of mathematical intuition. It makes a suspected structure inspectable outside the mind that found it.
That is powerful. Yet a proof still lives on a wider board.
It has history.
It has neighboring work.
It has a community capable of understanding it badly or well.
It has students who may inherit a technique without inheriting the path that made the technique intelligible.
It has institutional incentives capable of rewarding volume, priority, prestige, difficulty, novelty, beauty, usefulness, fashion, or whatever else the local scoring system can see.
Artificial intelligence changes the visible pieces on that board dramatically.
It may become a stronger problem solver than any individual human.
That does not erase the invisible board.
It makes the invisible board more important.
The question moves from:
into:
That is a Field Intelligence Gap question.
Artificial intelligence can become extraordinarily strategically intelligent inside mathematics. It can search, combine, calculate, formalize, test, and perhaps increasingly prove.
The field-intelligent question arrives one level up.
That is the board.
◆
There is no need to choose between artificial intelligence and mathematics.
There is a need to choose the relation.
The Better configuration is almost embarrassingly available already.
Let artificial intelligence search aggressively.
Let it formalize.
Let it test conjectures, generate counterexamples, explore long-tail problems, combine literatures, perform routine derivations, build candidate proofs, and expose structures no human would have found on the same schedule.
Then, preserve the rest of the field around that capability.
Tao has even begun pushing priority in this direction. In his current account, the race should move away from who can make an artificial-intelligence system produce the first raw proof and toward who can provide a genuinely public scientific exposition of the result.
That is a profound change in the win condition.
The first person to reach the destination is less important than the first person who can build a road other mathematicians can actually use.
Competition survives, now answerable to inquiry again.
The same move works at larger scale.
Artificial intelligence companies can compete to build stronger systems.
Mathematicians can compete to solve hard problems.
Journals can compete for excellent work.
Students can compete at Olympiads.
None of those contests needs to become the constitution of mathematics.
The larger game is cooperative.
Everyone is trying to gain better contact with structure that no participant owns.
A theorem cannot be conquered.
It can be found.
Then understood.
Then connected.
Then taught.
Then used to find something else.
◆
OpenAI's Astra results are impressive.
The strongest criticism does not require pretending otherwise.
If the model generated genuinely new, correct mathematical arguments, then artificial intelligence has crossed another important threshold in research mathematics. Formal verification makes the event stronger. The ability to produce ten substantial results at low marginal inference cost changes what kinds of mathematical search are reachable.
The complaints about attribution are also serious.
A correct proof does not erase the path that made it possible. Citation preserves mathematical memory. Literature review keeps the discovery connected to the field that generated it. Priority without provenance turns inquiry into a race whose finish line moves faster than its mapmakers.
Terence Tao sees the larger fracture.
For a long time, mathematical competition and mathematical inquiry were coupled tightly enough that winning one contest often advanced several other goods at once. Solve the hard problem. Build the theory. Train the mathematician. Explain the technique. Add to the literature. Give the next generation a stronger starting point.
Artificial intelligence can separate those goods.
A lab can win the theorem race while the community inherits verification load, citation reconstruction, exposition work, review pressure, and canonicalization debt.
A model can generate the answer while other people still have to discover what the answer means.
The model can enlarge mathematics while exposing the scoreboard's insufficiency.
The theorem does not care who won.
The theorem does not know OpenAI from UCLA. It does not know whether a human spent twelve years finding the bridge or a model searched the relevant space before dinner. It does not award moral points for suffering. It does not protect anybody's professional identity.
The relation either holds or it does not.
Inquiry cares about the path.
It cares because later understanding depends on knowing where the result came from, what it changed, which instruments survived correction, who can explain it, what it connects, what remains open, and whether the field became more capable of continuing after the discovery.
Competition can serve that process.
Competition cannot own it.
If artificial intelligence becomes capable of winning the mathematical contest, mathematicians have been handed an unusually sharp opportunity to ask whether the contest was ever the purpose.
That answer will not be found on the leaderboard.
The leaderboard is one of the things now under review.